A Robust Data-Driven Supervised Ensemble Machine Learning Framework for Soil Fertility Classification in Precision Agriculture Using Ensemble Boosting Models |
Author(s): |
| Gautam Palash , SIRT, Bhopal; Dr. Mohit Singh Tomar, SIRT, Bhopal |
Keywords: |
| Precision Agriculture; Soil Fertility Classification; Gradient Boosting; XGBoost; LightGBM; CatBoost; SMOTE; Class Imbalance; Data Leakage; Permutation Test; Reproducibility; Explainable AI |
Abstract |
|
Soil fertility assessment is central to precision agriculture, yet conventional laboratory analysis is slow, costly, and difficult to scale across the many smallholdings that dominate agriculture in developing economies. Machine learning, and gradient-boosting ensembles in particular, have been proposed as a rapid, data-driven alternative for classifying fertility from routinely measured physicochemical attributes. This paper develops and rigorously evaluates a supervised ensemble learning framework for three-class (Low, Medium, High) soil fertility classification, benchmarking XGBoost, LightGBM, and CatBoost against five classical baselines and a soft-voting ensemble on a public 880-sample, fifteen-feature dataset. The pipeline enforces a leakage-free protocol in which stratified train-test partitioning precedes feature standardisation and Synthetic Minority Over-sampling Technique (SMOTE), the latter fitted exclusively on the training partition to counter a severe imbalance in which the High-fertility class represents only 5.6% of samples. Beyond conventional accuracy, precision, recall, F1, Cohen's kappa, and Matthew's correlation coefficient, the framework introduces a dataset-integrity diagnostic layer combining feature-target correlation analysis, mutual-information estimation, and a label-permutation test. The boosting ensembles, led by CatBoost (46.6% test accuracy) and LightGBM (45.5%), outperformed the classical baselines; however, no model exceeded the 47.7% majority-class baseline, kappa and Matthews correlation coefficients clustered near zero, and cross-validated accuracy on true labels (46.1%) was statistically indistinguishable from accuracy on randomly permuted labels (45.6%). Every model, including the SMOTE-balanced pipeline, achieved zero recall on the High-fertility class. These convergent diagnostics indicate that the dataset's fifteen physicochemical features carry negligible information about the fertility label, a conclusion that reframes the study's principal contribution: not a high-accuracy classifier, but a transparent, reproducible protocol for distinguishing genuine predictive skill from the leakage-inflated accuracies (often exceeding 90%) reported elsewhere in the literature. The paper argues that model sophistication cannot substitute for data quality, and proposes the described diagnostic suite, together with a deployment blueprint spanning web, mobile, and IoT integration, as a template for responsible machine learning in precision agriculture. |
Other Details |
|
Paper ID: IJSRDV14I60006 Published in: Volume : 14, Issue : 6 Publication Date: 01/09/2026 Page(s): 33-42 |
Article Preview |
|
|
|
|
