Stunting Prediction Without Derived Diagnostic Variables: A Comparison Of RandomForest, Lightgbm, Xgboost, And Logistic Regression

Authors

  • Muhammad Resha Universitas Teknologi AKBA Makassar, Makassar, South Sulawesi, Indonesia Author
  • Apriana Toding Department of Electrical Engineering, Universitas Kristen Indonesia Paulus, South Sulawesi, Indonesia Author

Keywords:

Anthropometry, machine learning, stunting, Z Score

Abstract

Z-score-based stunting screening requires reference standards and calculation procedures that may constrain the efficiency of community health services. Evidence remains limited on whether machine-learning models can classify stunting using only basic anthropometric measurements without derived diagnostic variables. This study developed and evaluated a Z-score-free, leakage-aware classification framework using age, sex, weight, and height. A total of 40,071 records of children aged 0–59 months from Jeneponto Regency, Indonesia, collected between 2021 and 2024 were analyzed. XGBoost, Random Forest, LightGBM, and Logistic Regression were compared using a stratified 80:20 hold-out split and stratified five-fold cross validation.
Median imputation and SMOTE were applied exclusively to the training data. In the hold-out evaluation, Random Forest achieved 95.81% accuracy, a 94.86% F1-score, an MCC of 0.913, and a Cohen’s kappa of 0.913, whereas LightGBM achieved the highest recall
(95.83%) and ROC-AUC (99.28%). Cross-validation produced a consistent pattern, with Random Forest obtaining the highest mean F1-score (94.84% ± 0.34). SHAP analysis of XGBoost identified height and age as the dominant contributors to prediction. These findings indicate that tree-based models may support preliminary screening without using Z-scores as predictors. External and temporal validation is required before broader implementation in community health posts.

Downloads

Published

2026-06-30