Explainable machine learning for predicting childhood anemia in Sub-Saharan Africa using population-based DHS Data (2016-2024).
Journal:
PLOS global public health
Published Date:
Aug 12, 2026
Abstract
Childhood anemia remains a major public health challenge in Sub-Saharan Africa, adversely affecting physical growth, cognitive development, and child survival. The pooled prevalence across 26 countries exceeds 60%, underscoring the need for accurate and scalable prediction tools to support targeted interventions.This study used pooled Demographic and Health Survey (DHS) data (2016-2024) from 26 Sub-Saharan African countries, including 110,251 children aged 6-59 months. Multiple machine learning models (Logistic Regression, Decision Tree, Extra Trees, Random Forest, XGBoost, LightGBM, CatBoost, and MLP) were trained using hyperparameter tuning with 5-fold cross-validation. Model performance was evaluated using accuracy, precision, recall, F1-score, and ROC-AUC, with 95% confidence intervals estimated via bootstrapping. Model comparisons were conducted using the DeLong test, and SHAP was used for model interpretability.The CatBoost model demonstrated the best overall performance (ROC-AUC = 0.84, 95% CI: 0.840-0.848), followed closely by XGBoost and LightGBM. All machine learning models significantly outperformed logistic regression (p < 0.001, DeLong test), although the absolute improvement in discrimination was modest (ΔAUC ≈ 0.04). The models demonstrated moderate discriminatory ability, with relatively low to moderate sensitivity (recall ≈ 0.33-0.41) depending on the model and classification threshold. SHAP analysis identified residence type, height-for-age z-score, country, and child age as the most influential predictors.Machine learning models demonstrated moderate to strong discriminatory performance in predicting childhood anemia using DHS data. Although improvements over traditional models were statistically significant, the relatively low sensitivity limits their effectiveness as standalone screening tools. These findings highlight the importance of early childhood nutrition (particularly during 6-23 months), reduction of chronic undernutrition, and context-specific public health strategies. Integration of predictive analytics into national health systems may support risk stratification and resource allocation in high-burden settings. However, findings should be interpreted in light of the cross-sectional design and lack of external validation.
Authors
Keywords
No keywords available for this article.