Early risk prediction of vitamin deficiency disorders by an interpretable machine learning model with varying levels of data abstraction.
Journal:
Nutrition research (New York, N.Y.)
Published Date:
Jun 1, 2026
Abstract
Nutritional deficiencies are a global health challenge that often remain undetected until clinical symptoms or biochemical abnormalities become pronounced. Early identification of individuals at moderate nutritional risk is essential for timely intervention and disease prevention. Detailed feature impact analysis and evaluation of nonbiochemical biomarkers remain open research questions. This study formulates and evaluates a three-class nutritional risk classification task using a Light Gradient Boosting-based machine learning framework across three feature groups: (1) full clinical model A trained on all features, (2) model B trained on features except biochemical, and (3) model C trained on features except biochemical, symptom details. Experimental results based on balanced accuracy, macro F1 score, and ROC-AUC reveal that (1) model A achieved 0.9475, 0.9329, and 0.9952, (2) model B achieved 0.9190, 0.8845, and 0.9851, (3) model C achieved 0.9111, 0.8574, and 0.9781, respectively. Comparison with baseline machine learning models revealed the supremacy of the proposed LightGBM model. Explainable AI analysis using SHapley Additive exPlanations (SHAP) supported experimental results. It revealed that biochemical markers are necessary for high-risk classification, while noninvasive parameters enable identification of moderate nutritional risk. The proposed framework demonstrates that machine learning models can predict nutritional risks from varying data abstraction levels of input information. Specifically, outcomes from the noninvasive model indicate its suitability for community health screening tasks where laboratory infrastructure is constrained. However, a minimal model is suitable for nutritional risk screening from telehealth applications. Hypothesis: We hypothesized that nutritional risk can be predicted from nonbiochemical features.
Authors
Keywords
No keywords available for this article.