Explainable TabNet for gestational diabetes prediction with physician-in-the-loop and multi-site clinical validation.
Journal:
International journal of medical informatics
Published Date:
May 16, 2026
Abstract
BACKGROUND: Gestational diabetes mellitus (GDM) affects 15-25% of pregnancies worldwide and poses serious risks of macrosomia, preeclampsia, neonatal hypoglycaemia, and long-term type 2 diabetes. Existing machine learning models lack prospective multi-site external validation and formal physician trust evaluation, limiting real-world applicability. OBJECTIVES: To develop a clinically validated, explainable deep learning framework for GDM prediction using routinely available first-antenatal-visit clinical features, and to evaluate clinical readiness through dual-stage physician-in-the-loop (PITL) validation. METHODS: A TabNet binary classifier was developed on 3,525 clinical records using a three-stage feature-tailored hybrid imputation strategy (GAIN for HDL and OGTT; MissForest for Systolic BP; Mean for BMI). To prevent data leakage, SMOTE-based class balancing was applied exclusively within the training folds of a 5-fold stratified cross-validation pipeline, with validation folds remaining untouched. Explainability was delivered through TabNet intrinsic feature masks, SHAP, and LIME. Two-stage clinical validation comprised: (1) blinded PITL review by four certified obstetricians evaluating 30 patient cases with XAI explanations; and (2) prospective external validation across three independent Kerala hospitals totaling 80 patients. RESULTS: The proposed TabNet model achieved 97.13% accuracy, 94.05% precision, 98.91% recall, and 96.22% F1-score, outperforming ten baseline classifiers including Random Forest, XGBoost, and SVM under identical preprocessing conditions. Compared to recent state-of-the-art GDM prediction studies, the proposed model consistently outperformed comparable methods-under a rigorous 5-fold cross-validation strategy with confidence intervals, while most existing studies rely on single train-test splits without cross-validation. PITL validation yielded 96.7% concordance, an average Cohen's kappa of 0.909, and Fleiss' kappa of 0.963, with no prior GDM study reporting such formal physician endorsement. External multi-site F1 scores ranged from 83.70% to 87.00% across all three hospitals, reflecting an expected performance reduction in prospective real-world data, partly attributed to inter-site variability in feature availability and clinical data recording protocols. SHAP analysis identified a strong model-level interaction between PCOS and prediabetes as the dominant combined GDM risk signals, independently corroborated by all four obstetricians. CONCLUSION: The proposed framework integrates explainable deep learning with prospective dual-stage clinical validation, demonstrating promising performance as a clinically oriented proof-of-concept for the assessment of risk of GDM using routine clinical variables. TRIAL REGISTRATION: Clinical Trials Registry India, CTRI/2024/08/073158.
Authors
Keywords
No keywords available for this article.