Depression prediction and key factors: A comparative analysis of logistic regression and machine learning models.
Journal:
PloS one
Published Date:
Aug 27, 2026
Abstract
OBJECTIVES: Factors associated with depression were explored in this study through logistic regression, and predictive performance was compared with various Machine Learning models. METHODS: WHO SAGE India wave 2 data were used with depression as the outcome variable. and predictors were sociodemographic, health, and psychosocial variables. Descriptive analysis and Logistic regression were estimated. Random Forest, XGBoost, Support Vector Machine, Logistic Regression, Bagging, Decision Tree, Naïve Bayes, Ridge Logistic Regression, Neural Networks, and K Nearest Neighbors are the ten Machine Learning algorithms that were used. Performance measures consisted of accuracy, Area Under Curve, precision, recall, F1 score, Hamming loss, Jaccard score, and Matthew's correlation coefficient. Random Forest and XGBoost were used to assess feature importance. RESULTS: Depression was also more prevalent among younger adults, women, and individuals with poor self-rated health, stress, and sleep disturbances. Logistic regression revealed age and feeling low or sad as a factor (p = 0.008, p = 0.021). Most models demonstrated only moderate discriminative ability, with the AUC below 0.70, with better-performing models being Ridge regression (AUC = 0.716) and Random Forest (AUC = 0.713). Feature importance universally identified age, perception of health, quality of life, and depressive symptoms as important predictors. CONCLUSIONS: Logistic regression provides interpretability, and Machine Learning increases predictive accuracy. Combining both can enhance depression prediction and screening in public health practice.
Authors
Keywords
No keywords available for this article.