MACHINE LEARNING VERSUS PENALISED LOGISTIC REGRESSION FOR PREDICTING IN-HOSPITAL MORTALITY IN RUPTURED ABDOMINAL AORTIC ANEURYSM: A COMPARISON USING A UNIFORM VALIDATION FRAMEWORK.
Journal:
Annals of vascular surgery
Published Date:
Aug 19, 2026
Abstract
BACKGROUND: Ruptured abdominal aortic aneurysm (rAAA) remains associated with substantial in-hospital mortality. Although machine-learning methods can model complex nonlinear relationships between admission characteristics and outcome, their incremental value over conventional regression remains uncertain. We compared three prediction models using identical admission variables and a uniform validation framework. METHODS: This retrospective single-centre cohort included 196 unique patients with rAAA managed between 2006 and 2024. Five prespecified admission predictors were evaluated: age, hypovolaemic shock, maximal aneurysm diameter, haemoglobin and systolic blood pressure. Missing data were imputed independently within each training partition. Penalised logistic regression, gradient boosting and a multilayer perceptron (MLP) were compared using nested five-fold cross-validation repeated five times. Secondary temporal validation was performed in the 181 patients with a recoverable treatment year: models were developed in the 2006-2018 cohort (n=110) and evaluated, without refitting or recalibration, in the 2019-2024 cohort (n=71). RESULTS: In-hospital mortality occurred in 99/196 patients (50.5%). Gradient boosting and penalised logistic regression achieved comparable moderate discrimination, with AUCs of 0.739 (95% CI 0.666-0.805) and 0.729 (95% CI 0.654-0.796), respectively. Their paired AUC difference was 0.010 (95% CI -0.024 to 0.045). The MLP showed lower discrimination (AUC 0.643, 95% CI 0.563-0.719), with a paired difference versus penalised logistic regression of -0.086 (95% CI -0.160 to -0.013). In temporal validation, AUCs were 0.800 (95% CI 0.686-0.904) for gradient boosting, 0.739 (95% CI 0.607-0.858) for penalised logistic regression and 0.632 (95% CI 0.496-0.762) for the MLP. Temporal calibration demonstrated overprediction of absolute mortality risk, with observed mortality of 40.8% compared with mean predicted mortality of 49.9%, 50.0%, and 63.4%, respectively. CONCLUSIONS: Using a uniform internal and temporal validation framework, gradient boosting achieved the highest numerical performance but did not demonstrate a conclusive advantage over penalised logistic regression. The MLP provided no incremental predictive benefit. These findings indicate that increasing algorithmic complexity does not necessarily improve mortality prediction in modest-sized emergency vascular datasets and support interpretable regression as an essential benchmark. External multicentre validation and more complete prospective data collection are required before clinical implementation.
Authors
Keywords
No keywords available for this article.