Deep Learning Versus LASSO-Based Machine Learning for 1-Year Survival After Surgery for Spinal Metastases: A JASA Multicenter Prospective Cohort Analysis.

Journal: Spine
Published Date:

Abstract

STUDY DESIGN: Multicenter prospective cohort study; secondary analysis. OBJECTIVE: To evaluate predictors associated with 1-year survival after surgery for spinal metastases by comparing a comprehensive 50-variable deep learning (DL) model with a previously published 5-variable LASSO-based machine learning (ML) model and applying DL-based permutation feature importance as an exploratory analytic lens. SUMMARY OF BACKGROUND DATA: Surgical decision-making for spinal metastases requires reliable survival estimates. Traditional scores such as those of Tokuhashi and Tomita and contemporary tools such as SORG and NESMS support prognostication, but performance and calibration may vary across cohorts. A parsimonious 5-variable JASA ML model is clinically practical, whereas DL may help identify prognostic signals embedded in detailed activities of daily living (ADLs), patient-reported outcomes (PROs), and scoring-system components. METHODS: We analyzed 401 complete-case patients who underwent surgery for spinal metastases at 35 Japanese institutions (2018-2021). A feed-forward neural network incorporating 50 preoperative variables was evaluated using five repeated random 8:2 train-test splits. Accuracy, AUROC, Brier score, and calibration summaries were reported and descriptively compared with the previously published 5-variable LASSO-based ML model. RESULTS: At 1 year, 269 of 401 patients were alive. The DL model achieved 75.5+/- 3.0% accuracy (95% confidence interval [CI], 71.8%-79.2%), held-out AUROC 0.789 (95% CI, 0.681-0.886), and Brier score 0.214. The ML model achieved 71.8% accuracy (Wilson 95% CI, 67.2%-76.0%) and apparent AUROC 0.762. Because the comparison was descriptive rather than paired, formal statistical superiority was not claimed. DL feature importance highlighted Vitality Index-On and Off Toilet, EQ-5D-5L total score and pain/discomfort, and individual Tokuhashi/Tomita components; the ML-selected Vitality Index-Wake Up item ranked 38th. CONCLUSIONS: The 50-variable DL model provided reasonable prediction and generated clinically plausible feature-importance hypotheses, but it did not demonstrate a clearly meaningful performance advantage over the simpler 5-variable ML model. DL may be most useful for research-based feature discovery and refinement of future parsimonious prognostic tools, whereas validated simple models remain more practical for bedside prognostication. LEVEL OF EVIDENCE: 2.

Authors

Keywords

No keywords available for this article.