Traditional Machine Learning versus Deep Learning Approaches for Multi-Outcome Perioperative Risk Prediction

Journal: medRxiv
Published Date:

Abstract

Perioperative risk tools such as the ACS NSQIP Surgical Risk Calculator rely on structured variables, but whether newer modeling approaches improve prediction is unclear. Holding 69 clinician-reviewed preoperative variables fixed, we benchmarked six tabular model families and an LLM-based embedding classifier for four 30-day postoperative outcomes, training on 4,995,670 ACS NSQIP cases from 2018-2022 and evaluating temporally on 963,565 cases from 2024, with 2023 excluded because outcome windows cross the year boundary. The FT-Transformer achieved the highest AUROC on all four outcomes (0.755-0.955) and logistic regression the lowest, separated by only 0.010-0.030; the multilayer perceptron and FT-Transformer were statistically indistinguishable, and no architecture led on both AUROC and AUPRC (0.095-0.461). Among three LLM-based embedding classifiers evaluated on a sampled subset of the same test cohort, discrimination was comparable to the tabular models but did not exceed them (AUROC 0.736-0.945, AUPRC 0.087-0.437).

Authors

  • Ding
  • Z.; Han
  • G.; Ma
  • T.; Wang
  • F.; Hajagos
  • J.; Kurc
  • T.; Kumar
  • A. R.

Categories