Traditional Machine Learning versus Deep Learning Approaches for Multi-Outcome Perioperative Risk Prediction
Journal:
medRxiv
Published Date:
Sep 7, 2026
Abstract
Perioperative risk tools such as the ACS NSQIP Surgical Risk Calculator rely on structured variables, but whether newer modeling approaches improve prediction is unclear. Holding 69 clinician-reviewed preoperative variables fixed, we benchmarked six tabular model families and an LLM-based embedding classifier for four 30-day postoperative outcomes, training on 4,995,670 ACS NSQIP cases from 2018-2022 and evaluating temporally on 963,565 cases from 2024, with 2023 excluded because outcome windows cross the year boundary. The FT-Transformer achieved the highest AUROC on all four outcomes (0.755-0.955) and logistic regression the lowest, separated by only 0.010-0.030; the multilayer perceptron and FT-Transformer were statistically indistinguishable, and no architecture led on both AUROC and AUPRC (0.095-0.461). Among three LLM-based embedding classifiers evaluated on a sampled subset of the same test cohort, discrimination was comparable to the tabular models but did not exceed them (AUROC 0.736-0.945, AUPRC 0.087-0.437).