Large Language Models as a Clinical Interface for Predictive Modeling: A Feasibility Study in Pediatric Appendicitis.
Journal:
The American surgeon
Published Date:
Oct 9, 2026
Abstract
PurposeLarge language models (LLMs) are increasingly used in clinical data workflows, but most clinicians lack the time, training, or support to build predictive models. We evaluated whether a ChatGPT-assisted workflow could enable clinicians to independently construct and execute a basic predictive model using routine perioperative data, using postoperative length of stay (LOS) in pediatric appendicitis as a feasibility use case.MethodsWe conducted a retrospective study of 228 children undergoing laparoscopic appendectomy. Ten routinely available preoperative variables were used. ChatGPT guided preprocessing and generated code for linear regression and random forest models using naïve and structured prompting. Models were trained using an 80/20 split and evaluated using MAE, RMSE, r, and R2.ResultsNaïve prompting performed poorly, while structured prompting improved model performance (MAE 0.34, RMSE 0.85, r 0.85). On the test set, linear regression achieved MAE 1.00 and random forest 0.77. Identified variables reflected expected clinical patterns. These findings demonstrate internal consistency of the workflow rather than new predictive insight.ConclusionsA structured ChatGPT workflow enabled clinicians to construct standard predictive models using routine data. These findings support feasibility of clinician-directed model development, but not clinical utility, of LLM-enabled workflows.
Authors
Keywords
No keywords available for this article.