Are artificial intelligence systems ready for pediatric surgical decision-making? A comparative evaluation of large language models versus pediatric surgeons.

Journal: Pediatric surgery international
Published Date:

Abstract

PURPOSE: This study aimed to compare the diagnostic, investigative, and treatment decisions generated by five contemporary large language models (LLMs) across 50 pediatric surgery clinical scenarios against a reference standard established by pediatric surgery experts. METHODS: Fifty guideline-based clinical scenarios representing a broad range of pediatric surgical subspecialties were developed. Diagnosis, investigation, and treatment decisions were assessed for each scenario. A consensus reference standard was established by two pediatric surgeons. ChatGPT (GPT-5.5), Claude (Sonnet 4.6), Gemini (3.5 Flash), DeepSeek-V3, and Perplexity (1.0) were evaluated independently using standardized prompts. Accuracy, observed agreement (Po), Gwet's AC1 with 95% confidence intervals, and McNemar tests were used to evaluate model performance and agreement with the pediatric surgeon. RESULTS: Overall, the pediatric surgeon correctly classified 144 of 150 decisions (Po = 0.960). Claude achieved the highest overall performance (145/150; Po = 0.967), followed by ChatGPT and Gemini (144/150 each; Po = 0.960). DeepSeek-V3 and Perplexity achieved Po values of 0.913 and 0.873, respectively. Claude achieved perfect diagnostic performance (50/50; Po = 1.000). ChatGPT and Gemini performed best for investigations (48/50; Po = 0.960), while Claude, ChatGPT, and Gemini each achieved 48/50 correct treatment decisions (Po = 0.960). No significant differences were identified between the pediatric surgeon and any AI model with respect to diagnostic, investigative, or treatment performance (all p > 0.05). In the overall analysis, only Perplexity demonstrated a performance level significantly different from that of the pediatric surgeon (p = 0.004). CONCLUSIONS: LLMs showed high accuracy and strong agreement with expert judgment under standardized guideline-based clinical scenarios, particularly Claude, ChatGPT, and Gemini. The clustering of errors within the investigation and treatment domains suggests that, although these models show considerable promise as clinical decision-support tools, they should be used under clinician supervision rather than as independent decision-makers. CLINICAL TRIAL REGISTRATION: Not applicable.

Authors

Keywords

No keywords available for this article.