Diagnostic Ambiguity in Appendicitis: Comparing Failure Modes of Large Language Models and Surgeons.
Journal:
The American surgeon
Published Date:
Jul 25, 2026
Abstract
BackgroundSuspected acute appendicitis is difficult to triage in atypical presentations, including pediatric and geriatric cases and scenarios with muted inflammatory markers. Frontier LLMs are increasingly evaluated in emergency decision-making, but vignette-based comparisons should not be interpreted as evidence of clinical readiness. This study evaluated diagnostic behavior, triage safety, and error profiles of frontier LLMs in standardized appendicitis scenarios, using general surgery specialists as a benchmark.MethodsWe performed a comparative, cross-sectional diagnostic-accuracy simulation study using 150 standardized emergency department vignettes (90 appendicitis and 60 non-surgical mimics), enriched for diagnostically difficult scenarios. Six frontier AI systems and two board-certified general surgery specialists independently evaluated each vignette using a fixed triage prompt. Outcomes were diagnostic accuracy, sensitivity, specificity, inter-rater agreement, triage-priority accuracy, and qualitative error taxonomy.ResultsOverall diagnostic performance differed across cohorts (P < 0.001). GPT-5 achieved the highest accuracy (95.3%), exceeding GPT-4o (82.0%; P < 0.001) and slightly surpassing specialist consensus (92.7%). In pediatric and geriatric vignettes, GPT-5 maintained 93.8% accuracy, compared with 90.0% for specialists and 71.2% for GPT-4o. Error analysis identified distinct failure modes, including search-induced hallucinations in web-augmented systems. These findings were interpreted alongside clinical error severity rather than as product-level superiority.ConclusionIn this vignette-based simulation study, LLM performance varied substantially across systems. The main contribution was characterization of model and human failure modes under diagnostic ambiguity, supporting supervised, diagnosis-specific evaluation rather than autonomous emergency triage deployment.
Authors
Keywords
No keywords available for this article.