A Rasch-based analysis comparing the performance of updated multimodal large language models and surgeons on the Japanese surgical specialist examination.

Journal: Surgery today
Published Date:

Abstract

PURPOSE: To compare four multimodal large language models (LLMs) with surgeons on the 2023 Japanese Surgical Specialist Examination using item-level surgeon correct answer rates as the benchmark. METHODS: In this retrospective cross-sectional study, GPT-4.1, Claude Opus 4, Gemini 2.5 Pro, and o3 Pro were evaluated using 98 valid multiple-choice items, including 43 image-based and 55 text-only questions. The accuracy was examined overall, by image presence, and by subspecialty. Physician-anchored Rasch modeling placed LLMs and surgeons on a common latent scale, and logistic regression assessed the association between surgeon accuracy and LLM correctness at the item-level. RESULTS: The accuracy ranged from 77.6% for GPT-4.1 to 85.7% for o3 Pro. All models showed lower accuracy on image-based items than on text-only items. A Rasch analysis showed that all LLMs remained below the surgeons' overall, with relative abilities ranging from - 1.29 to - 0.74. Performance varied according to difficulty and subspecialty. Gastroenterology was consistently the weakest domain, whereas some models matched or exceeded the surgeon benchmark in selected areas, including respiratory, pediatrics, breast/endocrine, and emergency/anesthesiology. CONCLUSIONS: For this previously analyzed item set, updated multimodal LLMs achieved an overall accuracy below the surgeon benchmark, with the largest deficits on image-based items, supporting supervised and domain-specific use.

Authors

Keywords

No keywords available for this article.