Large Language Models Outperform PGY-5 Residents on the Orthopaedic In-Training Examination: A Comparative Analysis of Six Cutting-Edge Large Language Models.

Journal: The Journal of the American Academy of Orthopaedic Surgeons
Published Date:

Abstract

INTRODUCTION: Large language models (LLMs), such as ChatGPT, are becoming increasingly prevalent, particularly in medical education and clinical assessments. Previous LLMs were seen to perform at the level of a first-year resident on the 2022 Orthopaedic In-Training Examination (OITE). With exponential advances in LLMs over the past 3 years, the true capabilities of these models remain unexplored. In addition, the addition of image processing further increases their clinical applicability. The purpose of this study was to evaluate the performance of six LLMs on the 2024 OITE. METHODS: Six LLMs were evaluated in this study: ChatGPT (GPT-4o), Gemini 2.0 Flash, Grok 3, Mistral Large 2.7, DeepSeek R1, and Llama. ChatGPT, Gemini, Grok, and Mistral could evaluate images and text while DeepSeek and Llama were limited to text. Accuracy, image interpretation, and logical consistency were assessed in 203 multiple-choice questions, stratified by difficulty and type of the question. Statistical analyses involved chi-square tests, Fisher exact tests, z-tests, and Cohen κ tests. RESULTS: ChatGPT performed with the highest accuracy (74.9%), followed by DeepSeek, Llama, Grok, Mistral, and Gemini. ChatGPT also led in logical consistency (72.4%) and image interpretation (73.8%). Logical consistency strongly correlated with accuracy and correctness ( P < 0.00001). As difficulty increased, performance declined across all models. CONCLUSION: ChatGPT consistently scored the highest in terms of accuracy across all metrics while also maintaining reasoning quality. Compared with resident averages, ChatGPT performed at a postgraduate year five level which indicates its potential for integration into orthopaedic clinics, electronic medical records, and surgical planning. Further development models would allow for better performance on difficult questions and creating orthopaedic focused models could enhance these results.

Authors

Keywords

No keywords available for this article.