Assessing the Third Wave of Generative AI: Performance of Advanced Models on Text-based Questions From the 2024 Orthopaedic In-training Examination.
Journal:
The Journal of the American Academy of Orthopaedic Surgeons
Published Date:
Dec 22, 2025
Abstract
INTRODUCTION: Learners are rapidly using generative artificial intelligence (AI) models in their education. We assessed the performance of recently released or updated models on the 2024 American Academy of Orthopaedic Surgeons Orthopaedic In-training Examination for their potential applications in orthopaedic education. METHODS: Eleven models with recent enhancements in reasoning and research capabilities were evaluated. A total of 119 text-based questions were entered verbatim into each model. Model outputs were recorded as correct or incorrect. Additional analyses included reasoning time, citation accuracy, confidence in answer selection, and comparison with orthopaedic resident performance. References generated by the top performing model were compared with American Academy of Orthopaedic Surgeon Recommended Readings for incorrectly answered questions. RESULTS: Ten of 11 AI models exceeded the American Board of Orthopaedic Surgery minimal passing standard (67.7%). Eight models surpassed PGY5 resident performance. OpenAI's o1Pro with Deep Research achieved the highest accuracy (90.8%), outperforming the mean performance of PGY5 residents by 17.8%. More than half of the ResStudy Recommended Readings were cited as supporting references by the top performing model on questions it answered incorrectly. Subspecialty performance varied, with highest accuracy in Shoulder and Elbow and Sports Medicine questions. Longer reasoning times generally correlated with improved accuracy. DISCUSSION: Advanced AI models demonstrated substantial improvements over previous generations, with more than half of the tested models exceeding senior orthopaedic resident performance. Improved reasoning and research capabilities highlight the evolving capabilities of AI in medical education, although increased understanding of their utilization by learners is needed. Variability among subspecialties may suggest differences in training data or reasoning capabilities. CONCLUSIONS: Modern AI models exhibit high proficiency on the Orthopaedic In-training Examination and may serve as valuable supplemental educational tools. Ongoing evaluation is warranted to understand their optimal integration into orthopaedic training while recognizing limitations in clinical reasoning, lived experience, and imaging interpretation.
Authors
Keywords
No keywords available for this article.