Evaluating the performance of general vs retrieval-augmented generation large language models on oral and maxillofacial surgery board examinations.

Journal: International journal of oral and maxillofacial surgery
Published Date:

Abstract

Large language models have shown remarkable performance on medical examinations, but their application in oral and maxillofacial surgery remains under-investigated. This study evaluated whether general-purpose models outperform domain-specific retrieval-augmented generation (RAG) models in specialized medical assessments. Five AI models were tested on official Phase A board examinations (2017-2024, 788 questions): three general-purpose models (Gemini 2.5, ChatGPT-5, Claude Sonnet 4) and two GPT-4o variants (non-tuned baseline and RAG-augmented with complete access to official examination materials). Gemini 2.5 achieved 83.8% accuracy, ChatGPT-5 73.2%, Claude Sonnet 4 72.3%, GPT-4o non-tuned 51.4%, and GPT-4o with RAG 50.9%. Pairwise McNemar's tests showed significant differences between Gemini 2.5 and both ChatGPT-5 (P < 0.001) and Claude Sonnet 4 (P < 0.001), while ChatGPT-5 and Claude Sonnet 4 did not differ (P = 0.609). All general-purpose models significantly outperformed both GPT-4o variants (all P < 0.001). Direct comparison between GPT-4o variants revealed no difference (P = 0.788), indicating that RAG with direct access to examination materials provided no benefit over baseline architecture. Error analysis revealed complementary failure patterns: humans struggled with factual recall, while AI struggled with contextual integration within regional medical curricula. Gemini 2.5 significantly exceeded the 70% passing threshold (P < 0.001); ChatGPT-5 marginally exceeded it (P = 0.048); Claude Sonnet 4 achieved a passing score without statistically exceeding the threshold (P = 0.153). Architectural sophistication, rather than domain-specific knowledge augmentation, emerged as the primary performance determinant. These findings reflect examination performance on text-based multiple-choice questions and should not be extrapolated to real clinical decision-making.

Authors

Keywords

No keywords available for this article.