Benchmarking GPT-5, Gemini 2.5 Pro, Grok 4, and other LLMs on pediatric dentistry questions from a dental specialization exam.

Journal: Scientific reports
Published Date:

Abstract

Artificial intelligence (AI), particularly large language models (LLMs), is an increasingly prominent tool in medical and dental education. Trained with deep learning and NLP techniques, these models interpret meaning, generate text, and manage complex information. They hold potential for practical educational applications, such as supporting exam preparation and personalized learning. Moreover, their performance in clinical case recognition suggests an emerging potential for use in diagnostic decision-support systems. This study aimed to evaluate the performance of state-of-the-art large language models (LLMs) on pediatric dentistry questions from the Dentistry Specialization Examination (DUS) in Türkiye, a high-stakes national exam for postgraduate training. A total of 119 pediatric dentistry questions from the past ten years of the DUS were compiled and presented to 11 recently developed LLMs (17 with reasoning mode activations), including GPT-5, Gemini 2.5 Pro, Grok-4, and DeepSeek R1. Each model's accuracy (%) and average response generation time (seconds) were calculated and compared. The Gemini 2.5 Pro (92.44%) demonstrated significantly higher mean scores compared to all other models except GPT-4 (78.15%), GPT-5 (90.76%), GPTOSS (78.99, 75.63%), and Grok-4 (88.24%). Similar patterns were also observed for GPT-5 and Grok-4. In contrast, Qwen-3 (49.58 - Reasoning, 54.62 - No Reasoning), and MedGemma (58.82%) exhibited notably lower accuracy rates across most comparisons. Overall, these findings highlight Gemini 2.5 Pro and GPT-5 as achieving the highest accuracy levels among the models, whereas Qwen-3, Qwen-3 (R), and MedGemma demonstrated the weakest performance. While models such as Gemma, LLaMA, and Mistral demonstrated faster response times (<1 s), they exhibited relatively low accuracy. In contrast, reasoning-intensive modes (e.g., DeepSeek R1) improved accuracy but required excessively long generation times (up to 68 s). LLM performance on pediatric dentistry questions was highly variable. Top models, notably Gemini 2.5 Pro (92.44%) and GPT-5 (90.76%), approached expert-level accuracy, while others (e.g., Qwen-3) performed poorly. A critical speed-accuracy trade-off was evident: reasoning modes improved scores but were impractically slow, whereas faster models had low accuracy. This variability necessitates careful validation before LLMs are used in high-stakes dental education or assessment.

Authors

Keywords

No keywords available for this article.