Comparative evaluation of the performance of different AI chatbots in answering pediatric dentistry questions in the Turkish dental specialization exam.

Journal: BMC oral health
Published Date:

Abstract

BACKGROUND: Artificial intelligence-powered chatbots, whose use is becoming increasingly widespread, are frequently used in medicine and dentistry for tasks such as generating academic texts, student education, clinical diagnosis, and patient information. Given their inherent limitations, the accuracy of chatbot responses in the healthcare field is a crucial issue. Evaluating the performance of artificial intelligence in national examinations can provide objective data regarding its use as an educational support tool. This study aims to comparatively evaluate the performance of various artificial intelligence chatbots on questions from the pediatric dentistry section of the DUS (Dental Specialization Exam). METHODS: A total of 127 questions asked in the DUS in the field of pediatric dentistry between 2012 and 2021 were included in the study. The questions were divided into nine subject-area categories, two exam years, and two question types: clinical case-based and theoretical knowledge-based. The ChatGPT-5, ChatGPT-4o, Gemini 2.5 Pro, Gemini 2.5 Flash, Claude Opus 4, and Claude Sonnet 4 models were evaluated in the study. Descriptive data are presented as numbers and percentages. Statistical analyses were performed using Cochran's Q test (p < 0.05); pairwise comparisons were made using the McNemar test and Bonferroni correction. RESULTS: Gemini 2.5 Pro answered 93.7% of the questions correctly; ChatGPT-5 and Claude Opus 4 answered 88.9%, ChatGPT 4 and Gemini 2.5 Flash answered 87.4%, and Claude Sonnet 4 answered 83.4%. Gemini 2.5 Pro's performance in correctly answering questions was significantly higher than that of Claude Sonnet 4 (p = 0.001). No statistically significant differences were found in the chatbots' correct-answer performance when comparing by subject, year, and question type (clinical case and theoretical knowledge) (p > 0.05). CONCLUSIONS: This study demonstrates that chatbots perform well at answering exam questions in pediatric dentistry and may serve aspromising supplementary educational tool.

Authors

Keywords

No keywords available for this article.