Beyond exam accuracy: Tracking a persistent-failure set reveals visual dental reasoning gaps in multimodal LLMs.

Journal: Journal of dentistry
Published Date:

Abstract

OBJECTIVES: To benchmark late-2025 general-purpose multimodal large language models (LLMs) on the Japanese National Dental Examination (JNDE) and to reassess a previously identified persistent-failure set. METHODS: GPT-5.2T, Claude 4.5, and Gemini 3 were tested in January 2026 using a zero-shot protocol (no prompt engineering) on 350 publicly released JNDE-2025 questions (202 text-only; 148 visually-based) and on 33 questions in a persistent-failure set that had been answered incorrectly by all models in a prior JNDE-2024 evaluation. Responses were scored against the official answer key; paired accuracies were compared using Cochran's Q test followed by post hoc pairwise comparisons with Bonferroni-adjusted p values. RESULTS: Overall accuracy was high across all three models, but each model performed worse on visually-based than on text-only questions (71.0-79.7 % vs 91.6-95.5 %). Gemini 3 achieved the highest overall accuracy (88.9 %, 311/350), followed by GPT-5.2T (84.3 %, 295/350) and Claude 4.5 (84.0 %, 294/350). Twenty-one questions (6.0 %) were missed by all models and were predominantly visually-based (16/21), clustering in pediatric dentistry and orthodontics. In the persistent-failure set, 5/33 questions were answered correctly by all models, whereas 9/33 remained incorrect across all models. CONCLUSIONS: Licensing-examination benchmarking yields near-ceiling performance on text-only questions, but substantial gaps remain in visual dental reasoning. Whole-examination benchmarking should therefore be complemented by modality-stratified reporting and dentistry-specific challenge sets to track meaningful progress beyond aggregate examination accuracy. CLINICAL SIGNIFICANCE: High aggregate examination accuracy should not be interpreted as sufficient evidence to judge the reliability of image-conditioned dental support; more clinically grounded evaluations are required.

Authors

Keywords

No keywords available for this article.