Language-dependent performance variation in large language models for dental trauma management: a comparative evaluation of ChatGPT-5.2, Gemini 3.0, and Claude 4.5 Sonnet.

Journal: BMC oral health
Published Date:

Abstract

BACKGROUND: Large language models (LLMs) are increasingly evaluated for medical question answering and clinical information tasks, yet the impact of query language on their performance in specialized domains such as dental traumatology remains insufficiently studied. The primary objective was to evaluate whether query language (English vs. Turkish) affects LLM performance in a controlled scenario-based assessment of dental trauma management. Secondary objectives were to compare overall performance across three LLMs and to examine whether language effects are uniform across models or model-specific. METHODS: Twenty-seven clinical scenarios covering 13 dental trauma categories were presented to ChatGPT 5.2, Gemini 3.0, and Claude 4.5 Sonnet in both English and Turkish, generating 162 responses. Two blinded endodontists independently evaluated responses using a standardized rubric assessing accuracy (40%), completeness (35%), and safety (25%) against IADT 2020 Guidelines. Inter-rater reliability was assessed using intraclass correlation coefficient (ICC). Language effects were analyzed using Wilcoxon signed-rank tests; model comparisons employed Kruskal-Wallis and Mann-Whitney U tests with Bonferroni correction. RESULTS: Inter-rater reliability ranged from moderate to good across evaluation dimensions (ICC: 0.738-0.836). ChatGPT showed the strongest language effect with 9.14% higher performance in English (p < 0.001, r = 0.874). Gemini showed moderate English advantage (5.69%, p = 0.003, r = 0.572). Claude exhibited language independence with virtually identical performance in both languages (-0.02%, p = 0.220). In English, significant model differences emerged (H = 22.31, p < 0.001); however, model performance converged in Turkish (H = 2.89, p = 0.236). CONCLUSIONS: This exploratory scenario-based study demonstrates model-specific, rather than universal, language effects on LLM performance in dental trauma management. ChatGPT 5.2 achieved the highest performance in English but exhibited the most pronounced Turkish-language degradation, including substantial safety score decline. Gemini 3.0 showed an intermediate pattern with moderate English advantage. Claude 4.5 Sonnet demonstrated language-independent performance across all evaluated dimensions. These findings are based on a standardized scenario-based assessment and should not be extrapolated to real clinical environments or patient care settings.

Authors

Keywords

No keywords available for this article.