Large language models and chatbots in orthodontic patient education: A systematic review and meta-analysis.

Journal: International orthodontics
Published Date:

Abstract

INTRODUCTION: Patients now arrive at the orthodontic consultation having already consulted a chatbot, yet the accuracy, quality and readability of these tools lack quantitative synthesis. METHODS: Six databases and one trial register were searched on 24 April 2026 for LLM-based chatbots evaluated on patient-facing orthodontic content. RoB 2 was used for the RCT and an adapted JBI checklist for cross-sectional studies. Comparisons were pooled by instrument across accuracy, quality and readability with random-effects meta-analysis and the Hartung-Knapp-Sidik-Jonkman (HKSJ) small-sample correction; a ChatGPT-4-class primary analysis addressed mixed-generation pooling, and GRADE was applied to every pool. The review was single-reviewer with a pre-specified audit. RESULTS: Thirty-nine studies were included; risk of bias was low in 10, moderate in 10 and high in 19. ChatGPT-4-class and Gemini did not differ on Likert accuracy (MD -0.12 [-0.42 to +0.17]; I2=70.6%; p_HKSJ=0.533), and no accuracy or quality pool was significant under both the DerSimonian-Laird and HKSJ analyses. The largest pooled difference was in readability: across two studies, ChatGPT text was about 2.8 Flesch-Kincaid grades easier than Claude (MD -2.85 [-3.31 to -2.38]; p_DL<0.001), but only borderline under the conservative HKSJ (p_HKSJ=0.053) and therefore provisional. Five of 11 pools showed I2≥75%, and GRADE certainty was very low for every pool. CONCLUSION: No chatbot consistently dominates across accuracy, quality and readability; the largest difference, a provisional readability advantage of ChatGPT over Claude, rests on two studies. Chatbots cannot be relied upon as a safe complement to professional patient education and should never replace the information provided by the orthodontist, the more so because the same chatbot may not behave the same way within months.

Authors

Keywords

No keywords available for this article.