Performance of multimodal large language models in interpreting lateral cephalometric superimpositions: A comparative observer-performance study.

Journal: International orthodontics
Published Date:

Abstract

INTRODUCTION: Multimodal large language models (LLMs) can generate free-text interpretations of clinical images, but their performance on orthodontic cephalometric superimpositions is unknown. This study compared zero-shot interpretations from three LLMs with those of a second-year orthodontic resident. METHODS: Ninety lateral cephalometric superimposition images from a private orthodontic practice were analyzed, including 30 nongrowing, 30 growing, and 30 orthognathic cases. Each image included overall maxillary regional, and mandibular regional superimpositions. ChatGPT 5.5, Gemini 3.1 Pro, Claude Opus 4.8, and the resident interpreted the same images using the same prompt, with no case context provided. Two senior orthodontists scored each interpretation against adjudicated reference interpretations using a 16-item rubric, yielding total scores from 0 to 32 and four domain scores. Friedman tests compared methods; Wilcoxon signed-rank tests with Holm adjustment compared each LLM with the resident. RESULTS: Total scores differed significantly among methods (Friedman P<0.001; Kendall W=0.592). Median total scores were 30.5 (interquartile range [IQR]: 28-32) for the resident, 17 (IQR: 14-22) for ChatGPT, 12 (IQR: 9-16) for Gemini, and 9.5 (IQR: 4.25-14) for Claude. All LLMs scored significantly lower than the resident overall, by domain, and within each case type (adjusted P<0.001). CONCLUSIONS: In zero-shot cephalometric superimposition interpretation, all tested multimodal LLMs performed substantially below a second-year orthodontic resident. These models should not be used as stand-alone interpreters without expert review.

Authors

Keywords

No keywords available for this article.