An expert-led benchmark using common patient questions: evaluating large language models for adolescent idiopathic scoliosis education.

Journal: European spine journal : official publication of the European Spine Society, the European Spinal Deformity Society, and the European Section of the Cervical Spine Research Society
Published Date:

Abstract

AIM/BACKGROUND: Large language models (LLMs) are increasingly used by patients to obtain medical information. Adolescent idiopathic scoliosis (AIS), a chronic condition requiring long-term monitoring and treatment decisions, generates substantial demand for reliable and understandable patient education. Although LLMs may function as accessible explanatory tools, their suitability for patient-oriented use remains uncertain. This study aimed to perform an expert-led, patient-centered evaluation of two widely accessible LLMs, Claude Sonnet 4.5 and GPT 5.2, focusing on their ability to deliver accurate, clear, and conceptually adequate responses to common AIS-related patient questions. METHODS: A cross-sectional comparative design was used with 100 high-frequency patient questions covering ten clinical domains. Responses generated by both models using standardized zero-shot prompts were independently assessed by expert clinicians: factual accuracy by three raters (two orthopedic spine surgeons and one senior pediatric physiotherapist), and clarity and conceptual coverage by two raters (one surgeon and the physiotherapist). A structured evaluation framework examined three dichotomous dimensions relevant to patient education: factual accuracy, clarity and understandability, and conceptual coverage. Model performances were compared using McNemar's test, and inter-model agreement was assessed with Krippendorff's alpha. RESULTS: Both models demonstrated equally high factual accuracy (91%). However, clarity was limited, with only one-third of responses rated as sufficiently understandable. A significant difference was observed in conceptual coverage, with Claude Sonnet 4.5 outperforming GPT 5.2 (46% vs. 29%, p = 0.012), particularly in domains requiring integrative explanations. CONCLUSION: Despite strong factual accuracy, current LLMs show deficiencies in clarity and conceptual depth, limiting their reliability as standalone patient education tools for AIS. These findings highlight the necessity of clinician mediation and the importance of patient-centered evaluation criteria before clinical adoption. CLINICAL TRIAL REGISTRATION: As this study is not a clinical trial, clinical trial registration is not applicable.

Authors

Keywords

No keywords available for this article.