Are large language models such as ChatGPT, capable of supporting patients and general practitioners after spine surgery?

Journal: European spine journal : official publication of the European Spine Society, the European Spinal Deformity Society, and the European Section of the Cervical Spine Research Society
Published Date:

Abstract

PURPOSE: To evaluate whether large language models (LLMs) can provide accurate, complete, and audience-adapted answers to common spine-surgery-related questions for patients and family practitioners. METHODS: Ten frequently asked spine-surgery questions were collected at a level 1 trauma center and simplified linguistically. Five LLMs (ChatGPT, Claude 3.5 Sonnet, Gemini Advanced 1.5 Pro, Copilot Pro, and DeepSeek V3) were queried using zero-shot prompting with persona-specific instructions for family practitioners and middle-aged patients. Responses were assessed by spine surgeons and non-medical raters for correctness, completeness, adaptability, and empathy using five-point Likert scales. Readability was quantified using the Flesch Reading Ease Score (FRES). RESULTS: All LLMs generated largely correct and usable responses. ChatGPT and Claude showed the highest correctness and completeness, particularly for practitioner-directed answers. Gemini and Copilot achieved superior readability and empathy for patient-facing responses. DeepSeek demonstrated balanced performance across all domains. Readability differed substantially between practitioner- and patient-oriented outputs. CONCLUSION: LLMs can support communication and education following spine surgery when used with structured prompting. Clinical oversight remains essential to mitigate risks related to inaccuracies and hallucinations. LEVEL OF EVIDENCE: III.

Authors

Keywords

No keywords available for this article.