Large language models for lumbar spondylolisthesis detection: a multi-center pilot comparative radiographic accuracy study.

Journal: European journal of orthopaedic surgery & traumatology : orthopedie traumatologie
Published Date:

Abstract

INTRODUCTION: Lumbar spondylolisthesis remains a radiographic diagnosis. With the improvement of large language models (LLMs), such as ChatGPT, in interpreting multimodal images, many patients have turned to LLMs for diagnostic insight. Yet the diagnostic reliability of emerging LLMs for spinal pathologies remains understudied. This study sought to evaluate the performance of both ChatGPT-4o and ChatGPT-5 in detecting lumbar spondylolisthesis on standing lateral lumbar radiographs from a public imaging dataset evaluated by expert spine surgeon consensus. METHODS: From the VinDr-SpineXR dataset library, we extracted 200 standing lateral lumbar radiographs, including 100 labeled as spondylolisthesis-positive and 100 labeled as spondylolisthesis-negative. Five fellowship-trained spine surgeons independently reviewed all 200 radiographs. Expert consensus was defined as agreement by at least three surgeons. These same 200 radiographs were independently analyzed by GPT-4o and GPT-5 using a standardized binary prompt to assess presence or absence of spondylolisthesis. Diagnostic performance was assessed relative to expert surgeon consensus. RESULTS: After spine surgeon consensus was established, spondylolisthesis was confirmed in 81% of VinDr-SpineXR spondylolisthesis-labeled positive radiographs, while 2% of VinDr-SpineXR spondylolisthesis-labeled negative radiographs were reclassified as positive. With expert surgeon consensus set as ground truth, ChatGPT-5 outperformed ChatGPT-4o evidenced by higher sensitivity (67.5% vs. 49.4%) and overall accuracy (61.5% vs. 55.5%), with minor difference in specificity (57.3% vs. 59.8%). Inter-rater agreement was higher with ChatGPT-5 (κ = 0.238) than GPT-4o (κ = 0.091). CONCLUSIONS: ChatGPT-5 outperformed ChatGPT-4o in detecting lumbar spondylolisthesis. Yet, both LLMs remained limited compared to fellowship trained spine surgeons.

Authors

Keywords

No keywords available for this article.