Performance of multimodal large language models versus clinicians for radiographic knee osteoarthritis grading: A multiobserver study.

Journal: Skeletal radiology
Published Date:

Abstract

OBJECTIVE: To compare the performance of clinicians and two generations of multimodal large language models (LLMs) in Kellgren-Lawrence (KL) grading of knee osteoarthritis (KOA), including feature-level assessment and intraobserver repeatability. MATERIALS AND METHODS: In this retrospective single-center study, 348 knee radiographs were graded by a senior musculoskeletal radiologist (reference standard), a radiologist, an orthopedic surgeon, a radiology resident, and LLMs (ChatGPT-4o and ChatGPT-5.0). Binary KOA detection (KL 0-1 vs ≥ 2), feature-level interpretation (joint space narrowing, osteophytes, subchondral sclerosis), and intraobserver repeatability were evaluated. Agreement metrics included weighted κ, accuracy, and standard diagnostic measures. RESULTS: Agreement with the reference standard was highest for the radiologist (κ = 0.87), followed by the orthopedic surgeon and radiology resident. Both LLMs demonstrated moderate agreement, with ChatGPT-5.0 outperforming ChatGPT-4o. For binary KOA detection, ChatGPT-5.0 showed very high sensitivity (0.96) but reduced specificity. Per-grade classification was most accurate for KL 0 and KL 4, but remained limited for KL 1-2. Feature-level concordance was modest across all radiographic findings. Intraobserver repeatability was highest for the reference reader (κ = 0.881), followed by the orthopedic surgeon (κ = 0.634) and radiologist (κ = 0.626), while lower agreement was observed for the resident (κ = 0.478) and LLMs, with ChatGPT-5.0 showing higher consistency than ChatGPT-4o (κ = 0.591 vs. 0.485). CONCLUSION: Although ChatGPT-5.0 outperformed ChatGPT-4o, both models remained inferior to clinicians in detailed KL grading, feature-level interpretation, and reproducibility. Current multimodal LLMs show high sensitivity but limited specificity and are not suitable for standalone radiographic KOA assessment.

Authors

Keywords

No keywords available for this article.