Moderate-to-substantial agreement of ChatGPT-5 for Kellgren-Lawrence grading on synthetic knee radiographs: a controlled cross-sectional observer agreement study.

Journal: Rheumatology international
Published Date:

Abstract

AI models are increasingly explored for radiographic assessment of knee osteoarthritis, but their reliability as an experimental model for KL grading remains uncertain. This study evaluated ChatGPT-5 agreement with expert consensus for KL grading and compared its performance with a secondary AI model. This controlled cross-sectional observer agreement analysis used 630 hand-compiled synthetic posteroanterior fixed flexion knee radiographs selected for technical adequacy, anatomical suitability, and interpretability. Three experienced clinicians, one orthopedic surgeon and two rheumatologists, independently graded all images, and a three-rater majority-consensus expert reference standard was established. ChatGPT-5 assessed the same images using a standardized prompt and image-level protocol. Gemini 2.5 Flash assessed a 112-image subset. Agreement was analyzed using weighted Cohen's κ, Gwet's AC2, and mean absolute error (MAE). Binary classification behavior was evaluated at the KL ≥ 2 threshold. Inter-rater agreement before consensus formation was high: Fleiss' κ was 0.884, Gwet's AC2 was 0.921, and pairwise weighted κ ranged from 0.902 to 0.929. Complete three-rater agreement was observed in 520 radiographs (82.5%), and only 14 cases (2.2%) required adjudication. Compared with the expert reference standard, ChatGPT-5 showed moderate-to-substantial agreement, with weighted κ = 0.680 (95% CI 0.616-0.737), MAE = 0.589, and 60.9% exact five-grade agreement. Exact agreement was highest for KL 0 (81.5%) and KL 4 (66.9%) and lower for intermediate grades. For binary KL ≥ 2 classification, ChatGPT-5 correctly classified 513/630 radiographs, with 81.4% accuracy (95% CI 78.2-84.3), 82.1% sensitivity, 80.3% specificity, 87.6% PPV, and 72.6% NPV. A small but significant shift toward higher grading was observed (p = 0.013), with up-grading more frequent in KL 0-1 and down-grading more frequent in KL 3-4. In the 112-image exploratory subset, Gemini 2.5 Flash showed limited agreement with the expert reference standard (weighted κ = 0.23, MAE = 0.66, exact agreement = 50.0%) and poor agreement with ChatGPT-5 (weighted κ = 0.06). For KL ≥ 2 classification, Gemini achieved 87.5% accuracy, 100% sensitivity, 0% specificity, and 87.5% PPV; NPV was not estimable because no KL < 2 classifications were generated. In this controlled synthetic radiograph study, ChatGPT-5 showed moderate-to-substantial agreement with expert consensus under standardized experimental conditions. However, the curated synthetic dataset, limited technical reproducibility of web-based inference, systematic grading tendencies, exploratory nature of the secondary-model analysis, and limited ecological validity preclude any inference of clinical readiness or real-world diagnostic performance without validation on real-world, multicenter, prevalence-based knee radiograph datasets.

Authors

Keywords

No keywords available for this article.