Performance of Large Language Models in Answering Juvenile Systemic Lupus Erythematosus Questions: A Blinded Expert-based Comparative Study.

Journal: Journal of clinical rheumatology : practical reports on rheumatic & musculoskeletal diseases
Published Date:

Abstract

BACKGROUND: Large language models are increasingly used to obtain medical information, but their performance in juvenile systemic lupus erythematosus remains insufficiently characterized. OBJECTIVE: To compare expert-rated response quality and clinical relevance across 4 artificial intelligence (AI) models using standardized juvenile systemic lupus erythematosus-related questions. DESIGN: Cross-sectional expert-based comparative study. SETTING: Freely accessible consumer-facing web interfaces evaluated on December 30, 2025. PARTICIPANTS: Ten pediatric rheumatology experts. INTERVENTION: Not applicable. MAIN OUTCOME MEASURES: Twenty standardized questions generated 80 AI responses, rated on a 5-point Likert Scale. Model performance was compared using the Friedman test with Kendall W and Bonferroni-corrected post hoc analyses; interrater agreement was assessed using intraclass correlation coefficients (ICCs). RESULTS: DeepSeek achieved the highest median rating [5 (1 to 5)], followed by Gemini [4 (2 to 5)], ChatGPT [4 (3 to 5)], and Copilot [4 (3 to 5)]. Overall expert-rated performance differed significantly among models [Friedman χ2(3) = 63.553, p < 0.001; Kendall W = 0.106]. DeepSeek outperformed ChatGPT, Gemini, and Copilot (all p < 0.001), while Gemini performed better than ChatGPT (p = 0.007) and Copilot (p < 0.001). Interrater agreement was highest for DeepSeek (ICC =0 .734) and lowest for Copilot (ICC = 0.043). Trust scores remained unchanged before and after evaluation [4 (3 to 5) vs. 4 (3 to 5), p=0.414]. LIMITATION: Single-time-point, single-response testing did not assess within-model reproducibility, and the question set was not formally validated. CONCLUSION: AI models differed in expert-rated performance and interrater agreement; outputs should be interpreted cautiously and with specialist oversight.

Authors

Keywords

No keywords available for this article.