Large language models in the diagnosis and treatment of immunosuppression-related infections in rheumatic diseases: a comparative, vignette-based study.
Journal:
Rheumatology international
Published Date:
Oct 9, 2026
(1)
Abstract
This study compared the diagnostic accuracy, initial test selection, and initial treatment choices of three current-generation Large Language Models (LLMs) with those of two infectious disease specialists in managing infectious complications in patients with immunosuppressed rheumatic disease. This cross-sectional comparative study utilized 100 standardized Turkish-language clinical vignettes, developed by an expert committee, encompassing ten infection categories: viral, bone/joint, respiratory, central nervous system, gastrointestinal, genitourinary, skin and soft tissue, rare diseases, vaccination and prophylaxis, and adverse drug reactions. A prespecified reference standard, based on current clinical practice guidelines (ACR, EULAR, KLİMİK, IDSA, ATS, WHO, SANJO/EBJIS), was established prior to data collection. Three LLMs (Gemini 3.1 Pro, ChatGPT 5.5, Claude 4.7 Opus) and two blinded infectious disease specialists, operating under closed-book conditions, answered the same vignettes across three domains: diagnosis, initial diagnostic work-up, and initial treatment. Responses were evaluated dichotomously as correct or incorrect according to the reference standard. Diagnostic accuracy was 91% and 93% for the two human experts, compared to 99%, 97%, and 100% for Gemini 3.1 Pro, ChatGPT 5.5, and Claude 4.7 Opus, respectively. This pattern was consistent across all three domains (Cochran Q, all p ≤ 0.003). Human experts achieved complete case management in 83% and 90% of vignettes, while the LLMs achieved 97-100% (p < 0.001). Of 40 pairwise comparisons, 10 showed Bonferroni-corrected significance, all in human-versus-model comparisons. In this vignette-based comparison, all three LLMs answered a higher proportion of items correctly than the two participating infectious disease specialists in all three domains. As the cases were standardized written vignettes with a single prespecified correct answer, and only two human comparators were available, this reflects performance on a structured written task rather than superiority in clinical care. These findings suggest that LLMs could serve as a complementary decision-support tool, especially for identifying rare or opportunistic diagnoses, but should not replace specialists without further prospective, real-world validation.
Authors
Keywords
No keywords available for this article.