Development and Nationwide Multicentre Evaluation of Guideline-Grounded Large Language Model Chatbots to Support Patient Self-Management and Education in Rheumatology.

Journal: Journal of medical systems
Published Date:

Abstract

Patients with rheumatic diseases have persistent information needs that are not fully addressed in routine care. We developed and evaluated guideline-grounded, large language model (LLM) chatbots to support patient self-management and education in rheumatology.Ten disease-specific chatbots based on German guidelines were co-developed and deployed through 13 rheumatology centres and six patient organisations. Chatbot users rated responses and completed a questionnaire. User questions, feedback, and response characteristics were analysed using category-based coding and a six-dimensional LLM-as-a-judge assessment, with LLM-based ratings compared with rheumatologist ratings in random subsets.Between September 2025 and January 2026, 6291 questions were recorded. Thirteen question categories were identified, most commonly disease-specific questions (50.2%), medication and monitoring (37.3%) and diagnostics (28.2%). The chatbots were unable to answer in 263 interactions (4.2%). Of 2671 responses rated by users, 2481 (92.9%) received a positive rating. Insufficient detail was the most common reason for negative ratings (125/190, 65.8%). Among 602 questionnaire respondents, 84.6% reported that the chatbot was easy to use, 84.1% that answers were easy to understand, and 80.2% that it was a useful addition to patient education. In the LLM-based evaluation, 95.3% of answers were rated as completely safe and 79.1% as completely correct. Guideline adherence was assessed separately, with 45.0% rated as fully adherent; agreement with physician assessment was weak.Guideline-grounded chatbots received predominantly positive user feedback in real-world use, while LLM-based evaluation suggested that most responses were safe and correct. User questions and feedback may help guide iterative improvements to source content and patient education materials. Further studies are needed to evaluate educational effectiveness and independently validate response quality and clinical safety.

Authors

Keywords

No keywords available for this article.