Evaluating Healthcare Provider and Artificial Intelligence Chatbot Responses to Patient Messages from a Health System Using the CREATE TRUST Framework.

Journal: Journal of health communication
Published Date:

Abstract

Integration of artificial intelligence chatbots into healthcare requires rigorous, patient-centered evaluation. This study implements the CREATE TRUST framework-a novel tool evaluating both clinical substance and communication style-to compare responses from healthcare providers (HCPs) and two AI chatbots (GPT-4, Mixtral) to complex clinical questions. In this cross-sectional study, 189 real-world clinical messages from patients with cancer were retrospectively collected from an electronic health record. Anonymized and randomly ordered responses from the HCP, GPT-4, and Mixtral were blindly evaluated in triplicate by a team of oncologists. Evaluators rated each response on every attribute of CREATE TRUST (Correct, Referenced, Empathic, Authentic, Thorough, Engaging, Tailored, Respectful, Understandable, Safe) and provided an overall preference ranking. While HCP responses were ranked first most often (46% of evaluations), there was no significant difference in the overall CREATE TRUST score between HCPs (mean=28.9), GPT-4 (28.4), and Mixtral (28.1). HCPs performed significantly better on Authentic and Tailored. Chatbots scored higher on Empathic and Referenced. Performance was comparable for all other attributes. HCPs demonstrated greater performance variability, authoring a higher proportion of both high- and low-quality responses. HCPs and AI chatbots exhibit comparable overall quality but possess distinct, complementary strengths. HCPs excel in authentic, tailored communication, while chatbots provide more empathic and referenced responses. These findings suggest potential for a synergistic workflow where AI could enhance human-authored messages, improving targeted aspects of communication and mitigating low-quality responses.

Authors

Keywords

No keywords available for this article.