Accuracy and Readability of Generative Artificial Intelligence for Vascular Surgery Patients: A Specialist Based Evaluation Highlighting the Current Landscape of Safety Risks and Accessibility Gaps.

Journal: European journal of vascular and endovascular surgery : the official journal of the European Society for Vascular Surgery
Published Date:

Abstract

OBJECTIVE: Patients frequently seek health information and medical advice from chatbots instead of consulting their physicians or referring to credible patient education resources provided by medical societies. A study to evaluate the quality, readability, and clinical appropriateness of ChatGPT generated answers to common vascular surgery questions was designed. METHODS: Sixteen questions, written in a layperson style, were developed covering four major vascular conditions: abdominal aortic aneurysm; carotid stenosis; peripheral arterial occlusive disease; and varicose veins. The questions were categorised across four domains (signs/symptoms, natural history, medical advice, and best treatment) and further grouped for analysis into symptom related and treatment related queries. They were addressed to ChatGPT (gpt-3.5-turbo-0125) in dedicated sessions. Tone, complementarity, urgency, and uncertainty were assessed using adapted QUEST, DISCERN, and urgency scales. The Flesch Reading Ease (FRE) and Flesch-Kincaid Grade Level (FKGL) were used to measure readability. The accuracy, comprehensiveness, and clarity of the information was rated on a Likert scale by three board certified vascular surgeons. RESULTS: No significant differences were found in tone or complementarity across disease categories or question types. Urgency did not differ significantly between symptom related and treatment related queries overall; however, urgency varied significantly across subtypes of symptom related questions (p = .03), consistent with inconsistent escalation recommendations. The mean FRE was 32.3 ± 12.1 and FKGL 13.5 ± 2, corresponding to a university level reading requirement. Symptom related responses were more readable than treatment related ones (FRE 38.1 ± 10.3 vs. 26.4 ± 11.4; p = .025). Among the 16 outputs, only seven (44%) were judged clinically appropriate by all reviewers; clarity was rated adequate in 81% of responses, whereas only 50% reached acceptable accuracy and 69% acceptable comprehensiveness. Treatment related answers were particularly weak, with only 25% deemed appropriate. In symptom related questions, misalignment of urgency recommendations emerged as a potential patient safety concern. CONCLUSION: In this specialist evaluation, ChatGPT outputs related to vascular surgery were clear, but mostly clinically inappropriate and inaccurate, and required 13th grade reading skills. The combination of low accuracy and poor accessibility presents a serious patient safety concern for unsupervised patient education. As they stand, generalist large language models cannot provide patient facing information about vascular surgery. A rigorously validated, domain specific, and knowledge locked artificial intelligence system based on curated vascular guidelines may be more appropriate to ensure safety and comprehension.

Authors

Keywords

No keywords available for this article.