A Comparative Evaluation of ChatGPT and Google Search Output for Caregiver Education in Tetralogy of Fallot.

Journal: Cardiology and therapy
Published Date:

Abstract

INTRODUCTION: Caregivers of children with tetralogy of Fallot (TOF) frequently seek medical information online to better understand their child's diagnosis, treatment, and long-term care. Recently, generative artificial intelligence (AI) tools, including Google AI Overview and ChatGPT, have emerged as widely accessible sources of health information. However, the quality, readability, and accuracy of these platforms for caregiver education in congenital heart disease remain poorly characterized. METHODS: Forty frequently asked caregiver questions regarding TOF were identified using Google's "People Also Ask" feature. Each question was submitted to ChatGPT-5.5 and entered as a Google search query, generating 80 total responses. Google search output comprised 37 AI Overviews and three organic search excerpts. Responses were evaluated for word count, Flesch Reading Ease (FRE), and Flesch-Kincaid grade level (FKGL). Two independent reviewers assessed response accuracy using a 5-point Likert scale, while one reviewer evaluated educational quality using a modified Ensuring Quality Information for Patients (EQIP) instrument. RESULTS: Compared with Google search output (37 AI Overviews and three organic search excerpts), ChatGPT-5.5 generated significantly longer responses (327.9 ± 117.0 vs. 192.5 ± 85.9 words, p < 0.001), with a higher reading level (FKGL: 23.7 ± 8.5 vs. 15.4 ± 3.6, p < 0.001) and lower readability (FRE: 17.3 ± 14.9 vs. 27.4 ± 11.7, p < 0.001). ChatGPT-5.5 demonstrated higher reviewer-rated accuracy (4.95 ± 0.19 vs. 4.40 ± 0.44, p < 0.001) and higher modified EQIP scores (86.44% ± 7.01% vs. 59.62% ± 13.99%, p < 0.001). CONCLUSIONS: Under the study's prompting conditions, ChatGPT-5.5 responses received higher reviewer-rated accuracy and modified EQIP scores than Google search output but were longer and had higher calculated reading-grade levels. These findings do not establish improved caregiver comprehension, which was not directly assessed.

Authors

Keywords

No keywords available for this article.