Evaluating the reliability and information quality of ChatGPT responses on vaccine hesitancy: an expert panel study.
Journal:
BMC public health
Published Date:
Jul 21, 2026
Abstract
BACKGROUND: Vaccine hesitancy is a major public health problem that threatens global immunisation programmes. The use of artificial intelligence-based chatbots to access health information is increasing. This study aimed to evaluate the responses of ChatGPT-4o to frequently asked questions about vaccine hesitancy in terms of health information quality, based on expert opinion. METHODS: This cross-sectional expert evaluation study included 20 questions on vaccine hesitancy. These questions were submitted to ChatGPT-4o, and the responses were evaluated by nine experts from the fields of public health, paediatrics, and infectious diseases. The evaluation was based on six criteria synthesised from HONcode, DISCERN, JAMA Benchmarks, CRAAP, GQS, and QUEST: scientific accuracy, comprehensiveness, understandability, correction of misinformation, source attribution, and actionability. Each criterion was scored on a scale from 1 to 10. Inter-rater reliability was assessed using the intraclass correlation coefficient (ICC) and, as a prevalence-robust sensitivity analysis, Gwet's AC2 with quadratic weights. RESULTS: The highest mean scores were found for understandability (9.4 ± 1.1) and scientific accuracy (8.9 ± 1.3), while the lowest mean score was found for source attribution (4.5 ± 3.4). At the question level, the highest scores were obtained for Question 4, on aluminium in vaccines, and Question 14, on claims that pharmaceutical companies endanger children's health by producing and promoting vaccines. The lowest scores were obtained for Question 5, on live vaccines during pregnancy, and Question 6, on natural immunity. ICC analysis showed moderate agreement only for source attribution (ICC = 0.582; p < 0.001); the low ICC values for the other criteria were attributable to a ceiling effect, as confirmed by Gwet's AC2, which indicated substantially higher agreement for the high-scoring criteria. CONCLUSIONS: ChatGPT-4o can generate scientifically accurate and understandable responses to questions about vaccine hesitancy. However, it shows a systematic shortcoming in source attribution. Although the model has potential as a supportive tool in public health communication, its outputs should be reviewed by experts and supported with verifiable sources.
Authors
Keywords
No keywords available for this article.