Comparison of AI-assisted and human-produced podcasts derived from Cochrane Plain Language Summaries: protocol for a randomised non-inferiority trial (Health Information Effectiveness Trials-2).
Journal:
Journal of clinical epidemiology
Published Date:
Apr 20, 2026
Abstract
BACKGROUND AND OBJECTIVES: Podcasts can make health evidence easier to follow, but it is unclear whether artificial intelligence (AI)-assisted production can match human production when both use the same audio format. This trial will evaluate whether a vetted AI workflow can match human communicators on comprehension, quality, safety, accessibility, and trust when both deliver podcasts derived from the same evidence base. METHODS: We will run a randomized, two-arm, noninferiority trial comparing AI-assisted podcasts with human-produced podcasts. Adults (≥18 years; English-proficient) will be recruited from the general public via Prolific, an online research participant recruitment platform, and randomly allocated 1:1 to listen to three short episodes (6-8 minutes each) based on the same Cochrane Plain Language Summaries. The AI arm uses Wondercraft AI in a human-in-the-loop workflow; the human arm features experienced communicators working to an identical brief. In both arms, content is limited to the plain language summary, with authorship masked for participants and expert raters. The primary outcome is comprehension, measured by a 10-item test per episode, with the primary analysis using the participant-level mean score across the three episodes, aligned with the "Understanding" dimension of the Quality, Understanding, Expression, Safety, and Trust framework. Secondary outcomes include format accessibility (listenability), quality of information, perceived trust, and safety. Noninferiority margins are prespecified; for comprehension, the margin is 1 point on the 10-item scale. If noninferiority is shown, we will also assess superiority. We plan to recruit 458 participants. Differences between arms will be estimated using appropriate repeated-measures models, with two-sided 95% confidence intervals. RESULTS: This is a study protocol; results will be reported on completion of the trial. CONCLUSION: By providing head-to-head evidence in the same audio format, the study will address a practical question faced by journals and health organizations already experimenting with AI tools: can AI generate clear, safe, and trusted audio content at scale, and where does human input remain essential?
Authors
Keywords
No keywords available for this article.