Assessing the scientific quality of artificial intelligence generated graphical abstracts in respiratory medicine: an expert-based multicentre pilot study.
Journal:
Journal of visual communication in medicine
Published Date:
Sep 1, 2026
Abstract
Graphical abstracts are increasingly used to enhance scientific communication, yet their quality remains variable. Generative artificial intelligence (AI) tools can produce graphical abstracts, but their scientific reliability has not been systematically evaluated. To assess the scientific quality of AI-generated graphical abstracts in respiratory medicine. This multicentre study included a pilot phase (5 abstracts; 15 graphical abstracts) and a main phase (20 abstracts; 60 graphical abstracts). For each text abstract, three graphical abstracts were generated via ChatGPT (version 5.2), Claude Sonnet (version 4.5), and Gemini (version 3), using a standardised prompt. Six experts evaluated each graphical abstract using a 7-item scoring grid (total score/28). Inter-rater reliability and comparisons between models were assessed.The median total score was 21 [16-25]. Single-measure intraclass correlation coefficients (ICCs) indicated moderate agreement for the total score (ICC = 0.606), while average-measure ICC showed excellent reliability (ICC = 0.902). Significant differences were observed between models (p < 0.001), with a consistent ranking of Claude > ChatGPT > Gemini. Differences were observed across all evaluation criteria, with large effect sizes (Kendall's W up to 0.93). AI-generated graphical abstracts demonstrate moderate-to-high quality but remain heterogeneous across models. While promising as assistive tools, their use requires expert validation to ensure scientific accuracy and interpretative safety.
Authors
Keywords
No keywords available for this article.