Evaluating the Feasibility of Artificial Intelligence in Generating Visual Abstracts: A Pilot Study.
Journal:
The Journal of surgical research
Published Date:
Aug 5, 2026
Abstract
INTRODUCTION: Visual abstracts (VAs) have emerged as an efficient tool to disseminate scientific research, offering faster processing and improved information retention compared to traditional text-only abstracts. Concurrently, artificial intelligence (AI) has seen rapid advancements in language models and image generation. This study aims to evaluate the feasibility of using GPT-4o, and its more advanced model GPT-5.3, available on GPT-Pro paid version to generate VAs for scientific literature. METHODS: A pilot study was conducted using a dataset comprised of 26 pairs of visual and textual abstracts. The dataset was split into a training set (n = 20, 77%) and a testing set (n = 6, 23%) to approximate the commonly used 80-20 rule in machine learning while accommodating our sample size. GPT4o and GPT 5.3 were used as the AI models. The system was fine-tuned on the training set, with each input including the textual abstract and corresponding VA. In the testing phase, GPT4o and GPT 5.3 were provided the six-test set textual abstracts and a template generated from expert guidelines, prompting it to generate VAs. GPT 5.3 output was further refined with two iterative prompts after its first generated image. Our research team then compared the AI-generated VAs to the journal-created VAs, focusing on adherence to the template, accuracy of information representation, visual clarity, and overall design consistency. RESULTS: VAs produced by GPT4o demonstrated significant limitations. They consistently deviated from the provided template, exhibiting excessive imagery and text that often obscured the main message. All six generated VAs contained on average 17 instances (range 12-22) of misspelled and distorted text, impeding readability and comprehension. Notably, there was a lack of consistent structure across the AI-generated VAs, with each output varying considerably in format and style. Information accuracy was also compromised, with the AI output misrepresenting or omitting key data points from the original abstracts for the entire test cohort. The visual clarity was suboptimal, with cluttered layouts and poor color schemes that failed to effectively highlight the most important information. In contrast, VAs generated by GPT 5.3, before and after the iterative prompts, improved widely upon all four domains. While it still made errors regarding journal logo and author name, the rest of the design and consistency was comparable to the human-generated VAs. CONCLUSIONS: Our pilot study demonstrates that while the GPT4o model lacks the capability to produce VAs suitable for scientific knowledge dissemination, a stark improvement was noticed with the most recent GPT-5.3 model, with its ability to adhere to the standardized template and produce images with clarity, accuracy, and visual consistency, required for effective research communication, rendering it useful for generation of a first draft of VAs.
Authors
Keywords
No keywords available for this article.