Evaluation of Generative Artificial Intelligence Models in Producing Anatomically Accurate Illustrations: A Comparative Study of Text-to-Image Generators.
Journal:
Clinical anatomy (New York, N.Y.)
Published Date:
Jul 20, 2026
Abstract
Generative artificial intelligence (GAI) is increasingly being applied in biomedical sciences and medical education, including anatomy, where text-to-image generators may facilitate rapid creation of visual materials. However, the anatomical accuracy of such generated illustrations remains uncertain. The present study evaluated the performance of selected GAI models in generating anatomically correct representations of human anatomical structures. Six text-to-image generators (DALL-E, Gemini, Freepik, Midjourney, DeepAI, and Canva) were assessed using a standardized prompt ("Generate the most anatomically accurate image of the human [name of anatomical structure]"). For each anatomical region, two images were generated. The evaluated structures included the liver, femur, scapula, aortic arch, arm muscles, sacrum, humerus, thigh muscles, celiac trunk, kidney, brainstem, and lumbar plexus. All images were analyzed using reference tables based on Terminologia Anatomica, with individual structures assessed for presence and physiological correctness by four independent reviewers. Statistical analysis included calculation of structure-level proportions with 95% confidence intervals, pairwise inter-rater agreement using Cohen's κ, and exploratory logistic regression with robust standard errors clustered by generated image to evaluate the effects of model, anatomical region, and evaluator. A total of 144 images were analyzed. DALL-E demonstrated the highest overall proportion of structure-level entries scored as present (56.7%, 95% CI: 53.8-59.5) as well as the highest proportion of physiologically correct entries among those present (66.0%, 95% CI: 62.2-69.5). When both criteria were combined, the proportion of fully correct structure-level entries remained limited across all models, with the highest value observed for DALL-E (37.4%). Considerable variability in performance was noted across anatomical regions, with particularly low accuracy observed for complex structures such as the lumbar plexus and celiac trunk. Inter-rater agreement ranged from moderate to almost perfect. Overall, current GAI models demonstrate substantial limitations in producing anatomically accurate illustrations. Although certain models outperform others, the reliability of generated images remains insufficient, and expert verification is necessary before their application in medical education or scientific contexts.
Authors
Keywords
No keywords available for this article.