Accuracy of ChatGPT-4o in Identifying Anatomical Structures on Cadaveric Images: A Practical Anatomy Examination Study.
Journal:
Clinical anatomy (New York, N.Y.)
Published Date:
Mar 23, 2026
Abstract
The rapid expansion of large language models (LLMs), including ChatGPT, has generated interest in their potential role in medical education. Although prior studies have evaluated LLM performance in theoretical assessments and selected imaging tasks, their ability to recognize anatomical structures in a cadaver-based practical examination setting remains unclear. This study assessed the accuracy of ChatGPT-4o in identifying anatomical structures on photographs of cadaveric specimens marked in the same manner as during practical anatomy examinations. A total of 265 anatomical structures were labeled on cadaveric specimens from the Department of Anatomy of the Jagiellonian University Medical College and photographed. Using a standardized prompt, the free version of ChatGPT-4o was asked to identify each marked structure, with up to three attempts permitted and standardized feedback provided after incorrect responses. Identification was considered correct only when a valid anatomical term precisely corresponding to the indicated structure was provided. The overall accuracy was 22.26%. Correct identification occurred on the first attempt in 33 cases, on the second in 15, and on the third in 11. Accuracy was highest for osteological structures (64.71% correct within three attempts) and lowest for isolated thoracic organs (8.82%). The model frequently misidentified anatomical regions and occasionally generated non-existent anatomical terms. At its current stage of development, ChatGPT-4o does not appear to be a reliable tool for cadaver-based anatomical structure recognition or practical anatomy examination support.
Authors
Keywords
No keywords available for this article.