Large Language Models Provide Accurate but Potentially Unsafe Answers to Multimodal Critical Care Medicine Board Review Questions.
Journal:
Critical care medicine
Published Date:
Jun 18, 2026
Abstract
OBJECTIVES: To evaluate the performance of Chat Generative Pre-Trained Transformer-4 Omni (ChatGPT-4o) in answering multimodal critical care board review questions, with a focus on accuracy, image interpretation, reasoning quality, and potential for harm. DESIGN: Observational study using a validated item bank of multiple-choice questions accompanied by clinical images, analyzed through a custom ChatGPT-4o profile built using a validated framework. SETTING: Simulated environment mimicking critical care board examination conditions, with artificial intelligence responses reviewed by a panel of experienced critical care clinicians. SUBJECTS: One hundred eighty-three board-style questions from the Society of Critical Care Medicine item bank, representing a range of critical care domains and imaging modalities. ChatGPT-4o was evaluated on its responses, which were assessed by 14 clinical reviewers (physicians, advanced practice providers, and pharmacists). INTERVENTIONS: None. MEASUREMENTS AND MAIN RESULTS: ChatGPT-4o answered 74.9% of questions correctly, higher than pooled clinician responses (71.1%; p = 0.03). It showed strengths in question comprehension (87.4% correct) but lower performance in image interpretation (61.7%), reasoning (68.3%), and supporting information (66.1%). ChatGPT-4o excelled in pulmonary disease (91.7%), surgery and trauma (87.5%), and neurologic disorders (81.8%), and underperformed in critical care ultrasound (51.1%). Notably, 33.3% of its responses were associated with potential for clinical harm, often due to incorrect image interpretation and treatment recommendations. CONCLUSIONS: ChatGPT-4o demonstrates performance slightly above pooled clinician benchmarks on critical care board-style questions but has substantial limitations in multimodal question interpretation. Despite its high comprehension, deficiencies in ChatGPT-4o's reasoning and image analysis may lead to harmful clinical conclusions in high-stakes clinical decision-making or clinical education.
Authors
Keywords
No keywords available for this article.