Prospective evaluation of a large language model clinical decision support system in the emergency department.
Journal:
Nature medicine
Published Date:
Aug 19, 2026
Abstract
Prospective evidence for artificial intelligence (AI)-based clinical decision support in emergency departments remains limited. Here we conducted a DECIDE-AI stage 1 evaluation of SHAKED, a clinical decision support system built on multiple large language models, in a tertiary emergency department. Over 4 weeks, 1,138 patients were analyzed across two parallel units-one using SHAKED and one following routine rotations. Clinical adoption of SHAKED declined from 68% to 30%, owing to workload-sensitive disengagement (OR = 0.72 per shift hour, 95% CI 0.62 to 0.83). Physicians preferred the use of SHAKED for radiology consultations (OR = 2.98, 95% CI 1.58 to 5.63). No adverse events were detected, and expert review rated 99 of 100 sampled outputs as clinically appropriate. Emergency department length of stay did not differ between wings (4.9 h in both, P = 0.99). Intention-to-treat analysis showed a non-significant trend toward shorter consultation cycle time (-9.4 min, P = 0.077). These findings suggest that sustained clinician engagement, rather than algorithmic accuracy, may be the key barrier to effective clinical AI use in emergency departments. They inform randomized trial design but do not justify clinical deployment of AI clinical decision support at this stage. ClinicalTrials.gov identifier: NCT06902675 .
Authors
Keywords
No keywords available for this article.