Diagnostic accuracy of electronic medical record retrieval methods and a large language model for identifying cardiovascular events: a multisite retrospective validation study in a medical system in the United States.

Journal: BMJ open
Published Date:

Abstract

OBJECTIVE: To compare the diagnostic accuracy of four available automated electronic medical record (EMR) retrieval methods, including a large language model (LLM)-assisted workflow, against manual chart adjudication for identifying cardiovascular events. DESIGN: Retrospective diagnostic accuracy study. SETTING: Three sites within a single US tertiary health system. PARTICIPANTS: Two adult cohorts with previously adjudicated cardiovascular outcomes were included. Cohort 1 included 2258 patients treated with immune checkpoint inhibitors, and Cohort 2 included 1426 patients who underwent transcatheter aortic valve replacement. PRIMARY AND SECONDARY OUTCOME MEASURES: The reference standard was clinician manual chart adjudication. Outcomes included ischaemic stroke or transient ischaemic attack, myocardial infarction (MI), heart failure (HF) exacerbation or hospitalisation and a composite major adverse cardiovascular events (MACE) outcome. Automated retrieval methods included International Classification of Diseases (ICD) codes, primary diagnosis, problem list and a zero-shot LLM workflow. Area under the (receiver operating characteristic) curve (AUC), sensitivity, specificity and net reclassification improvement were assessed. RESULTS: In Cohort 1, the LLM achieved the highest AUC for stroke (0.920; 95% CI 0.881 to 0.958), MI (0.938; 95% CI 0.905 to 0.971) and composite MACE (0.880; 95% CI 0.854 to 0.907), whereas ICD-based retrieval had a higher AUC for HF (0.882; 95% CI 0.845 to 0.918 vs 0.873; 95% CI 0.831 to 0.914). In Cohort 2, the LLM achieved the highest AUC for all evaluated outcomes: stroke (0.915; 95% CI 0.862 to 0.968), MI (0.928; 95% CI 0.839 to 1.000), HF (0.844; 95% CI 0.803 to 0.884) and composite MACE (0.862; 95% CI 0.829 to 0.895). In Cohort 1, differences in AUC between the LLM and ICD methods were not statistically significant across outcomes, whereas in Cohort 2 the LLM showed significantly higher AUC for stroke and composite MACE. CONCLUSION: In this multisite retrospective validation study, the LLM-assisted workflow showed strong but context-dependent performance for identifying cardiovascular events from the EMR. Performance varied by outcome and cohort, and ICD-based retrieval remained competitive for some use cases. These findings support a complementary role for LLM-assisted extraction in retrospective cardiovascular outcomes research.

Authors

Keywords

No keywords available for this article.