OHCA-EXTRACT: Evaluating the Accuracy of a Large Language Model Pipeline for Out-of-Hospital Cardiac Arrest Case Identification and Utstein Variable Extraction
Journal:
medRxiv
Published Date:
Jul 1, 2026
Abstract
Background: Manual identification and abstraction of out-of-hospital cardiac arrest (OHCA) cases and Utstein template variables from electronic health records is resource-intensive and limits scalable measurement for observational research and registry participation. We evaluated a large language model (LLM)-assisted pipeline to identify true OHCA encounters within an ICD-coded emergency department (ED) cohort and to also extract a limited set of Utstein variables from unstructured documentation, with iterative pipeline refinement and final performance evaluation conducted on the same validation cohort. Methods: We conducted a study of ICD-flagged OHCA encounters within a large urban academic health system (2015?2024). A two-module pipeline was developed that included an identification module and a variable-extraction module. The pipeline was evaluated against independent manual chart abstraction on a validation sample (n = 176) with physician adjudication. The identification module performed binary OHCA classification along with etiology classification. The variable-extraction module extracted five Utstein variables: witnessed status, EMS defibrillation, first recorded rhythm, arrest location, and bystander response. We calculated sensitivity, specificity, PPV, NPV, and F1 scores (Clopper?Pearson 95% CIs) using the manual chart abstraction as the reference gold standard. Results: Of 176 processed encounters, 152 had complete gold-standard classification and were included in identification analyses (OHCA prevalence 81.6%; n = 124 true OHCA, n = 28 non-OHCA). The identification module achieved an accuracy of 0.91 (95% CI 0.85?0.95), sensitivity 0.94 (0.89?0.98), specificity 0.75 (0.55?0.89), PPV 0.94 (0.89?0.98), NPV 0.75 (0.55?0.89), and F1 0.94, compared with accuracy 0.82 for a prevalence-only baseline. Variable-level accuracy ranged from 0.63 to 0.75 (Kappa range 0.31?0.51) across the five extracted Utstein variables, with bystander response showing the lowest agreement (accuracy 0.63, Kappa = 0.31). Conclusions: Within-sample performance estimates indicate that this LLM pipeline can identify true OHCA encounters within an ICD-flagged ED cohort with accuracy and PPV above an administrative-coding baseline. Variable-level extraction accuracy ranged from fair to moderate across the five evaluated Utstein variables, indicating that further methodological development is required before automated extraction is feasible. With external validation and integration with targeted human verification, this approach may support human-augmented workflows using LLMs, however independent external validation is required before our findings can be generalized.