Scalable discovery and validation of order-specific electronic health record event trajectories for interpretable adverse-outcome risk estimation.

Journal: Journal of biomedical informatics
Published Date:

Abstract

OBJECTIVE: Clinicians must estimate patients' risk of adverse outcomes, yet many electronic health record tools represent health history as an unordered "bag of codes," discarding temporal order. Machine learning tools that leverage sequence order may offer better predictions but often lack interpretability for clinical audit. We introduce an interpretable, order-aware framework that mines frequent event pairs (A → B) and tests whether their ordering provides prognostic information for estimating risk of adverse outcomes (C) beyond the same codes without order. METHODS: Using the NIH All of Us Controlled Tier, 432,617 eligible participants were split 70/30 into discovery and confirmation. We combined 340,687 A → B pairs with nine outcomes (C) to create 3,066,183 candidate trajectories and tested 85,535 trajectories in this pilot using observation-aware indexing, a 90-day latency and 5-year A → B gap, inverse-probability weighting, weighted competing-risk cumulative incidence, four prespecified comparators (primary: B with no prior A), and Benjamini-Hochberg false discovery rate (FDR) in a discovery/confirmation design. RESULTS: After discovery FDR, 663 unique trajectories advanced and 171 validated for ≥ 1 horizon in confirmation. At 5 years, the median risk ratio (RR) was 2.02 versus the primary comparator and 3.89 versus a calendar-time baseline (median absolute risk difference 2.49 percentage-points). Reverse-order checks were feasible for ∼ 87% of hypotheses (median A → B → C vs B → A → C RR 1.18) at 5 years. CONCLUSION: Ordered trajectories can provide interpretable, audit-ready signal beyond unordered features and can augment white-box risk estimation for adverse clinical outcomes.

Authors

Keywords

No keywords available for this article.