High-Throughput Observational Evidence Generation Using Linked Electronic Health Record and Claims Data

Journal: medRxiv
Published Date:

Abstract

Background: Many of the most consequential treatment decisions concern patients and comparisons that randomized trials never address: off-label and head-to-head choices in complex, comorbid patients who are routinely excluded from trials, post-approval safety questions, and treatment-effect differences across subgroups. For these questions the problem is not merely the absence of a trial but the absence of any comparative evidence at all. Observational data contain these patients, yet existing analyses are hard to synthesize: discordant findings stem from differences in cohort definitions, confounder capture, and follow-up windows, and studies typically report a narrow subset of outcomes. Even when these are addressed, past systems have stopped after the data preparation step, limiting usability of outputs. Methods: We developed a high-throughput evidence-generation workflow using linked EHR and claims data. The cornerstone is a pre-specified measurement architecture optimized for causal analysis, applied uniformly across clinical scenarios: three post-index windows (acute to two-year follow-up); 28 comorbidities; 14 healthcare resource utilization (HCRU) categories; 30 laboratory measures with 57 binary thresholds; and 43 adverse event categories. We leveraged machine learning to generate confounding-adjusted actionable treatment comparisons using scalable collaborative targeted learning (C-TMLE), with unadjusted comparisons as a transparency benchmark, across ~1,038 outcomes per scenario, including effect-measure modification (EMM) assessments on up to 185 baseline features per scenario. Results: Across 135 clinical scenarios, the workflow produced 210,584,509 outcome evaluations. An evaluation is one outcome x follow-up window x treatment contrast x population stratum x estimator, with uncertainty bounds and supporting diagnostics. These evaluations were synthesized into 7,500 narrative summaries which underwent structured clinical and statistical quality control. Conclusions. Standardized, high-throughput workflows can shift evidence generation away from fragmented studies toward comprehensive evidence packages. This shared evidence base supports precision medicine by making treatment-effect heterogeneity visible across clinically meaningful subpopulations, thereby reducing the need for redundant, stakeholder-specific studies.

Authors

  • Gombar
  • S.; Shah
  • N.; Sanghavi
  • N.; Coyle
  • J.; Mukerji
  • A.; Chappelka
  • M.