On the use of natural language processing to implement the target trial framework using unstructured data from the electronic health record.

Journal: Global epidemiology

Published Date: Jun 1, 2025

Abstract

The increasing availability and accessibility of electronic health record (EHR) data has made it a rich secondary source to conduct comparative effectiveness studies. To perform such studies, many researchers are turning to the target trial framework (TTF) to emulate the hypothetical randomized clinical trial. The quality of this emulation depends, in part, on the availability and accessibility of data for each component of the TTF. Yet one overarching challenge with using EHR data is that unstructured fields, such as clinical encounter notes, contain copious details on the patient yet require additional steps to extract if needed in the conduct of the study. Natural language processing (NLP) represents a spectrum of methods to assist with automating this extraction, from simpler rule-based methods to machine learning and artificial intelligence approaches that can handle complex language structures. What follows is a discussion on how NLP methods can augment information and data for researchers looking to estimate a treatment effect using EHR data via the TTF to emulate the hypothetical clinical trial. We conclude with recommendations for researchers interested in using NLP methods to obtain data stored in the free text of the EHR as well as considerations regarding the quality and validity of this data for the TTF.

Authors

Nicole Rafalko

Department of Epidemiology and Biostatistics, Drexel University Dornsife School of Public Health, Philadelphia, PA, USA.
Milena Gianfrancesco

Division of Rheumatology, Department of Medicine, University of California, San Francisco.
Neal D Goldstein

Department of Epidemiology and Biostatistics, Drexel University Dornsife School of Public Health, Philadelphia, PA, USA.

Keywords

No keywords available for this article.

External Resources

View on PubMed Access via DOI PubMed (40476041)

On the use of natural language processing to implement the target trial framework using unstructured data from the electronic health record.

Abstract

Authors

Keywords

External Resources

Popular Topics

Recent Journals

On the use of natural language processing to implement the target trial framework using unstructured data from the electronic health record.

Abstract

Authors

Keywords

External Resources

Stay Ahead of Medical AI

Popular Topics

Recent Journals