Natural Language Processing to Identify Substance Use in Electronic Health Records: A Scoping Review.
Journal:
Substance use & addiction journal
Published Date:
Jul 21, 2026
Abstract
BACKGROUND: This scoping review aimed to characterize natural language processing (NLP) techniques deployed for identifying substance use in electronic health records (EHRs) and to compare the performance of these techniques by substance type. METHODS: We conducted a systematic search of PubMed, the Cochrane Library, Embase, Web of Science, ACM Digital Library, IEEE Xplore, and Scopus for peer-reviewed original research published in English. Studies were eligible if they applied NLP to identify non-prescription or problematic substance use in EHRs, provided full-text access, and reported quantitative performance metrics. RESULTS: A total of 86 studies met the inclusion criteria. NLP-identified substance use included tobacco (n = 42), alcohol (n = 26), non-prescription opioids (n = 30), cannabinoids (n = 9), stimulants (n = 6), and polysubstance use and/or other drugs (n = 16). NLP techniques included rule-based (n = 46), conventional machine learning (n = 39), deep learning (n = 17), and large language/transformer-based models (n = 22), with some studies applying multiple techniques. Annotation guidelines were available for 30 studies, and only 22 published their codes. Most studies reported performance metrics exceeding 0.80. CONCLUSIONS: Tobacco, opioids, and alcohol were the most frequently identified substances, whereas stimulants and cannabinoids were markedly underrepresented. Across the reviewed literature, transparency and reproducibility were limited, with few studies publishing code or detailed model specifications. Nevertheless, reported performance metrics were generally high. NLP techniques showed high performance in identifying tobacco, opioids, and alcohol use in EHRs, but stimulants and cannabinoids remain underrepresented. Transparency and reproducibility remain limited, underscoring the need for routine sharing of code, datasets, and model specifications.
Authors
Keywords
No keywords available for this article.