Natural language processing tool for extracting information about opioid overdoses in the USA from case narratives in the violent death reporting system.

Journal: Injury prevention : journal of the International Society for Child and Adolescent Injury Prevention
Published Date:

Abstract

BACKGROUND: Improving the infrastructure for drug overdose surveillance is critical for identifying new threats and responding to emerging trends. We aimed to develop a prototype tool using the principles of natural language processing that can extract information from the death records of drug overdose victims. METHODS: Data were obtained from the Violent Death Reporting System on drug overdose deaths. Narratives were manually labelled for 12 attributes of interest, totalling 82 labels about the circumstances of the overdose. Narratives were passed through the 'Excel Extractor' to identify and extract a target phrase and subsequently map the extracted phrase to predetermined code values. The output from the Excel Extractor was compared with manually labelled data to determine accuracy. Performance was compared against multiple machine learning models. RESULTS: The Excel Extractor performed well across the attributes of interest, achieving an F1 Score over 0.8 on nine of the 12 attributes. The Excel Extractor was the highest performing model on seven of the 12 attributes. The Excel Extractor achieved an F1 Score of 0.8 or higher on 46 of 82 (56%) of the labels, and a score of 0.9 or higher on nearly one-third (25 out of 82) of the labels. CONCLUSION: This work demonstrates it is feasible to develop a spreadsheet-formula-based natural language processing tool to accurately extract information about drug overdose deaths from narratives; for most attributes, a rule-based search performs well or better than deep learning. The Excel Extractor has the potential to streamline data abstraction for epidemiologists gathering data about drug overdose deaths.

Authors

Keywords

No keywords available for this article.