Real-world deployment of machine learning models for opioid overdose and opioid use disorder: a systematic review of clinical and operational lessons for addiction medicine.
Journal:
Journal of addictive diseases
Published Date:
Jan 13, 2026
Abstract
BACKGROUND: Machine learning (ML) is increasingly explored for opioid overdose and opioid use disorder (OUD) detection and prevention. Regional burden is uneven: the United States currently has among the highest rates of drug-overdose deaths worldwide, underscoring the urgent need for operationalized ML tools in US clinical and public-health settings. While many models exist, few have been tested in live settings. Understanding how deployed systems perform and are governed is essential for safe, equitable use in addiction medicine. METHODS: We systematically searched PubMed, Embase, IEEE Xplore, and Scopus from January 2019 to August 2025 for studies describing real-world ML deployments for overdose or OUD. Eligible studies required live integration into clinical, public health, or consumer workflows with prospective or concurrent evaluation. Risk of bias and operational robustness were assessed using the PROBAST-R framework (a PROBAST extension for deployed ML; 'risk of bias' here denotes systematic error in reported performance due to study design, data, or reporting). RESULTS: Fifteen studies were included, spanning health systems, public health surveillance, emergency services, and wearable detection. Most achieved good discrimination (AUC > 0.80) or high precision, though tradeoffs between sensitivity and specificity were common. Governance structures were more consistently reported in large deployments, yet no study described automated monitoring, and only two examined subgroup performance. Fairness audits and human-factors evaluations were rare. CONCLUSION: Deployed ML for opioid outcomes is feasible but uneven in maturity. Clinical implications: addiction medicine clinicians should (1) request subgroup performance results before adoption; (2) confirm post-deployment monitoring plans for drift and calibration; and (3) use human-in-the-loop safeguards (clinician or human verification before automated actions) to reduce harms from false positives/negatives.
Authors
Keywords
No keywords available for this article.