Natural language processing for detecting endoscopy-related adverse events in free-text procedure reports at a tertiary referral centre in Türkiye: a 10-year retrospective diagnostic accuracy study.
Journal:
BMJ open
Published Date:
Sep 3, 2026
Abstract
OBJECTIVES: Automated surveillance of endoscopy-related adverse events (AEs) from non-English free-text reports remains insufficiently evaluated. We compared rule-based and transformer-based natural language processing (NLP) for classifying index procedure reports according to 30-day clinician-adjudicated AE status. DESIGN: Retrospective, single-centre diagnostic accuracy study. SETTING: Tertiary endoscopy unit in Türkiye, August 2015 to August 2025. PARTICIPANTS: The source cohort comprised 140 385 reports. All 1512 lexicon-positive and 2000 randomly sampled lexicon-negative reports underwent clinician adjudication, identifying 1208 AE-positive reports. A 13 333-report corpus was allocated at patient level to training (n=9333), validation (n=2000) and a locked, fully adjudicated test set (n=2000; 180 AE-positive). PRIMARY AND SECONDARY OUTCOME MEASURES: The primary outcome was confirmation of an attributable AE within 30 days. Models analysed the index report only, whereas adjudication included subsequent documentation. Secondary outcomes were AE category and documentation timing. Diagnostic accuracy measures included sensitivity, specificity, predictive values, F1 score and transformer precision-recall area under the curve (PR-AUC). RESULTS: The rule-based model achieved 84.4% sensitivity (95% CI 78.4% to 89.0%) and 99.5% specificity (95% CI 99.1% to 99.7%). The transformer achieved higher sensitivity (92.2%, 95% CI 87.4% to 95.3%; McNemar's test, p=0.01) and 98.4% specificity (95% CI 97.7% to 98.9%). Positive predictive values were 94.4% and 85.1%, and negative predictive values were 98.5% and 99.2%, for the rule-based and transformer models, respectively; both had F1 scores of 0.89. Transformer PR-AUC was 0.91 (95% CI 0.87 to 0.94). The transformer reduced false negatives from 28 to 14 but increased false positives from 9 to 29. Cohort-wide AE incidence was not estimated because lexicon-negative reports were only partially verified. CONCLUSIONS: Both approaches had high specificity. The transformer reduced false-negative classifications but increased false-positive classifications. These findings support clinician-supervised retrospective case identification and quality assurance, but not autonomous diagnosis, prospective prediction or point-of-care use. External validation is required.
Authors
Keywords
No keywords available for this article.