Diagnostic accuracy of a DenseNet-121 deep learning algorithm for chest radiograph triage in health assessment applicants: a prospective shadow-mode validation study in Nepal

Journal: medRxiv
Published Date:

Abstract

Objectives: To evaluate the diagnostic accuracy of a publicly available DenseNet-121 convolutional neural network (TorchXRayVision) for triaging chest radiographs of health assessment applicants at a tertiary hospital in Nepal. Design: Prospective, single-centre, shadow-mode diagnostic accuracy validation study. Reported in accordance with the STARD 2015 checklist and STARD-AI/DECIDE-AI guidelines. Setting: Department of Radiology and Imaging, Patan Academy of Health Sciences / Patan Hospital, Lalitpur, Nepal. Participants: 826 consecutive health assessment applicants (foreign employment Pre-Departure Medical Examination and student migration) undergoing chest radiography between 5 June and 20 June 2026. Two cases were excluded due to DICOM technical failure. Index test: DenseNet-121 algorithm (TorchXRayVision library, densenet121-res224-all pretrained weights). A maximum aggregated pathology probability score was derived per radiograph and compared against a post-hoc derived threshold of 0.6258 (selected as the highest threshold achieving the pre-specified >=95% sensitivity criterion). Reference standard: Single-reader-per-case review by one of three radiologists - two board-certified radiodiagnosticians (LS: 276 cases; DM: 275 cases) and one radiology resident (UB: 275 cases) - each blinded to AI output, using a standardised data collection worksheet capturing binary classification (abnormal/normal) and free-text findings. Results: Of 826 radiographs, 41 (4.97%) were classified as abnormal by the reference standard. At the post-hoc derived threshold of 0.6258, the DenseNet-121 algorithm achieved: sensitivity 95.12% (95% CI 83.9-98.7%), specificity 77.2% (95% CI 74.1-80.0%), area under the receiver operating characteristic curve (AUROC) 0.9583 (95% bootstrap CI 0.9225-0.9843), NPV 99.67% (95% Wilson CI 98.8-99.9%), PPV 17.89% (95% Wilson CI 13.4-23.5%), and Cohen's {kappa} 0.237 (95% bootstrap CI 0.174-0.304). Brier score was 0.3621 (null Brier 0.0472) and ECE was 0.564, confirming calibration failure due to score compression (range 0.52-0.72) despite preserved discrimination. Cross-validated results: Ten-fold cross-validation yielded bias-corrected sensitivity 95.12% (95% Wilson CI 83.9-98.7%; optimism 0.00 pp) and specificity 75.80% (95% Wilson CI 72.7-78.7%; optimism +1.40 pp), confirming primary metrics are not materially inflated by circular optimisation. Conclusions: The DenseNet-121 algorithm demonstrated high sensitivity and excellent discrimination for chest radiograph triage in a Nepali health-assessment population, supporting its potential as a rule-out tool (NPV 99.67%). Systematic score compression - preserved discrimination despite calibration shift - is a quantifiable marker of LMIC distributional shift. Prospective local calibration studies are warranted before operational deployment.

Authors

  • Shrestha
  • L.; Maharjan
  • D.; Bista
  • U.