Teacher-Referenced Heatmap Coverage for Auditing Attribution Shifts in Updated CT Classifiers.

Journal: Journal of imaging informatics in medicine
Published Date:

Abstract

Model updates may change the attribution patterns of computed tomography (CT) classifiers without producing corresponding changes in predictive performance. This study evaluated whether peak-aware heatmap measures could provide a more informative signal than thresholded overlap for auditing such changes. Gradient-weighted class activation mapping (Grad-CAM) maps from baseline and decision-distilled ResNet18 classifiers were compared with a frozen, model-derived reference. The primary Heatmap Coverage Metric (HCM) endpoints measure missed circular coverage and nearest-peak distance from the reference to the evaluated map; the reverse direction was analyzed separately. Three held-out CT cohorts were evaluated at the slice level with three training seeds. Fixed-threshold intersection over union was frequently near zero. On the designated seed-42 tests, decision distillation increased circular discrepancy on Priv-EGFR by 0.083 (95% confidence interval, 0.032-0.137), decreased Euclidean discrepancy on Chest CT-Scan by 0.026 (95% confidence interval, - 0.037 to - 0.016 ), and produced inconclusive changes for both RADGEN endpoints. Both primary Chest endpoints were numerically lower under decision distillation in all three seeds, whereas Priv-EGFR and RADGEN showed endpoint- or seed-dependent behavior. Direction reversal changed some small or uncertain effects. Direction is an essential part of the HCM specification. The corrected endpoints complement thresholded overlap by detecting and characterizing update-related attribution shifts under a fixed audit protocol; they measure internal consistency rather than clinical correctness.

Authors

Keywords

No keywords available for this article.