Cross-cohort Mamography]{Representation-dependent domain adaptation for cross-cohort generalization of mammography diagnostic models

Journal: bioRxiv
Published Date:

Abstract

Background: Artificial intelligence (AI) models for mammography can achieve high diagnostic performance when training and test data originate from similar populations and imaging environments, but performance often deteriorates under cross- cohort domain shift. Differences in acquisition systems, image characteristics, preprocessing, population composition, and disease prevalence can introduce distributional variation that is poorly captured by conventional within- dataset validation. Recent studies have further shown that Foundation Model representations can retain substantial dataset-specific information. However, it remains unclear how image representation influences diagnostic transfer to completely unseen cohorts and whether domain-adversarial learning provides comparable benefits across different representation types. Methods: We systematically evaluated cross-cohort mammography classification using seven independent datasets spanning diverse geographic, institutional, scanner, and imaging environments. Three image representations were compared: handcrafted radiomic features, EfficientNet-B0-derived features, and embeddings from the medical vision-language model MedSigLIP-448. Each was evaluated under three preprocessing conditions: unprocessed images, Otsu-based breast-region extraction, and Otsu extraction combined with Contrast Limited Adaptive Histogram Equalization (CLAHE). Patient-level classification was performed using a mean-pooled multilayer perceptron (MLP) baseline and an attention-based Domain-Adversarial Neural Network (DANN). Generalization was evaluated using a multi-source Leave-One-Domain-Out (LODO) framework in which six datasets were used for model development and the seventh was completely withheld as an unseen test cohort, iteratively across all seven datasets. Performance was primarily assessed by area under the receiver operating characteristic curve (AUC-ROC). Results: Cross-cohort performance depended strongly on image representation and model architecture. MedSigLIP consistently provided the highest performance on unseen cohorts. With MedSigLIP features, mean LODO AUC-ROC increased from 0.598 with MLP to 0.685 with DANN for unprocessed images, from 0.607 to 0.678 after Otsu preprocessing, and from 0.603 to 0.659 after CLAHE+Otsu preprocessing. When DANN-MLP differences were evaluated across held-out datasets and preprocessing conditions, a significant gain was observed for MedSigLIP after correction for multiple comparisons, whereas corresponding gains were not significant for Radiomics or EfficientNet-B0. A mixed-effects analysis similarly showed a positive MedSigLIP X DANN interaction, although this narrowly missed conventional statistical significance. Unprocessed images produced the highest numerical mean performance for MedSigLIP-DANN, but preprocessing itself had no significant overall effect. Notably, linear probing showed strong dataset discriminability in MedSigLIP and EfficientNet representations but little linearly recoverable dataset information in Radiomics, indicating that low dataset separability alone did not predict better diagnostic generalization. Conclusion: Cross-cohort mammography generalization depends strongly on the choice of image representation and its interaction with the classification strategy. MedSigLIP provided the most transferable representation in our multi-source unseen-domain evaluation, and the performance gain associated with DANN was concentrated in MedSigLIP rather than consistently observed across representation types. These findings suggest that the effectiveness of domain-adversarial learning is representation-dependent and that neither conventional preprocessing nor low dataset discriminability alone ensures robust transfer to unseen clinical cohorts.

Authors

  • Pandey
  • A. K.; Ahmad
  • S.

Categories