CARE-MVLM: Counterfactual Abstention and Region-Grounded Evidence in Mammography Vision-Language Models

Journal: medRxiv
Published Date:

Abstract

Mammography interpretation requires tying each finding to supporting evidence and withholding judgment when that evidence is unavailable. Vision-language models are increasingly used for image-based medical question answering, yet existing benchmarks suggest that they struggle to jointly support accurate answers, evidence grounding, and appropriate abstention. We propose CARE-MVLM, built on Qwen2.5-VL-7B-Instruct, to address this gap through separate modules for answer prediction, answer-conditioned evidence generation, and answerability selection. On 7,358 triplets curated from CBIS-DDSM, MIAS, and VinDr-Mammo, CARE-MVLM substantially improves lesion-sensitive selective behavior over same VLM backbone baselines and exhibits markedly stronger lesion-sensitive abstention than GPT-5.6 and Gemini-3.5-flash while attaining the highest joint answer, grounding, and abstention reliability in the comparison.

Authors

  • Qu
  • B.; Liu
  • W.; Murrow
  • M.; Burger
  • M.; Guo
  • X.; Vaidya
  • M. S.; Rose
  • S. L.; Kantarcioglu
  • M.; Malin
  • B. A.; Yin
  • Z.