🛰️ Daily AI Frontier
‹ back to 2026-07-19

Region-Grounded Vision-Language Learning for Detection-Guided Mammographic Lesion Classification

arXiv cs.CV Medical/Healthcare AI Zhengbo Zhou, Jiren Li, Dooman Arefan, Margarita Zuley, Shandong Wu 2026-07-17

TL;DR - A region-grounded vision-language method for mammographic lesion classification that aligns localized lesion features with clinical text descriptors, mirroring how radiologists work. It matters because standard whole-image contrastive alignment dilutes the small, subtle lesion cues critical for malignancy assessment.

  • Region-text contrastive pretraining aligns lesion-specific features with structured clinical descriptors from radiology metadata, rather than aligning at the whole-image level.
  • Multi-component objective combats semantic collapse and background bias in low-vocabulary settings via positive alignment, fine-grained semantic hard negatives, and background suppression.
  • Auxiliary lesion detection head is jointly optimized with contrastive classification to preserve spatial sensitivity and enable localization-aware malignancy classification.
  • Evaluated on CBIS-DDSM and VinDr-Mammo, reporting superior performance vs. related methods across in-domain, cross-dataset, and transfer-learning settings (specific metrics not provided).

view merged work →