SeVeR: Selective Visual Exposure and Retrieval for 3D Medical Image Question Answering
TL;DR - SeVeR is a selective visual retrieval framework for 3D medical visual question answering that reduces redundant visual-token exposure while improving answer performance. The work also introduces BreMRIs-VQA, a large, clinically curated breast MRI benchmark.
- BreMRIs-VQA contains 1.19 million free-text and multiple-choice QA pairs from 71,000 MRI sequences across 12,900 patients.
- SeVeR compresses dense 3D volumes into modality-specific prototypes, then retrieves complementary evidence at multiple levels during decoding.
- Change-aware gated attention and a marginal-utility self-consistency objective suppress retrieval that does not improve reasoning.
- Experiments on BreMRIs-VQA and public benchmarks report better discriminative and generative performance with substantially fewer exposed visual tokens.