🛰️ Daily AI Frontier
‹ back to 2026-08-27

SeVeR: Selective Visual Exposure and Retrieval for 3D Medical Image Question Answering

Research Medical/Healthcare AI

Ranking

Overall 83
Content 100
Popularity 42

Observed public metrics from 1 member.

Representative image for SeVeR: Selective Visual Exposure and Retrieval for 3D Medical Image Question Answering

Merged summary

TL;DR - SeVeR is a selective visual retrieval framework for 3D medical visual question answering that reduces redundant visual-token exposure while improving answer performance. The work also introduces BreMRIs-VQA, a large, clinically curated breast MRI benchmark.

  • BreMRIs-VQA contains 1.19 million free-text and multiple-choice QA pairs from 71,000 MRI sequences across 12,900 patients.
  • SeVeR compresses dense 3D volumes into modality-specific prototypes, then retrieves complementary evidence at multiple levels during decoding.
  • Change-aware gated attention and a marginal-utility self-consistency objective suppress retrieval that does not improve reasoning.
  • Experiments on BreMRIs-VQA and public benchmarks report better discriminative and generative performance with substantially fewer exposed visual tokens.

Sources (1)

SeVeR: Selective Visual Exposure and Retrieval for 3D Medical Image Question Answering

arXiv cs.CV Yaojun Hu, Danyang Tu, Yang Liu, Jiajin Zhang, Wei Fang, Zhiqiang Liu, Chunlai Dong, Yingda Xia, Haochao Ying, Jian Wu, Ling Zhang 2026-08-26 arXiv:2608.25630
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-29 14:13:41.580883 UTC

TL;DR - SeVeR is a selective visual retrieval framework for 3D medical visual question answering that reduces redundant visual-token exposure while improving answer performance. The work also introduces BreMRIs-VQA, a large, clinically curated breast MRI benchmark.

  • BreMRIs-VQA contains 1.19 million free-text and multiple-choice QA pairs from 71,000 MRI sequences across 12,900 patients.
  • SeVeR compresses dense 3D volumes into modality-specific prototypes, then retrieves complementary evidence at multiple levels during decoding.
  • Change-aware gated attention and a marginal-utility self-consistency objective suppress retrieval that does not improve reasoning.
  • Experiments on BreMRIs-VQA and public benchmarks report better discriminative and generative performance with substantially fewer exposed visual tokens.
item →