🛰️ Daily AI Frontier
‹ back to 2026-08-27

Retrieval Heads Meet Vision: Uncovering How VLMs Locate and Extract Visual Information

Research Multimodal & Generative

Ranking

Overall 78
Content 95
Popularity 37

Observed public metrics from 1 member.

Merged summary

TL;DR - The paper identifies Visual Retrieval Heads (VRHs), a small subset of attention heads that enable vision-language models to connect text descriptions to relevant image regions. Their causal importance and transfer across tasks and related architectures reveal a shared, sparse mechanism for visual grounding.

  • VRHs comprise roughly 1.7–2.6% of attention heads across 11 VLMs and five referring-expression benchmarks.
  • Masking the top 20 VRHs reduces grounding accuracy by up to 80 percentage points, whereas masking random heads has little effect.
  • Heads discovered through bounding-box prediction remain causal for attribute, spatial, counting, and visual-math tasks.
  • VRHs preserve output formatting while disrupting localization and transfer between VLMs sharing an LLM backbone despite differences in their visual components and instruction tuning.

Sources (1)

Retrieval Heads Meet Vision: Uncovering How VLMs Locate and Extract Visual Information

arXiv cs.CV Chanho Park, Daehyeon Choi, Jihyun Lee, Minhyuk Sung 2026-08-27 arXiv:2608.27417
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-21 14:30:15.341633 UTC

TL;DR - The paper identifies Visual Retrieval Heads (VRHs), a small subset of attention heads that enable vision-language models to connect text descriptions to relevant image regions. Their causal importance and transfer across tasks and related architectures reveal a shared, sparse mechanism for visual grounding.

  • VRHs comprise roughly 1.7–2.6% of attention heads across 11 VLMs and five referring-expression benchmarks.
  • Masking the top 20 VRHs reduces grounding accuracy by up to 80 percentage points, whereas masking random heads has little effect.
  • Heads discovered through bounding-box prediction remain causal for attribute, spatial, counting, and visual-math tasks.
  • VRHs preserve output formatting while disrupting localization and transfer between VLMs sharing an LLM backbone despite differences in their visual components and instruction tuning.
item →