🛰️ Daily AI Frontier
‹ back to 2026-08-01

ReToken: One Token to Improve Vision-Language Models for Visual Retrieval

Research Multimodal & Generative

Ranking

Overall 86
Content 95
Popularity 65

Observed public metrics from 1 member.

Merged summary

TL;DR - ReToken is a learnable embedding that retrieves sparse, query-relevant visual tokens from a pre-filled KV cache, helping vision-language models handle long image and video contexts efficiently. It delivers substantial benchmark gains while fitting training and long-video inference on one H100.

  • Uses a single explicit retrieval token to filter distracting visual context.
  • Improves Qwen3VL-8B and InternVL3.5 on Visual Haystacks by 13.4 and 12.4 points, respectively.
  • Transfers zero-shot to long-video retrieval, improving Qwen3VL-8B by 8.0 points on LVBench.
  • Requires training on only a small image-QA dataset.

Sources (1)

ReToken: One Token to Improve Vision-Language Models for Visual Retrieval

arXiv cs.CV Yao Xiao, Reuben Tan, Zhen Zhu, Yuqun Wu, Jianfeng Gao, Derek Hoiem 2026-07-30 arXiv:2607.28627
Public signals Hugging Face upvotes 9
Providers: Hugging Face · Upvotes 9 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-08-31 14:30:08.084581 UTC

TL;DR - ReToken is a learnable embedding that retrieves sparse, query-relevant visual tokens from a pre-filled KV cache, helping vision-language models handle long image and video contexts efficiently. It delivers substantial benchmark gains while fitting training and long-video inference on one H100.

  • Uses a single explicit retrieval token to filter distracting visual context.
  • Improves Qwen3VL-8B and InternVL3.5 on Visual Haystacks by 13.4 and 12.4 points, respectively.
  • Transfers zero-shot to long-video retrieval, improving Qwen3VL-8B by 8.0 points on LVBench.
  • Requires training on only a small image-QA dataset.
item →