🛰️ Daily AI Frontier
‹ back to 2026-08-01

ReToken: One Token to Improve Vision-Language Models for Visual Retrieval

arXiv cs.CV Multimodal & Generative Yao Xiao, Reuben Tan, Zhen Zhu, Yuqun Wu, Jianfeng Gao, Derek Hoiem 2026-07-30

TL;DR - ReToken is a learnable embedding that retrieves sparse, query-relevant visual tokens from a pre-filled KV cache, helping vision-language models handle long image and video contexts efficiently. It delivers substantial benchmark gains while fitting training and long-video inference on one H100.

  • Uses a single explicit retrieval token to filter distracting visual context.
  • Improves Qwen3VL-8B and InternVL3.5 on Visual Haystacks by 13.4 and 12.4 points, respectively.
  • Transfers zero-shot to long-video retrieval, improving Qwen3VL-8B by 8.0 points on LVBench.
  • Requires training on only a small image-QA dataset.

view merged work →