🛰️ Daily AI Frontier
‹ back to 2026-08-25

Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents

arXiv cs.CV Multimodal & Generative Wenqi Liu, Shijie Ma, Yunxiao Wang, Meng Liu, Qile Su, Han Liu, Bohan Hou, Xuanyu Zheng, Changyi Liu, Tianke Zhang, Haonan Fan, Kaiyu Jiang, Yingxin Li, Jiankang Chen, Xu Wang, Bin Wen, Tingting Gao, Han Li, Jianhua Yin, Yinwei Wei, Xuemeng Song 2026-08-24

TL;DR - VideoRover is an open-world video agent that coordinates targeted video inspection, multimodal search, and web browsing to answer questions requiring both sparse visual evidence and external knowledge. Its unified approach shows that active grounding, retrieval, and long-horizon reinforcement learning provide complementary gains.

  • Iteratively chooses among video cropping, multimodal search, and webpage browsing based on prior tool results.
  • Uses an automated curation pipeline yielding 26K verified supervised fine-tuning trajectories and 3K challenging reinforcement-learning instances.
  • Introduces VideoRover-Bench, stratified by video duration and research difficulty.
  • VideoRover-8B-RL matches proprietary models in direct answering without tools and outperforms larger open-source models given the same tool suite on the evaluated benchmarks.

view merged work →