🛰️ Daily AI Frontier
‹ back to 2026-09-10

VLX-VR: An Agentic-Aware Video Reasoning Model

arXiv cs.CL Multimodal & Generative Sheng Li, Peng Liu, Qianqian Zhang, Tiancheng Zhao 2026-09-09
Representative image for VLX-VR: An Agentic-Aware Video Reasoning Model

TL;DR - VLX-VR is an agentic video-reasoning model that adaptively gathers multimodal evidence and manages memory through a Think–Memory–Observation loop. It reaches 78.79% accuracy on MINERVA while maintaining broadly stable performance across video durations.

  • Uses reinforcement learning on videos and agent trajectories to learn evidence acquisition, memory operations, and when to terminate reasoning.
  • Achieves 76.70%, 78.73%, and 80.92% accuracy across MINERVA’s three duration groups, with a cross-duration variance of 2.97 pp².
  • On correctly answered samples, 96.20% of reasoning traces align with MINERVA’s reference evidence and reasoning; 75.80% of all samples satisfy both correctness and trace-grounding criteria.
  • Counting, state changes, causal reasoning, and spatial perception remain challenging.

view merged work →