VLX-VR: An Agentic-Aware Video Reasoning Model
TL;DR - VLX-VR is an agentic video-reasoning model that adaptively gathers multimodal evidence and manages memory through a Think–Memory–Observation loop. It reaches 78.79% accuracy on MINERVA while maintaining broadly stable performance across video durations.
- Uses reinforcement learning on videos and agent trajectories to learn evidence acquisition, memory operations, and when to terminate reasoning.
- Achieves 76.70%, 78.73%, and 80.92% accuracy across MINERVA’s three duration groups, with a cross-duration variance of 2.97 pp².
- On correctly answered samples, 96.20% of reasoning traces align with MINERVA’s reference evidence and reasoning; 75.80% of all samples satisfy both correctness and trace-grounding criteria.
- Counting, state changes, causal reasoning, and spatial perception remain challenging.