VLX-VR: An Agentic-Aware Video Reasoning Model
Ranking
Overall
78
Content
95
Popularity
37
Observed public metrics from 1 member.
Merged summary
TL;DR - VLX-VR is an agentic video-reasoning model that adaptively gathers multimodal evidence and manages memory through a Think–Memory–Observation loop. It reaches 78.79% accuracy on MINERVA while maintaining broadly stable performance across video durations.
- Uses reinforcement learning on videos and agent trajectories to learn evidence acquisition, memory operations, and when to terminate reasoning.
- Achieves 76.70%, 78.73%, and 80.92% accuracy across MINERVA’s three duration groups, with a cross-duration variance of 2.97 pp².
- On correctly answered samples, 96.20% of reasoning traces align with MINERVA’s reference evidence and reasoning; 75.80% of all samples satisfy both correctness and trace-grounding criteria.
- Counting, state changes, causal reasoning, and spatial perception remain challenging.
Sources (1)
VLX-VR: An Agentic-Aware Video Reasoning Model
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - VLX-VR is an agentic video-reasoning model that adaptively gathers multimodal evidence and manages memory through a Think–Memory–Observation loop. It reaches 78.79% accuracy on MINERVA while maintaining broadly stable performance across video durations.
- Uses reinforcement learning on videos and agent trajectories to learn evidence acquisition, memory operations, and when to terminate reasoning.
- Achieves 76.70%, 78.73%, and 80.92% accuracy across MINERVA’s three duration groups, with a cross-duration variance of 2.97 pp².
- On correctly answered samples, 96.20% of reasoning traces align with MINERVA’s reference evidence and reasoning; 75.80% of all samples satisfy both correctness and trace-grounding criteria.
- Counting, state changes, causal reasoning, and spatial perception remain challenging.