OmniReasoner: Thinking with Long Audio-Video via Native Tool Use
Merged summary
TL;DR - OmniReasoner trains omnimodal LLMs to reason over long audio-video streams by selectively zooming into relevant time intervals. This improves accuracy and temporal grounding while focusing expensive high-fidelity processing on informative segments.
- Uses supervised fine-tuning and reinforcement learning to teach models when and where to invoke a zoom-in tool.
- Combines a low-cost global preview with dense local inspection of selected audio-video clips.
- Introduces TimeAnchor to keep temporal tool arguments consistent across different sampling granularities.
- Generates training trajectories without manual interval labels using synthetic video editing and composition.
Sources (1)
OmniReasoner: Thinking with Long Audio-Video via Native Tool Use
TL;DR - OmniReasoner trains omnimodal LLMs to reason over long audio-video streams by selectively zooming into relevant time intervals. This improves accuracy and temporal grounding while focusing expensive high-fidelity processing on informative segments.
- Uses supervised fine-tuning and reinforcement learning to teach models when and where to invoke a zoom-in tool.
- Combines a low-cost global preview with dense local inspection of selected audio-video clips.
- Introduces TimeAnchor to keep temporal tool arguments consistent across different sampling granularities.
- Generates training trajectories without manual interval labels using synthetic video editing and composition.