🛰️ Daily AI Frontier
‹ back to 2026-07-22

OmniReasoner: Thinking with Long Audio-Video via Native Tool Use

Research Multimodal & Generative

Merged summary

TL;DR - OmniReasoner trains omnimodal LLMs to reason over long audio-video streams by selectively zooming into relevant time intervals. This improves accuracy and temporal grounding while focusing expensive high-fidelity processing on informative segments.

  • Uses supervised fine-tuning and reinforcement learning to teach models when and where to invoke a zoom-in tool.
  • Combines a low-cost global preview with dense local inspection of selected audio-video clips.
  • Introduces TimeAnchor to keep temporal tool arguments consistent across different sampling granularities.
  • Generates training trajectories without manual interval labels using synthetic video editing and composition.

Sources (1)

OmniReasoner: Thinking with Long Audio-Video via Native Tool Use

arXiv cs.CV Yu Chen, Caorui Li, Ziyu Xiong, Yidong Wang, Mingqi Gao, Shuman Liu, Biao Liu, Chunfeng Yang, Anxiang Zeng, Haibo Zhang, Chaofan Chen 2026-07-21 arXiv:2607.19339

TL;DR - OmniReasoner trains omnimodal LLMs to reason over long audio-video streams by selectively zooming into relevant time intervals. This improves accuracy and temporal grounding while focusing expensive high-fidelity processing on informative segments.

  • Uses supervised fine-tuning and reinforcement learning to teach models when and where to invoke a zoom-in tool.
  • Combines a low-cost global preview with dense local inspection of selected audio-video clips.
  • Introduces TimeAnchor to keep temporal tool arguments consistent across different sampling granularities.
  • Generates training trajectories without manual interval labels using synthetic video editing and composition.
item →