🛰️ Daily AI Frontier
‹ back to 2026-09-09

Omni Interaction Agent Technical Report

arXiv eess.AS Multimodal & Generative Orantqing, Shengpeng Ji, Junlong Tong, Jialong Zuo, Dongjie Fu, Di Cao, Yangzhuo Li, Shangda Wu, Franz, Evan, Theron Veyra, Changhao Pan, Jingyu Lu, Dongchao Yang, Zhifei Xie, Yang Tan, Xiaoyu Shen, Xiaoda Yang, Wenfu Wang, Teddysun, Steveyves, Zhou Zhao, Bryanytian 2026-09-08
Representative image for Omni Interaction Agent Technical Report

TL;DR - Gander is an end-to-end multimodal agent designed for continuous, full-duplex interaction over streaming video, speech, and text. It combines low-latency conversation with higher-level reasoning and agentic workflows, allowing interruptions, proactive feedback, and follow-up questions.

  • A “Cerebellum-Brain” architecture separates real-time interaction from complex reasoning and agentic tasks, coordinating them through tool calls and an orchestration runtime.
  • Its streaming Thinker-Talker design represents chunked user inputs and model outputs as one ordered token stream to support continuous, low-latency exchanges.
  • Evaluations cover conversation, multimodal understanding, interactivity, and agentic intelligence, including noisy, multi-party, and backchannel scenarios.
  • The authors report competitive omni-interaction performance in internal human evaluations and release the models, code, and data.

view merged work →