🛰️ Daily AI Frontier
‹ back to 2026-08-06

豆包视频通话率先上线「全模态全双工」,AI终于能边看、边听、边说了

WeChat: 夕小瑶科技说 Multimodal & Generative 2026-08-05
Representative image for 豆包视频通话率先上线「全模态全双工」,AI终于能边看、边听、边说了

TL;DR - ByteDance's Doubao app has shipped what it claims is the first consumer "omni-modal full-duplex" video calling experience, powered by a new native audio-video model called SeedRealtime. It matters because it extends real-time voice interaction (as in OpenAI's GPT-Live) with continuous vision, so the assistant can watch, listen, and speak simultaneously.

  • New base model: SeedRealtime, a native audio-video full-duplex LLM, replaces the prior stack; the article contrasts it with GPT-Live, which OpenAI documents as launching without video/screen-sharing support (audio-only full duplex), and with Thinking Machines' omni-modal work still at the research stage.
  • Interaction directionality: the model reportedly fuses speaker identity, face orientation, gaze, context and scene relations to decide whether an utterance is addressed to it — the author observed it stayed silent during side conversations with a friend, and that this ability disappeared when the camera was covered.
  • Turn-taking and proactivity: it waits through mid-sentence pauses and self-corrections instead of barging in after 2–3 seconds of silence, and conversely speaks up unprompted — e.g. flagging an item sorted into the wrong box under user-defined rules, with persistent task context and no repeated wake-up needed.
  • Concurrent multimodal grounding: while co-watching a podcast clip, it attributed statements to individual speakers (Papi酱, 罗翔, LKS), identified people by clothing, separated background audio (a hair dryer) from speech, and answered live questions without mistaking on-screen voices for the user.
  • Note: all claims are first-person product-trial impressions from a WeChat tech account, not benchmarked results.

view merged work →