🛰️ Daily AI Frontier
‹ back to 2026-08-24

Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

arXiv cs.CV Multimodal & Generative Nan Duan, Haoyang Huang, Weiyang Jin, Haoran Li, Yaowei Li, Yuming Li, Yijun Liu, Xin Lu, Xiaoxiao Ma, Yanwen Ma, Yaofeng Su, Yilang Sun, Haoyu Wang, Zeyue Xue, Songchun Zhang, Junhao Zhuang 2026-08-24
Representative image for Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

TL;DR - JoyAI-Echo-1.5 is a unified audio-visual generation system designed for coherent long-form video and interactive worlds. It combines cross-shot memory, geometric camera control, and rollout-aware training to preserve characters, voices, and scene stability over long horizons.

  • Its long-video variant aggregates visual history and speech-derived speaker cues to maintain character appearance and voice identity across shots.
  • Its world-model variant translates diverse navigation inputs into metric 6-DoF camera trajectories for controller-agnostic viewpoint control.
  • Progressive teacher forcing and Self-Gradient Forcing convert a bidirectional backbone into an efficient causal few-step generator.
  • The system improves multiple long-video metrics over existing baselines and reports first place on WBench with an average score of 81.7.

view merged work →