Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds
TL;DR - JoyAI-Echo-1.5 is a unified audio-visual generation system designed for coherent long-form video and interactive worlds. It combines cross-shot memory, geometric camera control, and rollout-aware training to preserve characters, voices, and scene stability over long horizons.
- Its long-video variant aggregates visual history and speech-derived speaker cues to maintain character appearance and voice identity across shots.
- Its world-model variant translates diverse navigation inputs into metric 6-DoF camera trajectories for controller-agnostic viewpoint control.
- Progressive teacher forcing and Self-Gradient Forcing convert a bidirectional backbone into an efficient causal few-step generator.
- The system improves multiple long-video metrics over existing baselines and reports first place on WBench with an average score of 81.7.