🛰️ Daily AI Frontier
‹ back to 2026-07-29

Wonder: Video World Model Done Better

arXiv cs.CV Multimodal & Generative Jiacong Xu, Hanwen Jiang, Zhixin Shu, Kalyan Sunkavalli, Vishal M. Patel, Yiqun Mei 2026-07-28
Representative image for Wonder: Video World Model Done Better

TL;DR - Wonder is a real-time video world model that turns images or videos into camera-controllable environments. Its coordinated control, memory, and distillation techniques enable coherent minute-scale exploration at 16 FPS.

  • Dense coordinate-field conditioning provides spatially aligned camera-motion and orientation cues.
  • Sparse-attention memory retrieves relevant context tokens efficiently regardless of total context length.
  • An improved self-forcing distillation pipeline better preserves control adherence, generation diversity, and long-term memory.
  • Supports both image-to-video world creation and real-time re-shooting of existing dynamic scenes.

view merged work →