🛰️ Daily AI Frontier
‹ back to 2026-08-25

EchoWM: Open and Enterable Omnimodal World Models

arXiv cs.CV Multimodal & Generative Songchun Zhang, Yaowei Li, Junhao Zhuang, Weiyang Jin, Haoyu Wang, Xin Lu, Yilang Sun, Shiyi Zhang, Haoran Li, Xiaoxiao Ma, Yuming Li, Yijun Liu, Yaofeng Su, Yanwen Ma, Haoyu Wu, Zihan Su, Yue Ma, Lvmin Zhang, Haoyang Huang, Zeyue Xue, Anyi Rao, Nan Duan 2026-08-24
Representative image for EchoWM: Open and Enterable Omnimodal World Models

TL;DR - EchoWM is an omnimodal world model that turns continuous navigation inputs into interactive 720p video with synchronized environmental sound, music, and speech. It advances “enterable” generative media by supporting controlled, long-horizon first- and third-person experiences.

  • Unifies discrete commands and continuous camera poses as metric-scale relative 6-DoF trajectories.
  • Uses dataset-level calibration to preserve motion magnitude across heterogeneous training data.
  • Combines a complementary data engine, progressive training, and autoregressive post-training for audiovisual control and long-horizon generation.
  • Evaluations report strong trajectory following and visual quality, with synchronized sound and speech across varied scenes.

view merged work →