EchoWM: Open and Enterable Omnimodal World Models
TL;DR - EchoWM is an omnimodal world model that turns continuous navigation inputs into interactive 720p video with synchronized environmental sound, music, and speech. It advances “enterable” generative media by supporting controlled, long-horizon first- and third-person experiences.
- Unifies discrete commands and continuous camera poses as metric-scale relative 6-DoF trajectories.
- Uses dataset-level calibration to preserve motion magnitude across heterogeneous training data.
- Combines a complementary data engine, progressive training, and autoregressive post-training for audiovisual control and long-horizon generation.
- Evaluations report strong trajectory following and visual quality, with synchronized sound and speech across varied scenes.