M$^\text{4}$World: A Multi-view Multimodal Driving World Model for Interactive Object Manipulation and Minute-long Streaming
TL;DR — M⁴World is a multi-view, multimodal generative driving world model that synthesizes future surround-view video plus synchronized LiDAR, with interactive object-level control and stable minute-long streaming for scalable autonomous-driving simulation.
- Generates surround-view video streams and synchronized LiDAR scans; a flexible conditioning interface enables explicit control over individual objects' spatial layout and visual appearance.
- Achieves stable minute-long streaming via a multi-stage training framework enabling online causal generation in only four denoising steps while maintaining coherent long-rollout dynamics.
- Adds efficient few-clip post-training and visual reference-conditioned models for rare-case/long-tail customization, plus a VLM-based judging pipeline evaluating condition adherence, view-wise controllability, and cross-view object consistency.
- Reported to deliver high generation quality, precise controllability, and stability, supporting downstream long-tail augmentation and scene editing (specific quantitative metrics not provided in the abstract).