MM-Future: Multi-Mode Joint World-Action Modeling for Autonomous Driving
Ranking
Overall
82
Content
95
Popularity
N/A
No observed public metrics; popularity remains neutral/archived.
Merged summary
TL;DR - MM-Future is a world-action model for autonomous driving that jointly generates multiple possible future scenes and corresponding driving actions. Its paired, multimodal predictions improve planning performance under uncertainty compared with single-mode and action-only approaches.
- Initializes each hypothesis from a structured action prior and an independent future-scene source, then co-evolves them with a modality-aware diffusion Transformer.
- Compresses multi-view video into planning-oriented “MM-Tokens” for efficient multi-mode rollout.
- Uses a future-conditioned scorer to rank trajectory proposals based on shared history and each proposal’s predicted future.
- Reports 94.0 PDMS and 91.5 EPDMS on NAVSIM navtest, plus a 32.3 HD-Score in zero-shot closed-loop HUGSIM evaluation.
Sources (1)
MM-Future: Multi-Mode Joint World-Action Modeling for Autonomous Driving
Public signals
N/A
TL;DR - MM-Future is a world-action model for autonomous driving that jointly generates multiple possible future scenes and corresponding driving actions. Its paired, multimodal predictions improve planning performance under uncertainty compared with single-mode and action-only approaches.
- Initializes each hypothesis from a structured action prior and an independent future-scene source, then co-evolves them with a modality-aware diffusion Transformer.
- Compresses multi-view video into planning-oriented “MM-Tokens” for efficient multi-mode rollout.
- Uses a future-conditioned scorer to rank trajectory proposals based on shared history and each proposal’s predicted future.
- Reports 94.0 PDMS and 91.5 EPDMS on NAVSIM navtest, plus a 32.3 HD-Score in zero-shot closed-loop HUGSIM evaluation.