🛰️ Daily AI Frontier
‹ back to 2026-09-18

MM-Future: Multi-Mode Joint World-Action Modeling for Autonomous Driving

Research Autonomous Driving

Ranking

Overall 82
Content 95
Popularity N/A

No observed public metrics; popularity remains neutral/archived.

Representative image for MM-Future: Multi-Mode Joint World-Action Modeling for Autonomous Driving

Merged summary

TL;DR - MM-Future is a world-action model for autonomous driving that jointly generates multiple possible future scenes and corresponding driving actions. Its paired, multimodal predictions improve planning performance under uncertainty compared with single-mode and action-only approaches.

  • Initializes each hypothesis from a structured action prior and an independent future-scene source, then co-evolves them with a modality-aware diffusion Transformer.
  • Compresses multi-view video into planning-oriented “MM-Tokens” for efficient multi-mode rollout.
  • Uses a future-conditioned scorer to rank trajectory proposals based on shared history and each proposal’s predicted future.
  • Reports 94.0 PDMS and 91.5 EPDMS on NAVSIM navtest, plus a 32.3 HD-Score in zero-shot closed-loop HUGSIM evaluation.

Sources (1)

MM-Future: Multi-Mode Joint World-Action Modeling for Autonomous Driving

arXiv cs.CV Shuai Liu, Hechangle Gong, Hao Jiang, Runlin He, Junxiang Zhan, Kai Huang, Sheng Yang, Shaoqing Ren 2026-09-17 arXiv:2609.20377
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-26 14:15:27.300510 UTC

TL;DR - MM-Future is a world-action model for autonomous driving that jointly generates multiple possible future scenes and corresponding driving actions. Its paired, multimodal predictions improve planning performance under uncertainty compared with single-mode and action-only approaches.

  • Initializes each hypothesis from a structured action prior and an independent future-scene source, then co-evolves them with a modality-aware diffusion Transformer.
  • Compresses multi-view video into planning-oriented “MM-Tokens” for efficient multi-mode rollout.
  • Uses a future-conditioned scorer to rank trajectory proposals based on shared history and each proposal’s predicted future.
  • Reports 94.0 PDMS and 91.5 EPDMS on NAVSIM navtest, plus a 32.3 HD-Score in zero-shot closed-loop HUGSIM evaluation.
item →