🛰️ Daily AI Frontier
‹ back to 2026-09-18

MM-Future: Multi-Mode Joint World-Action Modeling for Autonomous Driving

arXiv cs.CV Autonomous Driving Shuai Liu, Hechangle Gong, Hao Jiang, Runlin He, Junxiang Zhan, Kai Huang, Sheng Yang, Shaoqing Ren 2026-09-17
Representative image for MM-Future: Multi-Mode Joint World-Action Modeling for Autonomous Driving

TL;DR - MM-Future is a world-action model for autonomous driving that jointly generates multiple possible future scenes and corresponding driving actions. Its paired, multimodal predictions improve planning performance under uncertainty compared with single-mode and action-only approaches.

  • Initializes each hypothesis from a structured action prior and an independent future-scene source, then co-evolves them with a modality-aware diffusion Transformer.
  • Compresses multi-view video into planning-oriented “MM-Tokens” for efficient multi-mode rollout.
  • Uses a future-conditioned scorer to rank trajectory proposals based on shared history and each proposal’s predicted future.
  • Reports 94.0 PDMS and 91.5 EPDMS on NAVSIM navtest, plus a 32.3 HD-Score in zero-shot closed-loop HUGSIM evaluation.

view merged work →