🛰️ Daily AI Frontier
‹ back to 2026-09-04

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

arXiv cs.CV Multimodal & Generative Kang Liao, Yihang Luo, Xiao-Ming Wu, Linyi Jin, Size Wu, Chunyu Lin, Yao Zhao, Fei Wang, Wei Li, Chen Change Loy 2026-09-03
Representative image for Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

TL;DR - Puffin-World is a unified multimodal model that natively represents physics, geometry, and appearance to generate, reconstruct, and interact with physically consistent 3D worlds. Its integrated design supports closed-loop exploration without relying on external offline modules.

  • Jointly models gravity and latitude, depth, and images as native 3D world states.
  • Uses an Omni-Camera representation to support varied camera configurations, tasks, and motion patterns.
  • Couples future-view synthesis with geometry reconstruction while propagating physical dynamics across frames.
  • Scales training with Puffin-16M, containing 15 million vision-language-camera triplets and 1 million motion trajectories; code, models, and data were released.

view merged work →