🛰️ Daily AI Frontier
‹ back to 2026-09-04

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

Research Multimodal & Generative

Ranking

Overall 87
Content 95
Popularity 68

Observed public metrics from 1 member.

Representative image for Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

Merged summary

TL;DR - Puffin-World is a unified multimodal model that natively represents physics, geometry, and appearance to generate, reconstruct, and interact with physically consistent 3D worlds. Its integrated design supports closed-loop exploration without relying on external offline modules.

  • Jointly models gravity and latitude, depth, and images as native 3D world states.
  • Uses an Omni-Camera representation to support varied camera configurations, tasks, and motion patterns.
  • Couples future-view synthesis with geometry reconstruction while propagating physical dynamics across frames.
  • Scales training with Puffin-16M, containing 15 million vision-language-camera triplets and 1 million motion trajectories; code, models, and data were released.

Sources (1)

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

arXiv cs.CV Kang Liao, Yihang Luo, Xiao-Ming Wu, Linyi Jin, Size Wu, Chunyu Lin, Yao Zhao, Fei Wang, Wei Li, Chen Change Loy 2026-09-03 arXiv:2609.04196
Public signals Hugging Face upvotes 70
Providers: Hugging Face · Upvotes 70 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-25 14:24:04.981356 UTC

TL;DR - Puffin-World is a unified multimodal model that natively represents physics, geometry, and appearance to generate, reconstruct, and interact with physically consistent 3D worlds. Its integrated design supports closed-loop exploration without relying on external offline modules.

  • Jointly models gravity and latitude, depth, and images as native 3D world states.
  • Uses an Omni-Camera representation to support varied camera configurations, tasks, and motion patterns.
  • Couples future-view synthesis with geometry reconstruction while propagating physical dynamics across frames.
  • Scales training with Puffin-16M, containing 15 million vision-language-camera triplets and 1 million motion trajectories; code, models, and data were released.
item →