李飞飞发布:全球首个多模态世界模型
TL;DR - World Labs introduced Atlas, a multimodal world model that generates camera-controlled imagery, reconstructs 3D scenes, and simulates spatial-temporal environments from images or video. It could support applications ranging from visual effects to scalable real-to-sim training for robots.
- Atlas uses a multimodal autoregressive diffusion Transformer to process text, images, camera poses, and depth maps within a shared 3D spatial context.
- From one or more images, it can synthesize new views, output explicit 3D representations, and generate up to one minute of 1440p camera-controlled video.
- World Labs reports that Atlas outperformed evaluated state-of-the-art video models on camera-controlled generation and specialized open-source models on sparse-view 3D reconstruction.
- The model can produce realistic RGB and depth observations from a few photos, enabling varied simulated environments for robot training and testing; early access is limited to selected partners.