🛰️ Daily AI Frontier
‹ back to 2026-08-29

SpatialCrafter: Single Image World Modeling with Generative 3D Proxies

arXiv cs.CV Multimodal & Generative Chuan Fang, Lingteng Qiu, Yixun Liang, Rui Chen, Kunming Luo, Zhaohua Zheng, Tongyuan Bai, Feipeng Tian, Zilong Dong, Zihan Zhou, Ping Tan 2026-08-27
Representative image for SpatialCrafter: Single Image World Modeling with Generative 3D Proxies

TL;DR - SpatialCrafter generates explorable 3D scenes from a single image by first constructing a globally consistent 3D proxy, then refining it with photorealistic details. This two-stage approach reduces hallucination and long-term geometric drift under large viewpoint changes.

  • Its Point-anchored Sparse Structure Flow module predicts a spatially aligned, geometrically consistent 3D proxy.
  • A video diffusion model acts as a Generative Deferred Refiner, adding high-frequency appearance details while following the proxy geometry.
  • Parallel Geometry Injection and Proxy-Aware Corruption improve integration with pretrained video diffusion models and robustness to proxy artifacts.
  • The authors introduce a 115K-scene hybrid dataset and report improvements over prior methods on synthetic and real-world data.

view merged work →