🛰️ Daily AI Frontier
‹ back to 2026-08-29

SpatialCrafter: Single Image World Modeling with Generative 3D Proxies

Research Multimodal & Generative

Ranking

Overall 80
Content 95
Popularity 45

Observed public metrics from 1 member.

Representative image for SpatialCrafter: Single Image World Modeling with Generative 3D Proxies

Merged summary

TL;DR - SpatialCrafter generates explorable 3D scenes from a single image by first constructing a globally consistent 3D proxy, then refining it with photorealistic details. This two-stage approach reduces hallucination and long-term geometric drift under large viewpoint changes.

  • Its Point-anchored Sparse Structure Flow module predicts a spatially aligned, geometrically consistent 3D proxy.
  • A video diffusion model acts as a Generative Deferred Refiner, adding high-frequency appearance details while following the proxy geometry.
  • Parallel Geometry Injection and Proxy-Aware Corruption improve integration with pretrained video diffusion models and robustness to proxy artifacts.
  • The authors introduce a 115K-scene hybrid dataset and report improvements over prior methods on synthetic and real-world data.

Sources (1)

SpatialCrafter: Single Image World Modeling with Generative 3D Proxies

arXiv cs.CV Chuan Fang, Lingteng Qiu, Yixun Liang, Rui Chen, Kunming Luo, Zhaohua Zheng, Tongyuan Bai, Feipeng Tian, Zilong Dong, Zihan Zhou, Ping Tan 2026-08-27 arXiv:2608.27073
Public signals Hugging Face upvotes 0
Providers: Hugging Face · Upvotes 0 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-25 14:26:25.142313 UTC

TL;DR - SpatialCrafter generates explorable 3D scenes from a single image by first constructing a globally consistent 3D proxy, then refining it with photorealistic details. This two-stage approach reduces hallucination and long-term geometric drift under large viewpoint changes.

  • Its Point-anchored Sparse Structure Flow module predicts a spatially aligned, geometrically consistent 3D proxy.
  • A video diffusion model acts as a Generative Deferred Refiner, adding high-frequency appearance details while following the proxy geometry.
  • Parallel Geometry Injection and Proxy-Aware Corruption improve integration with pretrained video diffusion models and robustness to proxy artifacts.
  • The authors introduce a 115K-scene hybrid dataset and report improvements over prior methods on synthetic and real-world data.
item →