SpatialCrafter: Single Image World Modeling with Generative 3D Proxies
TL;DR - SpatialCrafter generates explorable 3D scenes from a single image by first constructing a globally consistent 3D proxy, then refining it with photorealistic details. This two-stage approach reduces hallucination and long-term geometric drift under large viewpoint changes.
- Its Point-anchored Sparse Structure Flow module predicts a spatially aligned, geometrically consistent 3D proxy.
- A video diffusion model acts as a Generative Deferred Refiner, adding high-frequency appearance details while following the proxy geometry.
- Parallel Geometry Injection and Proxy-Aware Corruption improve integration with pretrained video diffusion models and robustness to proxy artifacts.
- The authors introduce a 115K-scene hybrid dataset and report improvements over prior methods on synthetic and real-world data.