🛰️ Daily AI Frontier
‹ back to 2026-09-08

WorldSculpt: Generating Compositional Worlds from Grounded Videos

arXiv cs.CV Multimodal & Generative Muyao Niu, Jixuan He, Ruihan Yu, Lian Fu, Yonghao Yu, Zheng-Hui Huang, Yifan Zhan, Fengbo Lan, Yongtao Ge, Yinqiang Zheng, Kaipeng Zhang, Zhixiang Wang 2026-09-04
Representative image for WorldSculpt: Generating Compositional Worlds from Grounded Videos

TL;DR - WorldSculpt generates cluttered 3D scenes as collections of individually grounded object meshes in a shared coordinate frame. It scales to hundreds of heavily occluded objects without scene-level training, enabling editable worlds for gaming, AR/VR, simulation, and robotics.

  • Extends the Pixal3D single-object generative prior with multi-view conditioning over posed observations.
  • Fine-tunes only on canonical single objects but generalizes to complex, densely cluttered scenes.
  • Introduces UE-MeshyScene, a photorealistic benchmark with per-object annotations and ground-truth meshes.
  • Outperforms prior methods across single-object and multi-object evaluations, with larger gains under greater complexity and occlusion.

view merged work →