🛰️ Daily AI Frontier
‹ back to 2026-09-08

WorldSculpt: Generating Compositional Worlds from Grounded Videos

Research Multimodal & Generative

Ranking

Overall 86
Content 95
Popularity 64

Observed public metrics from 1 member.

Representative image for WorldSculpt: Generating Compositional Worlds from Grounded Videos

Merged summary

TL;DR - WorldSculpt generates cluttered 3D scenes as collections of individually grounded object meshes in a shared coordinate frame. It scales to hundreds of heavily occluded objects without scene-level training, enabling editable worlds for gaming, AR/VR, simulation, and robotics.

  • Extends the Pixal3D single-object generative prior with multi-view conditioning over posed observations.
  • Fine-tunes only on canonical single objects but generalizes to complex, densely cluttered scenes.
  • Introduces UE-MeshyScene, a photorealistic benchmark with per-object annotations and ground-truth meshes.
  • Outperforms prior methods across single-object and multi-object evaluations, with larger gains under greater complexity and occlusion.

Sources (1)

WorldSculpt: Generating Compositional Worlds from Grounded Videos

arXiv cs.CV Muyao Niu, Jixuan He, Ruihan Yu, Lian Fu, Yonghao Yu, Zheng-Hui Huang, Yifan Zhan, Fengbo Lan, Yongtao Ge, Yinqiang Zheng, Kaipeng Zhang, Zhixiang Wang 2026-09-04 arXiv:2609.05416
Public signals Hugging Face upvotes 25
Providers: Hugging Face · Upvotes 25 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-25 14:23:04.082304 UTC

TL;DR - WorldSculpt generates cluttered 3D scenes as collections of individually grounded object meshes in a shared coordinate frame. It scales to hundreds of heavily occluded objects without scene-level training, enabling editable worlds for gaming, AR/VR, simulation, and robotics.

  • Extends the Pixal3D single-object generative prior with multi-view conditioning over posed observations.
  • Fine-tunes only on canonical single objects but generalizes to complex, densely cluttered scenes.
  • Introduces UE-MeshyScene, a photorealistic benchmark with per-object annotations and ground-truth meshes.
  • Outperforms prior methods across single-object and multi-object evaluations, with larger gains under greater complexity and occlusion.
item →