李飞飞高徒黄文龙:具身智能需要一次「脑内搜索」革命 | RSS 2026
TL;DR - Stanford researcher Wenlong Huang presented Point-World, a 3D world model intended to let robots predict counterfactual physical outcomes and search for novel behaviors before acting. The approach could reduce embodied AI’s dependence on costly demonstrations and support continual autonomous learning.
- Point-World represents observations and robot actions as 3D point-cloud flows, enabling a Transformer model to simulate interactions across different robot embodiments.
- The model reportedly handles rigid bodies, fluids, cloth, occlusion, and articulated objects without explicit object segmentation.
- The team relabeled roughly 3,500 hours of interaction data with depth and camera poses; Huang also described unpublished results where 2D pretraining plus 15 hours of 3D data outperformed an earlier 500-hour setup in dynamics prediction.
- Huang argues that world-model priors combined with counterfactual search could help robots invent behaviors absent from demonstrations and progressively incorporate them through lifelong learning.