顶会具身辩论赛实况:徐丹飞激辩,PI连输三轮,谷歌大佬倒戈
TL;DR - At an RSS 2026 workshop, robotics experts debated how world models should be structured, controlled, and trained. The emerging view favors modular models, action-grounded evaluation, and combining internet-scale human video with precise robot data.
- Separating world and policy models currently offers better real-time efficiency, debugging, and handling of failure versus success data.
- Action value, physical consistency, and semantic priors were favored over pixel-level realism for contact-rich tasks.
- Human videos provide scale and long-tail coverage, while embodied robot data supplies precise dynamics and control signals.
- A proposed hybrid data flywheel uses world models to bridge broad internet data with high-quality robot experience.