美国具身也没成熟!PI:中国公司何必总当“中国版XX”|RSS 2026
TL;DR - A report from RSS 2026 argues that embodied AI remains immature worldwide and that VLA models and world models are complementary rather than competing paths. Progress depends increasingly on curated data, model-generated training pipelines, and hardware–software co-design.
- World models must predict task-relevant futures that improve action planning; visually plausible video alone cannot replace action supervision.
- Data quality, mixture, and alignment may matter more than raw volume, but no validated scaling law or universal recipe exists yet.
- Existing simulation, planning, and vision models can act as teachers; Stanford’s VLK pipeline generated 48,000 synthetic trajectories and transferred a policy to a Unitree G1.
- China’s strongest current advantage may be robot hardware, while reliable software, data services, and whole-body intelligence remain open challenges globally.