现场直击:高昂真机成本锁死具身进化,Google 科学家用视频世界模型破局|RSS 2026
Merged summary
TL;DR - Google/NYU researcher Sherry Yang presented World Gym and World Gymnast, which use controllable video world models to evaluate and improve robot policies without costly physical trials. The approach enables reproducible safety testing and reinforcement learning for failure recovery in the cloud.
- World Gym uses a Diffusion Transformer trained on diverse robot data to predict future video from images and control actions; the reported Bridge model used two A100 GPUs.
- Tests of Octo, OpenVLA, and RT-1 showed world-model evaluations strongly correlated with physical-robot success rates.
- Image editing enables customizable out-of-distribution safety tests involving visual distractions, colors, and shapes.
- World Gymnast uses VLM-generated rewards and policy-gradient updates; physical tests on four held-out tasks reportedly outperformed supervised fine-tuning and the Simpler simulator baseline.
Sources (1)
现场直击:高昂真机成本锁死具身进化,Google 科学家用视频世界模型破局|RSS 2026
TL;DR - Google/NYU researcher Sherry Yang presented World Gym and World Gymnast, which use controllable video world models to evaluate and improve robot policies without costly physical trials. The approach enables reproducible safety testing and reinforcement learning for failure recovery in the cloud.
- World Gym uses a Diffusion Transformer trained on diverse robot data to predict future video from images and control actions; the reported Bridge model used two A100 GPUs.
- Tests of Octo, OpenVLA, and RT-1 showed world-model evaluations strongly correlated with physical-robot success rates.
- Image editing enables customizable out-of-distribution safety tests involving visual distractions, colors, and shapes.
- World Gymnast uses VLM-generated rewards and policy-gradient updates; physical tests on four held-out tasks reportedly outperformed supervised fine-tuning and the Simpler simulator baseline.