SeeAct AI穆尧:具身智能终局,一定是从“被训练”走向“自我进化”|物理AI50人
TL;DR - SeeAct AI founder Mu Yao argues that embodied intelligence must progress from static training to reward-driven, recursive self-improvement. The key scaling lever is not merely more data or larger models, but generating and converting large volumes of interaction experience into better robotic capabilities.
- “Experience scaling” lets robots learn from autonomous exploration, failures, environmental feedback, and corrective actions rather than relying solely on human demonstrations.
- Recursive self-improvement forms a loop: stronger policies generate richer experiences, which are evaluated through rewards and used to improve low-level control, planning, memory, and skills.
- The proposed data strategy combines limited real-world robot interaction with extensive virtual exploration using physics simulators and neural world models.
- SeeAct AI is designing efficient embodied foundation models—including discrete-diffusion vision-language-action models—to make subsequent reinforcement learning and reward optimization more practical.