WRC 2026|生数科技发布最新研究成果,提出通用世界模型五级发展路线
TL;DR - ShengShu Technology introduced a five-level roadmap for general world models that progresses from world generation to multi-agent orchestration. The framework matters because it treats understanding, prediction, and action as a unified feedback loop connecting generative models with embodied AI.
- The roadmap spans world generation (L1), interactive worlds (L2), physical action (L3), autonomous world agents (L4), and multi-agent/resource orchestration (L5).
- Its Motubrain model uses a Mixture-of-Transformers architecture to jointly process images, video, language, and robot actions for integrated perception, prediction, and control.
- ShengShu reports that Motubrain achieved 96.1 on RoboTwin 2.0, runs about 10× faster than Motus, and can adapt to a new robot embodiment using 50–100 human demonstrations.
- Reaching L4–L5 will require advances in joint evaluation, physical reasoning, persistent memory, online learning, efficient real-time deployment, and safety.