🛰️ Daily AI Frontier
‹ back to 2026-09-14

探索RSI,生数新世界模型让机器人开始自我进化

Industry & News Embodied AI

Ranking

Overall 66
Content 75
Popularity 45

Observed public metrics from 1 member.

Representative image for 探索RSI,生数新世界模型让机器人开始自我进化

Merged summary

TL;DR - ShengShu Technology unveiled Motus2, a multimodal world-action model that lets robots generate actions, predict outcomes, evaluate results, and improve their policies through a closed feedback loop. It marks an early, bounded exploration of recursive self-improvement for robotic manipulation rather than open-ended autonomous learning.

  • Motus2 combines action generation, an action-conditioned world model, and a value model; Best-of-N planning simulates and scores candidate actions before execution.
  • Planning plus model-based reinforcement learning raised average success on two real-robot tasks from 65% to 75%.
  • Adding tactile feedback improved paper-tearing and cup-extraction success from 60% to 72.5%, while observation memory supported tasks requiring historical context.
  • Training uses roughly 130,000 hours of human egocentric video plus robot-alignment data; robot-domain intermediate training raised five-task average success from 51% to 84%.

Sources (1)

探索RSI,生数新世界模型让机器人开始自我进化

量子位 林, 方舟 2026-09-14 arXiv:2608.30237
Public signals Hugging Face upvotes 5
Providers: Hugging Face · Upvotes 5 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-25 14:19:54.220216 UTC

TL;DR - ShengShu Technology unveiled Motus2, a multimodal world-action model that lets robots generate actions, predict outcomes, evaluate results, and improve their policies through a closed feedback loop. It marks an early, bounded exploration of recursive self-improvement for robotic manipulation rather than open-ended autonomous learning.

  • Motus2 combines action generation, an action-conditioned world model, and a value model; Best-of-N planning simulates and scores candidate actions before execution.
  • Planning plus model-based reinforcement learning raised average success on two real-robot tasks from 65% to 75%.
  • Adding tactile feedback improved paper-tearing and cup-extraction success from 60% to 72.5%, while observation memory supported tasks requiring historical context.
  • Training uses roughly 130,000 hours of human egocentric video plus robot-alignment data; robot-domain intermediate training raised five-task average success from 51% to 84%.
item →