🛰️ Daily AI Frontier
‹ back to 2026-09-26

World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal

Research LLM Agents

Ranking

Overall 78
Content 95
Popularity 38

Observed public metrics from 1 member.

Representative image for World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal

Merged summary

TL;DR - World Action Agent is a multi-agent framework that lets vision-language models rehearse, revise, and visually correct robot actions before and during execution. It achieves state-of-the-art manipulation results while generating traces that substantially improve smaller VLMs.

  • Its visual workspace combines geometry-selected contact views, editable action rehearsals, and closed-loop in-view corrections.
  • A Skill Agent retrieves multimodal procedural skills evolved from expert videos and human demonstrations under evidence-based review.
  • On LIBERO-Pro, WAA reaches 75.6% average success using skills learned only from LIBERO-90; those skills also transfer to robosuite without further training.
  • Fine-tuning Qwen3.5-9B on WAA interaction traces improves out-of-domain success from 1.7% to 43.3%.

Sources (1)

World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal

arXiv cs.RO Yehang Zhang, Haojian Huang, Yifan Chang, Jianchong Su, Bohan Zhou, Yingjie Xu, Wosong Chen, Tianhao Zhou, Chenxu Wang, Tianyi Zhang, Yangkai Wei, Wenqian Li, Shiyuan Deng, Yinchuan Li, Ying-Cong Chen, Zexi Li 2026-09-24 arXiv:2609.29964
Public signals Hugging Face upvotes 4 · Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · Upvotes 4 OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-26 14:03:40.795637 UTC

TL;DR - World Action Agent is a multi-agent framework that lets vision-language models rehearse, revise, and visually correct robot actions before and during execution. It achieves state-of-the-art manipulation results while generating traces that substantially improve smaller VLMs.

  • Its visual workspace combines geometry-selected contact views, editable action rehearsals, and closed-loop in-view corrections.
  • A Skill Agent retrieves multimodal procedural skills evolved from expert videos and human demonstrations under evidence-based review.
  • On LIBERO-Pro, WAA reaches 75.6% average success using skills learned only from LIBERO-90; those skills also transfer to robosuite without further training.
  • Fine-tuning Qwen3.5-9B on WAA interaction traces improves out-of-domain success from 1.7% to 43.3%.
item →