World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal
Ranking
Overall
78
Content
95
Popularity
38
Observed public metrics from 1 member.
Merged summary
TL;DR - World Action Agent is a multi-agent framework that lets vision-language models rehearse, revise, and visually correct robot actions before and during execution. It achieves state-of-the-art manipulation results while generating traces that substantially improve smaller VLMs.
- Its visual workspace combines geometry-selected contact views, editable action rehearsals, and closed-loop in-view corrections.
- A Skill Agent retrieves multimodal procedural skills evolved from expert videos and human demonstrations under evidence-based review.
- On LIBERO-Pro, WAA reaches 75.6% average success using skills learned only from LIBERO-90; those skills also transfer to robosuite without further training.
- Fine-tuning Qwen3.5-9B on WAA interaction traces improves out-of-domain success from 1.7% to 43.3%.
Sources (1)
World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal
Public signals
Hugging Face upvotes 4 · Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - World Action Agent is a multi-agent framework that lets vision-language models rehearse, revise, and visually correct robot actions before and during execution. It achieves state-of-the-art manipulation results while generating traces that substantially improve smaller VLMs.
- Its visual workspace combines geometry-selected contact views, editable action rehearsals, and closed-loop in-view corrections.
- A Skill Agent retrieves multimodal procedural skills evolved from expert videos and human demonstrations under evidence-based review.
- On LIBERO-Pro, WAA reaches 75.6% average success using skills learned only from LIBERO-90; those skills also transfer to robosuite without further training.
- Fine-tuning Qwen3.5-9B on WAA interaction traces improves out-of-domain success from 1.7% to 43.3%.