World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal
TL;DR - World Action Agent is a multi-agent framework that lets vision-language models rehearse, revise, and visually correct robot actions before and during execution. It achieves state-of-the-art manipulation results while generating traces that substantially improve smaller VLMs.
- Its visual workspace combines geometry-selected contact views, editable action rehearsals, and closed-loop in-view corrections.
- A Skill Agent retrieves multimodal procedural skills evolved from expert videos and human demonstrations under evidence-based review.
- On LIBERO-Pro, WAA reaches 75.6% average success using skills learned only from LIBERO-90; those skills also transfer to robosuite without further training.
- Fine-tuning Qwen3.5-9B on WAA interaction traces improves out-of-domain success from 1.7% to 43.3%.