🛰️ Daily AI Frontier
‹ back to 2026-09-26

World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal

arXiv cs.RO LLM Agents Yehang Zhang, Haojian Huang, Yifan Chang, Jianchong Su, Bohan Zhou, Yingjie Xu, Wosong Chen, Tianhao Zhou, Chenxu Wang, Tianyi Zhang, Yangkai Wei, Wenqian Li, Shiyuan Deng, Yinchuan Li, Ying-Cong Chen, Zexi Li 2026-09-24
Representative image for World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal

TL;DR - World Action Agent is a multi-agent framework that lets vision-language models rehearse, revise, and visually correct robot actions before and during execution. It achieves state-of-the-art manipulation results while generating traces that substantially improve smaller VLMs.

  • Its visual workspace combines geometry-selected contact views, editable action rehearsals, and closed-loop in-view corrections.
  • A Skill Agent retrieves multimodal procedural skills evolved from expert videos and human demonstrations under evidence-based review.
  • On LIBERO-Pro, WAA reaches 75.6% average success using skills learned only from LIBERO-90; those skills also transfer to robosuite without further training.
  • Fine-tuning Qwen3.5-9B on WAA interaction traces improves out-of-domain success from 1.7% to 43.3%.

view merged work →