🛰️ Daily AI Frontier
‹ back to 2026-09-17

In-Context Robot Learning with VLM Agents

Research LLM Agents

Ranking

Overall 85
Content 95
Popularity 62

Observed public metrics from 1 member.

Representative image for In-Context Robot Learning with VLM Agents

Merged summary

TL;DR - GPT-Policy is a framework that uses vision-language model agents to learn robot tasks from deployment-time context without gradient updates. Real-robot experiments suggest that human videos improve task completion, while aligned action references provide additional gains for contact-sensitive tasks.

  • A context compiler preserves task-relevant visual transitions from demonstrations, examples, and interaction feedback.
  • A VLM proposes robot-tool actions, while a constrained controller verifies and executes them and reports outcomes.
  • The study evaluates task success, efficiency, model differences, and the effects of removing context components.
  • Human demonstrations can help even without robot action labels, indicating a path toward more adaptable general-purpose robotic policies.

Sources (1)

In-Context Robot Learning with VLM Agents

arXiv cs.CV Dongzhou Cheng, Taoran Yi, Ye Fang, Xingwu Zhang, Fan Feng, Yixuan Li, Gengxiong Zhuang, Rongze Wang, Shuai Yang, Wei Song, Weizhi Xue, Minyan Wu, Jie Gui, Jiaqi Wang, Tong Wu 2026-09-16 arXiv:2609.19138
Public signals Hugging Face upvotes 22
Providers: Hugging Face · Upvotes 22 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-25 14:19:00.205351 UTC

TL;DR - GPT-Policy is a framework that uses vision-language model agents to learn robot tasks from deployment-time context without gradient updates. Real-robot experiments suggest that human videos improve task completion, while aligned action references provide additional gains for contact-sensitive tasks.

  • A context compiler preserves task-relevant visual transitions from demonstrations, examples, and interaction feedback.
  • A VLM proposes robot-tool actions, while a constrained controller verifies and executes them and reports outcomes.
  • The study evaluates task success, efficiency, model differences, and the effects of removing context components.
  • Human demonstrations can help even without robot action labels, indicating a path toward more adaptable general-purpose robotic policies.
item →