🛰️ Daily AI Frontier
‹ back to 2026-08-23

Beyond Imitation: Self-Improving Robot Policies via Off-Policy Q-Planning

Research Robot Learning

Ranking

Overall 87
Content 95
Popularity 68

Observed public metrics from 1 member.

Representative image for Beyond Imitation: Self-Improving Robot Policies via Off-Policy Q-Planning

Merged summary

TL;DR - Q-Planning enables large behavior-cloned robot policies to improve from successful and failed deployment rollouts by training a small off-policy Q-function while keeping the original policy frozen. It delivers stable gains in simulation and on contact-rich real-robot tasks without further human demonstrations.

  • The Q-function guides inference using a single-step, Q-weighted average over actions sampled from the behavior-cloned policy.
  • After ten self-improvement iterations, success rose from 93% to 99% on LIBERO-10 and from 83.8% to 91.4% on bimanual RoboTwin.
  • In five real-robot iterations without human intervention, stack-cups improved from 40% to 90% and insert-wallet from 25% to 80%.
  • Under the same online data budget, Q-Planning was the only tested method to improve stably from failures without training an auxiliary actor.

Sources (1)

Beyond Imitation: Self-Improving Robot Policies via Off-Policy Q-Planning

arXiv cs.RO Varun Giridhar, Anant Khandelwal, Jeremy A. Collins, Ignat Georgiev, Animesh Garg 2026-08-21 arXiv:2608.21204
Public signals Hugging Face upvotes 2
Providers: Hugging Face · Upvotes 2 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-22 14:33:06.028658 UTC

TL;DR - Q-Planning enables large behavior-cloned robot policies to improve from successful and failed deployment rollouts by training a small off-policy Q-function while keeping the original policy frozen. It delivers stable gains in simulation and on contact-rich real-robot tasks without further human demonstrations.

  • The Q-function guides inference using a single-step, Q-weighted average over actions sampled from the behavior-cloned policy.
  • After ten self-improvement iterations, success rose from 93% to 99% on LIBERO-10 and from 83.8% to 91.4% on bimanual RoboTwin.
  • In five real-robot iterations without human intervention, stack-cups improved from 40% to 90% and insert-wallet from 25% to 80%.
  • Under the same online data budget, Q-Planning was the only tested method to improve stably from failures without training an auxiliary actor.
item →