🛰️ Daily AI Frontier
‹ back to 2026-09-19

RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

arXiv cs.CL LLM Agents Yan Yu, Zhengxi Lu, Yizhou Liu, Yichen Pan, Aozhe Wang, Qipeng Chen, Hua Yang, Wenqi Zhang, Weiming Lu, Qianglong Chen, Yongliang Shen 2026-09-17
Representative image for RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

TL;DR - RetireOPD trains agentic language models with a skill-conditioned teacher that automatically “retires” once the student approaches its performance and their discrepancy plateaus. This adaptive combination of on-policy distillation and reinforcement learning substantially outperforms RL alone on ALFWorld and WebShop.

  • The teacher is optimized separately using environment rewards before jointly supervising a skill-free student alongside RL.
  • Adaptive Retirement replaces a fixed distillation schedule with criteria based on teacher-student discrepancy and relative task success.
  • Across Qwen2.5 models from 1.5B to 7B, gains over RL range from 14.1%–18.8% on ALFWorld and 11.8%–19.0% on WebShop.
  • The student ultimately surpasses its skill-conditioned teacher in every reported setting.

view merged work →