🛰️ Daily AI Frontier
‹ back to 2026-09-19

RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

Research LLM Agents

Ranking

Overall 86
Content 95
Popularity 65

Observed public metrics from 1 member.

Representative image for RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

Merged summary

TL;DR - RetireOPD trains agentic language models with a skill-conditioned teacher that automatically “retires” once the student approaches its performance and their discrepancy plateaus. This adaptive combination of on-policy distillation and reinforcement learning substantially outperforms RL alone on ALFWorld and WebShop.

  • The teacher is optimized separately using environment rewards before jointly supervising a skill-free student alongside RL.
  • Adaptive Retirement replaces a fixed distillation schedule with criteria based on teacher-student discrepancy and relative task success.
  • Across Qwen2.5 models from 1.5B to 7B, gains over RL range from 14.1%–18.8% on ALFWorld and 11.8%–19.0% on WebShop.
  • The student ultimately surpasses its skill-conditioned teacher in every reported setting.

Sources (1)

RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

arXiv cs.CL Yan Yu, Zhengxi Lu, Yizhou Liu, Yichen Pan, Aozhe Wang, Qipeng Chen, Hua Yang, Wenqi Zhang, Weiming Lu, Qianglong Chen, Yongliang Shen 2026-09-17 arXiv:2609.20784
Public signals Hugging Face upvotes 56
Providers: Hugging Face · Upvotes 56 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-25 14:18:05.375108 UTC

TL;DR - RetireOPD trains agentic language models with a skill-conditioned teacher that automatically “retires” once the student approaches its performance and their discrepancy plateaus. This adaptive combination of on-policy distillation and reinforcement learning substantially outperforms RL alone on ALFWorld and WebShop.

  • The teacher is optimized separately using environment rewards before jointly supervising a skill-free student alongside RL.
  • Adaptive Retirement replaces a fixed distillation schedule with criteria based on teacher-student discrepancy and relative task success.
  • Across Qwen2.5 models from 1.5B to 7B, gains over RL range from 14.1%–18.8% on ALFWorld and 11.8%–19.0% on WebShop.
  • The student ultimately surpasses its skill-conditioned teacher in every reported setting.
item →