RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
Ranking
Overall
86
Content
95
Popularity
65
Observed public metrics from 1 member.
Merged summary
TL;DR - RetireOPD trains agentic language models with a skill-conditioned teacher that automatically “retires” once the student approaches its performance and their discrepancy plateaus. This adaptive combination of on-policy distillation and reinforcement learning substantially outperforms RL alone on ALFWorld and WebShop.
- The teacher is optimized separately using environment rewards before jointly supervising a skill-free student alongside RL.
- Adaptive Retirement replaces a fixed distillation schedule with criteria based on teacher-student discrepancy and relative task success.
- Across Qwen2.5 models from 1.5B to 7B, gains over RL range from 14.1%–18.8% on ALFWorld and 11.8%–19.0% on WebShop.
- The student ultimately surpasses its skill-conditioned teacher in every reported setting.
Sources (1)
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
Public signals
Hugging Face upvotes 56
TL;DR - RetireOPD trains agentic language models with a skill-conditioned teacher that automatically “retires” once the student approaches its performance and their discrepancy plateaus. This adaptive combination of on-policy distillation and reinforcement learning substantially outperforms RL alone on ALFWorld and WebShop.
- The teacher is optimized separately using environment rewards before jointly supervising a skill-free student alongside RL.
- Adaptive Retirement replaces a fixed distillation schedule with criteria based on teacher-student discrepancy and relative task success.
- Across Qwen2.5 models from 1.5B to 7B, gains over RL range from 14.1%–18.8% on ALFWorld and 11.8%–19.0% on WebShop.
- The student ultimately surpasses its skill-conditioned teacher in every reported setting.