🛰️ Daily AI Frontier
‹ back to 2026-09-09

Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation

arXiv cs.LG LLMs & Foundation Models Youngrok Park, Sangmin Bae, Hojung Jung, Jongwoo Ko, Yunseon Choi, Young Jin Kim, Pashmina Cameron, Aaron Courville, Se-Young Yun 2026-09-08
Representative image for Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation

TL;DR - On-Policy Reverse Distillation (OPRD) helps stronger models learn from weaker teachers without inheriting their performance ceiling. It accelerates verifier-guided optimization by amplifying teacher-aligned, verifier-supported updates while preserving the underlying policy objective’s stationary points.

  • OPRD measures the teacher’s policy shift relative to its reference policy on student-generated rollouts.
  • It rescales only verifier-supported student updates aligned with that shift, allowing the student to improve beyond the teacher.
  • In successive transfer and multi-teacher distillation, OPRD reportedly achieves higher performance with fewer updates than existing reinforcement-learning and distillation methods.
  • OPRD students remain stylistically closer to verifier-RL models than to their teachers, suggesting that teacher guidance accelerates rather than redirects learning.

view merged work →