🛰️ Daily AI Frontier
‹ back to 2026-09-04

Rethinking On-Policy Distillation of Large Language Models II: One Training Example

arXiv cs.AI LLMs & Foundation Models Zixuan Fu, Bingxiang He, Yuxin Zuo, Haohuan Huang, Jinqian Zhang, Ruhang Xiao, Cheng Qian, Qinyu Luo, Huan-ang Gao, Yudong Wang, Zhiyuan Liu, Ning Ding, Chaojun Xiao 2026-09-03
Representative image for Rethinking On-Policy Distillation of Large Language Models II: One Training Example

TL;DR - On-policy distillation (OPD) can recover most full-dataset gains using only one training query, suggesting that student rollouts rapidly expose broad supervision. The main bottleneck is the student’s slow absorption of teacher signals, not insufficient data.

  • A single query reaches 71.5% of the states visited by full-data OPD, mostly within the first 100 training steps.
  • Sixteen semantically diverse queries achieve 98.9% state coverage and match full-data validation performance.
  • Student-teacher alignment slows similarly with one query and the full dataset, with even fixed states requiring hundreds of steps to absorb.
  • The result extends to multi-teacher OPD; content-light templates and off-domain WildChat queries also approach real-query performance.

view merged work →