🛰️ Daily AI Frontier
‹ back to 2026-08-13

Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling

arXiv cs.LG LLMs & Foundation Models Xinmu Ge, Zizhuo Zhang, Yu Huang, Jianing Zhu, Lin Yuan, Wanli Gu, Weichang Wu, Weiran Huang, Xiaolu Zhang, Bo Han, Jun Zhou, Jiangchao Yao 2026-08-12
Representative image for Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling

TL;DR - This study finds that on-policy distillation improves LLM reasoning mainly by making successful outputs more likely with small sampling budgets, rather than expanding the model’s ultimate reasoning capabilities.

  • OPD models sustain better avg@K across sampling budgets.
  • As K increases, pre-OPD models gradually overtake OPD models on pass@K.
  • Training shifts performance toward stronger small-K results at the expense of the large-K capability boundary.
  • At pass@1024, OPD makes more previously solvable problems unsolvable than it makes previously unsolvable problems solvable.

view merged work →