🛰️ Daily AI Frontier
‹ back to 2026-09-04

Rethinking On-Policy Distillation of Large Language Models II: One Training Example

Research LLMs & Foundation Models

Ranking

Overall 91
Content 100
Popularity 70

Observed public metrics from 1 member.

Representative image for Rethinking On-Policy Distillation of Large Language Models II: One Training Example

Merged summary

TL;DR - On-policy distillation (OPD) can recover most full-dataset gains using only one training query, suggesting that student rollouts rapidly expose broad supervision. The main bottleneck is the student’s slow absorption of teacher signals, not insufficient data.

  • A single query reaches 71.5% of the states visited by full-data OPD, mostly within the first 100 training steps.
  • Sixteen semantically diverse queries achieve 98.9% state coverage and match full-data validation performance.
  • Student-teacher alignment slows similarly with one query and the full dataset, with even fixed states requiring hundreds of steps to absorb.
  • The result extends to multi-teacher OPD; content-light templates and off-domain WildChat queries also approach real-query performance.

Sources (1)

Rethinking On-Policy Distillation of Large Language Models II: One Training Example

arXiv cs.AI Zixuan Fu, Bingxiang He, Yuxin Zuo, Haohuan Huang, Jinqian Zhang, Ruhang Xiao, Cheng Qian, Qinyu Luo, Huan-ang Gao, Yudong Wang, Zhiyuan Liu, Ning Ding, Chaojun Xiao 2026-09-03 arXiv:2609.04172
Public signals Hugging Face upvotes 101
Providers: Hugging Face · Upvotes 101 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-25 14:24:05.574649 UTC

TL;DR - On-policy distillation (OPD) can recover most full-dataset gains using only one training query, suggesting that student rollouts rapidly expose broad supervision. The main bottleneck is the student’s slow absorption of teacher signals, not insufficient data.

  • A single query reaches 71.5% of the states visited by full-data OPD, mostly within the first 100 training steps.
  • Sixteen semantically diverse queries achieve 98.9% state coverage and match full-data validation performance.
  • Student-teacher alignment slows similarly with one query and the full dataset, with even fixed states requiring hundreds of steps to absorb.
  • The result extends to multi-teacher OPD; content-light templates and off-domain WildChat queries also approach real-query performance.
item →