🛰️ Daily AI Frontier
‹ back to 2026-09-08

What Matters in On-Policy Distillation? A Perspective on Data Efficiency and Data Selection

Research LLMs & Foundation Models

Ranking

Overall 78
Content 95
Popularity 37

Observed public metrics from 1 member.

Merged summary

TL;DR - This empirical study finds that on-policy distillation can improve LLM reasoning with remarkably little data when examples elicit long chains of thought. Selecting just eight hard examples matched a 17K-example baseline across models ranging from 1.5B to 7B parameters.

  • One-shot distillation consistently improved performance across all sampled training examples, with harder problems often producing larger gains.
  • Improvement correlated with longer reasoning paths rather than high token entropy.
  • Long chains of thought better preserved teacher alignment over extended reasoning and exposed patterns such as reflection and alternative approaches.
  • Even problems beyond the teacher’s ability proved useful for hard-example selection and student training.

Sources (1)

What Matters in On-Policy Distillation? A Perspective on Data Efficiency and Data Selection

arXiv cs.AI Zhinan Hou, Jiaqi Zhang, Xunliang Cai, Keyou You 2026-09-04 arXiv:2609.05198
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-08 14:03:41.634398 UTC

TL;DR - This empirical study finds that on-policy distillation can improve LLM reasoning with remarkably little data when examples elicit long chains of thought. Selecting just eight hard examples matched a 17K-example baseline across models ranging from 1.5B to 7B parameters.

  • One-shot distillation consistently improved performance across all sampled training examples, with harder problems often producing larger gains.
  • Improvement correlated with longer reasoning paths rather than high token entropy.
  • Long chains of thought better preserved teacher alignment over extended reasoning and exposed patterns such as reflection and alternative approaches.
  • Even problems beyond the teacher’s ability proved useful for hard-example selection and student training.
item →