🛰️ Daily AI Frontier
‹ back to 2026-09-08

What Matters in On-Policy Distillation? A Perspective on Data Efficiency and Data Selection

arXiv cs.AI LLMs & Foundation Models Zhinan Hou, Jiaqi Zhang, Xunliang Cai, Keyou You 2026-09-04

TL;DR - This empirical study finds that on-policy distillation can improve LLM reasoning with remarkably little data when examples elicit long chains of thought. Selecting just eight hard examples matched a 17K-example baseline across models ranging from 1.5B to 7B parameters.

  • One-shot distillation consistently improved performance across all sampled training examples, with harder problems often producing larger gains.
  • Improvement correlated with longer reasoning paths rather than high token entropy.
  • Long chains of thought better preserved teacher alignment over extended reasoning and exposed patterns such as reflection and alternative approaches.
  • Even problems beyond the teacher’s ability proved useful for hard-example selection and student training.

view merged work →