🛰️ Daily AI Frontier
‹ back to 2026-09-19

What Does Privileged Information Add to On-Policy Self-Distillation?

arXiv cs.CL LLMs & Foundation Models XiuYu Zhang, Wei Chow, Junfeng Fang, Zhenkai Liang, Tat-Seng Chua 2026-09-17

TL;DR - This study isolates how much privileged answers or worked solutions contribute to on-policy self-distillation for mathematical reasoning. Most gains came from distillation itself rather than privileged references, with benefits varying by model and rollout format.

  • AMPLE-Math contains 5,319 problems with six reasoning views sharing the same answer, enabling matched comparisons against reference-free distillation.
  • For Qwen3-1.7B, reference-free distillation explained much of the improvement under thinking-enabled evaluation; polished solutions provided only modest additional benefit.
  • Complete reasoning traces improved SmolLM3-3B by two percentage points at step 50, showing that reference value depends on the student model.
  • Replacing short direct responses with long thinking-enabled rollouts reversed gains into losses, suggesting OPSD primarily improves access to existing reasoning capabilities across inference modes.

view merged work →