What Does Privileged Information Add to On-Policy Self-Distillation?
TL;DR - This study isolates how much privileged answers or worked solutions contribute to on-policy self-distillation for mathematical reasoning. Most gains came from distillation itself rather than privileged references, with benefits varying by model and rollout format.
- AMPLE-Math contains 5,319 problems with six reasoning views sharing the same answer, enabling matched comparisons against reference-free distillation.
- For Qwen3-1.7B, reference-free distillation explained much of the improvement under thinking-enabled evaluation; polished solutions provided only modest additional benefit.
- Complete reasoning traces improved SmolLM3-3B by two percentage points at step 50, showing that reference value depends on the student model.
- Replacing short direct responses with long thinking-enabled rollouts reversed gains into losses, suggesting OPSD primarily improves access to existing reasoning capabilities across inference modes.