When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation
TL;DR - This paper identifies mismatched EOS-token preferences between students and teachers as a key cause of excessive response length in on-policy distillation. Treating functionally equivalent EOS tokens as one semantic stopping action substantially reduces length inflation across Qwen3, Llama, and Gemma models.
- Matching decoding stopping sets alone does not resolve the underlying termination mismatch.
- Teacher supervision can suppress the student’s preferred EOS token without successfully transferring the teacher’s alternative.
- Termination preferences can shift substantially across different stages of K2-Horizon training.
- Late-stage length inflation persists after EOS alignment, indicating that termination mismatch is important but not the only cause.