When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation
Ranking
Overall
84
Content
90
Popularity
69
Observed public metrics from 1 member.
Merged summary
TL;DR - This paper identifies mismatched EOS-token preferences between students and teachers as a key cause of excessive response length in on-policy distillation. Treating functionally equivalent EOS tokens as one semantic stopping action substantially reduces length inflation across Qwen3, Llama, and Gemma models.
- Matching decoding stopping sets alone does not resolve the underlying termination mismatch.
- Teacher supervision can suppress the student’s preferred EOS token without successfully transferring the teacher’s alternative.
- Termination preferences can shift substantially across different stages of K2-Horizon training.
- Late-stage length inflation persists after EOS alignment, indicating that termination mismatch is important but not the only cause.
Sources (1)
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation
Public signals
Hugging Face upvotes 107
TL;DR - This paper identifies mismatched EOS-token preferences between students and teachers as a key cause of excessive response length in on-policy distillation. Treating functionally equivalent EOS tokens as one semantic stopping action substantially reduces length inflation across Qwen3, Llama, and Gemma models.
- Matching decoding stopping sets alone does not resolve the underlying termination mismatch.
- Teacher supervision can suppress the student’s preferred EOS token without successfully transferring the teacher’s alternative.
- Termination preferences can shift substantially across different stages of K2-Horizon training.
- Late-stage length inflation persists after EOS alignment, indicating that termination mismatch is important but not the only cause.