🛰️ Daily AI Frontier
‹ back to 2026-09-19

When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation

arXiv cs.LG LLMs & Foundation Models Yuxiao Yang, Tianrun Yu, Shangzhe Li, Kaixiang Zhao, Xuchao Zhang, Chetan Bansal, Huaxiu Yao, Taylor W. Killian, Weitong Zhang 2026-09-17
Representative image for When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation

TL;DR - This paper identifies mismatched EOS-token preferences between students and teachers as a key cause of excessive response length in on-policy distillation. Treating functionally equivalent EOS tokens as one semantic stopping action substantially reduces length inflation across Qwen3, Llama, and Gemma models.

  • Matching decoding stopping sets alone does not resolve the underlying termination mismatch.
  • Teacher supervision can suppress the student’s preferred EOS token without successfully transferring the teacher’s alternative.
  • Termination preferences can shift substantially across different stages of K2-Horizon training.
  • Late-stage length inflation persists after EOS alignment, indicating that termination mismatch is important but not the only cause.

view merged work →