🛰️ Daily AI Frontier
‹ back to 2026-09-02

Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall

Research LLMs & Foundation Models

Ranking

Overall 88
Content 100
Popularity 60

Observed public metrics from 1 member.

Representative image for Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall

Merged summary

TL;DR - Standard forward-KL knowledge distillation during language-model mid-training improves reasoning but can slow factual recall acquisition. The proposed entropy-routed Switch Distillation preserves most factual recall while delivering stronger reasoning and knowledge performance than next-token prediction.

  • Teacher confidence is higher on procedural data than knowledge-intensive data, creating stage-dependent distillation effects as student knowledge evolves.
  • Switch Distillation uses teacher predictive entropy to distill confident tokens and applies cross-entropy elsewhere.
  • It achieves 1.61–1.71× reasoning and 1.13–1.19× knowledge and commonsense performance while retaining 96.7–96.8% of factual recall versus standard next-token prediction.
  • After post-training, it closes the factual-recall gap while retaining 1.25–1.32× reasoning and 1.13–1.20× knowledge and commonsense gains.

Sources (1)

Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall

arXiv cs.CL Jacqueline He, Howard Yen, Shuyue Stella Li, Margaret Li, Hanqing Zeng, Yinglong Xia, Benyu Zhang, Zhuokai Zhao, Qiang Zhang, Pang Wei Koh, Luke Zettlemoyer, Wen-tau Yih 2026-09-01 arXiv:2609.01532
Public signals Hugging Face upvotes 8
Providers: Hugging Face · Upvotes 8 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-25 14:24:55.721704 UTC

TL;DR - Standard forward-KL knowledge distillation during language-model mid-training improves reasoning but can slow factual recall acquisition. The proposed entropy-routed Switch Distillation preserves most factual recall while delivering stronger reasoning and knowledge performance than next-token prediction.

  • Teacher confidence is higher on procedural data than knowledge-intensive data, creating stage-dependent distillation effects as student knowledge evolves.
  • Switch Distillation uses teacher predictive entropy to distill confident tokens and applies cross-entropy elsewhere.
  • It achieves 1.61–1.71× reasoning and 1.13–1.19× knowledge and commonsense performance while retaining 96.7–96.8% of factual recall versus standard next-token prediction.
  • After post-training, it closes the factual-recall gap while retaining 1.25–1.32× reasoning and 1.13–1.20× knowledge and commonsense gains.
item →