🛰️ Daily AI Frontier
‹ back to 2026-09-09

Good Pretraining, Bad SFT: Checkpoint Quality Across the Training Stack

arXiv cs.AI LLMs & Foundation Models Sohir Maskey, Philipp Scholl, Jonas Knupp, Pit Neitemeier, Sascha Wirges 2026-09-08

TL;DR - A 30B mixture-of-experts study finds that checkpoints with better pretraining loss or benchmark scores do not necessarily produce the best models after supervised fine-tuning and the full downstream training stack. Robustness to local weight perturbations—termed higher solution density—better characterizes checkpoints that retain strong downstream performance.

  • Checkpoint rankings can change across different stages of model training.
  • Strong pretraining metrics alone may be insufficient for selecting downstream initialization checkpoints.
  • Better final checkpoints maintain downstream performance under local perturbations to their weights.
  • Solution density may offer an additional signal for checkpoint selection.

view merged work →