Good Pretraining, Bad SFT: Checkpoint Quality Across the Training Stack
Ranking
Observed public metrics from 1 member.
Merged summary
TL;DR - A 30B mixture-of-experts study finds that checkpoints with better pretraining loss or benchmark scores do not necessarily produce the best models after supervised fine-tuning and the full downstream training stack. Robustness to local weight perturbations—termed higher solution density—better characterizes checkpoints that retain strong downstream performance.
- Checkpoint rankings can change across different stages of model training.
- Strong pretraining metrics alone may be insufficient for selecting downstream initialization checkpoints.
- Better final checkpoints maintain downstream performance under local perturbations to their weights.
- Solution density may offer an additional signal for checkpoint selection.
Sources (1)
Good Pretraining, Bad SFT: Checkpoint Quality Across the Training Stack
TL;DR - A 30B mixture-of-experts study finds that checkpoints with better pretraining loss or benchmark scores do not necessarily produce the best models after supervised fine-tuning and the full downstream training stack. Robustness to local weight perturbations—termed higher solution density—better characterizes checkpoints that retain strong downstream performance.
- Checkpoint rankings can change across different stages of model training.
- Strong pretraining metrics alone may be insufficient for selecting downstream initialization checkpoints.
- Better final checkpoints maintain downstream performance under local perturbations to their weights.
- Solution density may offer an additional signal for checkpoint selection.