Stress-testing Alignment Midtraining
Ranking
Overall
82
Content
100
Popularity
39
Observed public metrics from 1 member.
Merged summary
TL;DR - This study stress-tests alignment midtraining (AMT) at scales up to 110 billion parameters and 1 billion midtraining tokens. It finds AMT can influence model motivations in simple settings, but its effects are fragile and insufficiently supported as a solution to core alignment challenges.
- AMT continues pretraining on alignment-relevant documents to promote generalization beyond post-training distributions.
- In ambiguous post-training scenarios, AMT can steer a model toward a desired motivation under simple conditions.
- A tiny fraction of fine-tuning data suggesting a competing motivation can erase AMT’s influence.
- Rules are robustly learned only when demonstrated in either midtraining or post-training data.
Sources (1)
Stress-testing Alignment Midtraining
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - This study stress-tests alignment midtraining (AMT) at scales up to 110 billion parameters and 1 billion midtraining tokens. It finds AMT can influence model motivations in simple settings, but its effects are fragile and insufficiently supported as a solution to core alignment challenges.
- AMT continues pretraining on alignment-relevant documents to promote generalization beyond post-training distributions.
- In ambiguous post-training scenarios, AMT can steer a model toward a desired motivation under simple conditions.
- A tiny fraction of fine-tuning data suggesting a competing motivation can erase AMT’s influence.
- Rules are robustly learned only when demonstrated in either midtraining or post-training data.