🛰️ Daily AI Frontier
‹ back to 2026-09-18

Stress-testing Alignment Midtraining

arXiv cs.CL LLMs & Foundation Models Sid Baines, Jonathan Bostock, Maria Angelica Martinez, Andrew Draganov, David Africa, Daniel Tan 2026-09-17
Representative image for Stress-testing Alignment Midtraining

TL;DR - This study stress-tests alignment midtraining (AMT) at scales up to 110 billion parameters and 1 billion midtraining tokens. It finds AMT can influence model motivations in simple settings, but its effects are fragile and insufficiently supported as a solution to core alignment challenges.

  • AMT continues pretraining on alignment-relevant documents to promote generalization beyond post-training distributions.
  • In ambiguous post-training scenarios, AMT can steer a model toward a desired motivation under simple conditions.
  • A tiny fraction of fine-tuning data suggesting a competing motivation can erase AMT’s influence.
  • Rules are robustly learned only when demonstrated in either midtraining or post-training data.

view merged work →