🛰️ Daily AI Frontier
‹ back to 2026-09-09

Everything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-Training

arXiv cs.AI LLMs & Foundation Models Yunpeng Xu, Kun Zheng 2026-09-08

TL;DR - Controlled experiments on multi-domain mid-training find that moderate coverage of each reasoning domain—roughly 10%–40%—outperforms extreme allocations. Subsequent supervised alignment improves overall accuracy but generally cannot repair domain gaps created during mid-training.

  • All five KOR-Bench domains exhibited interior coverage optima; fitted Qwen3-8B peaks ranged from 9.9% to 35.1%.
  • A fixed-budget SFT pass improved 116 of 120 evaluated cells by 4.32% on average, yet bridged none of 240 performance gaps at a 5% threshold and only 30 at 10%.
  • Omitting a domain caused mid-training accuracy to collapse, although a FineWeb-Edu control indicates that generic distributional drift partly confounds this result.
  • An exploratory optimized allocation produced the largest full-pipeline gain (+4.36 percentage points), but its advantage was only marginal under a Welch test.

view merged work →