Everything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-Training
TL;DR - Controlled experiments on multi-domain mid-training find that moderate coverage of each reasoning domain—roughly 10%–40%—outperforms extreme allocations. Subsequent supervised alignment improves overall accuracy but generally cannot repair domain gaps created during mid-training.
- All five KOR-Bench domains exhibited interior coverage optima; fitted Qwen3-8B peaks ranged from 9.9% to 35.1%.
- A fixed-budget SFT pass improved 116 of 120 evaluated cells by 4.32% on average, yet bridged none of 240 performance gaps at a 5% threshold and only 30 at 10%.
- Omitting a domain caused mid-training accuracy to collapse, although a FineWeb-Edu control indicates that generic distributional drift partly confounds this result.
- An exploratory optimized allocation produced the largest full-pipeline gain (+4.36 percentage points), but its advantage was only marginal under a Welch test.