Everything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-Training
Ranking
Overall
78
Content
95
Popularity
37
Observed public metrics from 1 member.
Merged summary
TL;DR - Controlled experiments on multi-domain mid-training find that moderate coverage of each reasoning domain—roughly 10%–40%—outperforms extreme allocations. Subsequent supervised alignment improves overall accuracy but generally cannot repair domain gaps created during mid-training.
- All five KOR-Bench domains exhibited interior coverage optima; fitted Qwen3-8B peaks ranged from 9.9% to 35.1%.
- A fixed-budget SFT pass improved 116 of 120 evaluated cells by 4.32% on average, yet bridged none of 240 performance gaps at a 5% threshold and only 30 at 10%.
- Omitting a domain caused mid-training accuracy to collapse, although a FineWeb-Edu control indicates that generic distributional drift partly confounds this result.
- An exploratory optimized allocation produced the largest full-pipeline gain (+4.36 percentage points), but its advantage was only marginal under a Welch test.
Sources (1)
Everything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-Training
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - Controlled experiments on multi-domain mid-training find that moderate coverage of each reasoning domain—roughly 10%–40%—outperforms extreme allocations. Subsequent supervised alignment improves overall accuracy but generally cannot repair domain gaps created during mid-training.
- All five KOR-Bench domains exhibited interior coverage optima; fitted Qwen3-8B peaks ranged from 9.9% to 35.1%.
- A fixed-budget SFT pass improved 116 of 120 evaluated cells by 4.32% on average, yet bridged none of 240 performance gaps at a 5% threshold and only 30 at 10%.
- Omitting a domain caused mid-training accuracy to collapse, although a FineWeb-Edu control indicates that generic distributional drift partly confounds this result.
- An exploratory optimized allocation produced the largest full-pipeline gain (+4.36 percentage points), but its advantage was only marginal under a Welch test.