Hyperparameter Scaling Laws Across MoE Sparsity
Ranking
Overall
81
Content
100
Popularity
37
Observed public metrics from 1 member.
Merged summary
TL;DR - A large-scale study derives unified hyperparameter scaling laws for Mixture-of-Experts models, showing that optimal learning rate and batch size depend explicitly on expert activation ratio. The laws accurately extrapolate to a held-out 12B-parameter model activating only 1/64 of its experts.
- Based on 1,800 pretraining runs spanning six activated-parameter scales, roughly 20 trillion tokens, and models with up to 6B non-embedding parameters.
- At fixed sparsity, optimal batch size scales with training tokens, while optimal learning rate scales with compute and is robust to how compute is divided between model size and data.
- Across sparsity levels, activation ratio contributes an additional multiplicative power-law factor that total or activated parameter count alone cannot explain.
- The scaling relationships outperform alternative functional forms and transfer across expert granularities.
Sources (1)
Hyperparameter Scaling Laws Across MoE Sparsity
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - A large-scale study derives unified hyperparameter scaling laws for Mixture-of-Experts models, showing that optimal learning rate and batch size depend explicitly on expert activation ratio. The laws accurately extrapolate to a held-out 12B-parameter model activating only 1/64 of its experts.
- Based on 1,800 pretraining runs spanning six activated-parameter scales, roughly 20 trillion tokens, and models with up to 6B non-embedding parameters.
- At fixed sparsity, optimal batch size scales with training tokens, while optimal learning rate scales with compute and is robust to how compute is divided between model size and data.
- Across sparsity levels, activation ratio contributes an additional multiplicative power-law factor that total or activated parameter count alone cannot explain.
- The scaling relationships outperform alternative functional forms and transfer across expert granularities.