🛰️ Daily AI Frontier
‹ back to 2026-08-21

Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts

Research Efficiency & Systems

Ranking

Overall 84
Content 90
Popularity 70

Observed public metrics from 1 member.

Representative image for Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts

Merged summary

TL;DR - This paper proposes a two-step method for transferring optimal learning rates from small proxy models to large Mixture-of-Experts models and trillion-token training runs. It could substantially reduce the cost of hyperparameter sweeps for large-scale pretraining.

  • Adapts Maximal Update Parameterization (μP) to MoE architectures using Multi-head Latent Attention and the Muon optimizer.
  • Demonstrates consistent learning-rate transfer across models of different widths.
  • Fits a token-budget scaling law that predicts optimal learning rates up to 10 trillion tokens with (R^2=0.95).
  • Validates the method by stably pretraining a 155B-parameter MoE model with 17B active parameters.

Sources (1)

Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts

arXiv cs.LG Nayeon Kim, Hojin Lee, Yunju Bak, Jaesun Park, Boseop Kim 2026-08-20 arXiv:2608.20061
Public signals Hugging Face upvotes 46
Providers: Hugging Face · Upvotes 46 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-19 14:25:39.722307 UTC

TL;DR - This paper proposes a two-step method for transferring optimal learning rates from small proxy models to large Mixture-of-Experts models and trillion-token training runs. It could substantially reduce the cost of hyperparameter sweeps for large-scale pretraining.

  • Adapts Maximal Update Parameterization (μP) to MoE architectures using Multi-head Latent Attention and the Muon optimizer.
  • Demonstrates consistent learning-rate transfer across models of different widths.
  • Fits a token-budget scaling law that predicts optimal learning rates up to 10 trillion tokens with (R^2=0.95).
  • Validates the method by stably pretraining a 155B-parameter MoE model with 17B active parameters.
item →