🛰️ Daily AI Frontier
‹ back to 2026-09-02

Efficiently Estimating Optimal Hyperparameter Scaling Laws through Power-Law Entropy Search

Research Efficiency & Systems

Ranking

Overall 88
Content 100
Popularity 61

Observed public metrics from 1 member.

Representative image for Efficiently Estimating Optimal Hyperparameter Scaling Laws through Power-Law Entropy Search

Merged summary

TL;DR - Power-Law Entropy Search (PLES) uses cost-aware, multi-fidelity Bayesian optimization to estimate optimal LLM training hyperparameter scaling laws efficiently. It achieves accurate estimates with less than one-tenth the computational budget of grid search and other baselines.

  • PLES selects experiments based on expected reduction in overall scaling-law uncertainty per unit of compute.
  • Its adaptive strategy naturally favors informative, lower-cost small-scale training runs.
  • Evaluations cover synthetic benchmarks, surrogates fitted to real LLM training data, and actual LLM pre-training runs.
  • The method targets scaling-law estimation rather than optimization of a single objective function.

Sources (1)

Efficiently Estimating Optimal Hyperparameter Scaling Laws through Power-Law Entropy Search

arXiv cs.LG Zhiliang Chen, Sebastian Ament, David Eriksson, Maximilian Balandat, Eytan Bakshy, Jihao Andreas Lin 2026-09-01 arXiv:2609.01431
Public signals Semantic Scholar citations 1 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 1 · Influential citations 0 X · N/A Fetched 2026-09-17 14:24:34.480390 UTC

TL;DR - Power-Law Entropy Search (PLES) uses cost-aware, multi-fidelity Bayesian optimization to estimate optimal LLM training hyperparameter scaling laws efficiently. It achieves accurate estimates with less than one-tenth the computational budget of grid search and other baselines.

  • PLES selects experiments based on expected reduction in overall scaling-law uncertainty per unit of compute.
  • Its adaptive strategy naturally favors informative, lower-cost small-scale training runs.
  • Evaluations cover synthetic benchmarks, surrogates fitted to real LLM training data, and actual LLM pre-training runs.
  • The method targets scaling-law estimation rather than optimization of a single objective function.
item →