🛰️ Daily AI Frontier
‹ back to 2026-08-13

Small-Scale Experiments: Are We There Yet?

arXiv cs.LG LLMs & Foundation Models Nicholas Lourie, Kyunghyun Cho, Karen Ullrich, Sanae Lotfi 2026-08-12

TL;DR - Small-model scaling laws become visible when hyperparameters are extensively tuned, suggesting inexpensive experiments can predict some large-scale model behavior. However, statistical limits still constrain extrapolation.

  • Hyperparameter tuning matters more than other tested scaling-law recipe components.
  • Hyperparameter sensitivity decreases with model scale as the loss surface becomes lower-dimensional.
  • Small-scale experiments correctly recover that transformer pre-normalization improves with increasing model size.
  • Reliable extrapolation requires a holistic methodology rather than scaling laws alone.

view merged work →