🛰️ Daily AI Frontier
‹ back to 2026-08-26

Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining

Research LLMs & Foundation Models

Ranking

Overall 89
Content 100
Popularity 64

Observed public metrics from 1 member.

Merged summary

TL;DR - This paper finds that language-model pretraining loss dynamics are governed mainly by the effective learning rate (ELR), defined as the learning rate relative to parameter norm. Using ELR as a shared coordinate could make scaling laws and training schedules more transferable across optimization and norm-control methods.

  • Matching ELR produces nearly identical loss trajectories despite substantially different raw learning rates and parameter norms.
  • The effect holds across optimizers, architectures, datasets, and model scales, with typical mean collapse errors of a few Ă— 10^-3.
  • Normalization design and the timescale of learning-rate and norm variation determine how precisely trajectories collapse.
  • ELR explains how weight decay and Hyperball shape loss dynamics and enables a fitted scaling law to transfer across norm-control methods, including delayed acceleration.

Sources (1)

Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining

arXiv cs.LG Zihan Liu, Ruiheng Zheng, Shaobo Zhang, Changxin Tian, Kunlong Chen, Zhiqiang Zhang, Lei Wu 2026-08-25 arXiv:2608.24814
Public signals Semantic Scholar citations 3 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 3 · Influential citations 0 X · N/A Fetched 2026-09-25 14:28:42.578514 UTC

TL;DR - This paper finds that language-model pretraining loss dynamics are governed mainly by the effective learning rate (ELR), defined as the learning rate relative to parameter norm. Using ELR as a shared coordinate could make scaling laws and training schedules more transferable across optimization and norm-control methods.

  • Matching ELR produces nearly identical loss trajectories despite substantially different raw learning rates and parameter norms.
  • The effect holds across optimizers, architectures, datasets, and model scales, with typical mean collapse errors of a few Ă— 10^-3.
  • Normalization design and the timescale of learning-rate and norm variation determine how precisely trajectories collapse.
  • ELR explains how weight decay and Hyperball shape loss dynamics and enables a fitted scaling law to transfer across norm-control methods, including delayed acceleration.
item →