Hyperball May Not Be a Free Lunch
Ranking
Overall
77
Content
80
Popularity
71
Observed public metrics from 1 member.
Merged summary
TL;DR - This paper argues that Hyperball-style optimizers derive much of their behavior from effective step-size evolution rather than inherently better update directions. Their performance still depends critically on learning-rate scheduling.
- Introduces an angular effective learning rate incorporating parameter-update angle, parameter norm, and update norm.
- Finds radial update components have limited direct impact on angular displacement under the tested settings.
- Schedule-matching experiments attribute differences between MuonH and MuonWD mainly to effective step-size dynamics.
- More aggressive decay improves MuonH early but can hurt later pretraining performance.
Sources (1)
Hyperball May Not Be a Free Lunch
Public signals
Semantic Scholar citations 1 · Semantic Scholar influential citations 0
TL;DR - This paper argues that Hyperball-style optimizers derive much of their behavior from effective step-size evolution rather than inherently better update directions. Their performance still depends critically on learning-rate scheduling.
- Introduces an angular effective learning rate incorporating parameter-update angle, parameter norm, and update norm.
- Finds radial update components have limited direct impact on angular displacement under the tested settings.
- Schedule-matching experiments attribute differences between MuonH and MuonWD mainly to effective step-size dynamics.
- More aggressive decay improves MuonH early but can hurt later pretraining performance.