Hyperball May Not Be a Free Lunch
TL;DR - This paper argues that Hyperball-style optimizers derive much of their behavior from effective step-size evolution rather than inherently better update directions. Their performance still depends critically on learning-rate scheduling.
- Introduces an angular effective learning rate incorporating parameter-update angle, parameter norm, and update norm.
- Finds radial update components have limited direct impact on angular displacement under the tested settings.
- Schedule-matching experiments attribute differences between MuonH and MuonWD mainly to effective step-size dynamics.
- More aggressive decay improves MuonH early but can hurt later pretraining performance.