🛰️ Daily AI Frontier
‹ back to 2026-08-27

Spectral Allocation: Why Muon Outperforms Adam, and How to Improve Muon

Research Efficiency & Systems

Ranking

Overall 88
Content 95
Popularity 71

Observed public metrics from 1 member.

Representative image for Spectral Allocation: Why Muon Outperforms Adam, and How to Improve Muon

Merged summary

TL;DR - Spectral analysis suggests Muon accelerates LLM pretraining by allocating updates more effectively than Adam across singular directions. The proposed Spectral-Aware Muon further exploits tolerant directions, reducing the training tokens needed to reach a target validation loss.

  • Transformer loss landscapes exhibit a stable anisotropic profile: a volatile spectral head requires small steps, while the tolerant bulk supports much larger ones.
  • Muon’s uniform scaling explains its advantage over Adam and SGD but still underuses the bulk directions.
  • SAMuon amplifies the bulk using a static spectral prior; SAMuon-lite approximates this with rank-one power iteration and near-zero wall-clock overhead.
  • Across 124M–1B parameter models, SAMuon used 13.3%–24.0% fewer tokens than Muon to reach the same validation loss.

Sources (1)

Spectral Allocation: Why Muon Outperforms Adam, and How to Improve Muon

arXiv cs.LG Xiaodong Wu, Wenyi Yu, Chao Zhang, Philip Woodland 2026-08-26 arXiv:2608.25990
Public signals Semantic Scholar citations 1 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 1 · Influential citations 0 X · N/A Fetched 2026-09-19 14:22:28.549645 UTC

TL;DR - Spectral analysis suggests Muon accelerates LLM pretraining by allocating updates more effectively than Adam across singular directions. The proposed Spectral-Aware Muon further exploits tolerant directions, reducing the training tokens needed to reach a target validation loss.

  • Transformer loss landscapes exhibit a stable anisotropic profile: a volatile spectral head requires small steps, while the tolerant bulk supports much larger ones.
  • Muon’s uniform scaling explains its advantage over Adam and SGD but still underuses the bulk directions.
  • SAMuon amplifies the bulk using a static spectral prior; SAMuon-lite approximates this with rank-one power iteration and near-zero wall-clock overhead.
  • Across 124M–1B parameter models, SAMuon used 13.3%–24.0% fewer tokens than Muon to reach the same validation loss.
item →