🛰️ Daily AI Frontier
‹ back to 2026-08-24

Scaling Muon for Diffusion Transformers

Research Efficiency & Systems

Ranking

Overall 78
Content 95
Popularity 39

Observed public metrics from 1 member.

Representative image for Scaling Muon for Diffusion Transformers

Merged summary

TL;DR - Periodic Row-wise Muon makes the Muon optimizer more practical for training Diffusion Transformers at 1.3B–15B parameters. It preserves Muon’s generative-quality gains while substantially reducing optimizer computation, communication, and total training time.

  • Muon improved best observed generative quality by 12.9–19.1% over AdamW across tested model scales.
  • The proposed method runs the full five-step Newton–Schulz spectral update only every (K) steps, using cheaper row-wise constrained updates between refreshes.
  • A sharding-aware distributed implementation reduces optimizer time by 46.9–54.3%, step time by 15.7–24.3%, and logical communication volume by 66.7%.
  • It reached its best generative quality with 33.7–64.8% less active training time than vanilla Muon, while staying within 0.5% quality on 1.3B–4B models and improving quality by 4.5% at 9B.

Sources (1)

Scaling Muon for Diffusion Transformers

arXiv cs.LG Chenghao Li, Xiao Han, Xinxin Huang, Wei Liu, Boyang Li, Bing Xiao, Heran Zhang, Juanma Perez Rua, Ke Xu, Kangning Liu, Linjun Kuang, Na Li, Tan Wang, Tian Xie, Wei Peng, Yang Pei, Yifan Xu, Yuanhao Zhai, Yuwei Lin, Zhe Wang, Zihao He, Daniel Li, Junbiao Tang, Ziyang Jiang, Dake Chen 2026-08-21 arXiv:2608.20818
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-28 14:15:39.556963 UTC

TL;DR - Periodic Row-wise Muon makes the Muon optimizer more practical for training Diffusion Transformers at 1.3B–15B parameters. It preserves Muon’s generative-quality gains while substantially reducing optimizer computation, communication, and total training time.

  • Muon improved best observed generative quality by 12.9–19.1% over AdamW across tested model scales.
  • The proposed method runs the full five-step Newton–Schulz spectral update only every (K) steps, using cheaper row-wise constrained updates between refreshes.
  • A sharding-aware distributed implementation reduces optimizer time by 46.9–54.3%, step time by 15.7–24.3%, and logical communication volume by 66.7%.
  • It reached its best generative quality with 33.7–64.8% less active training time than vanilla Muon, while staying within 0.5% quality on 1.3B–4B models and improving quality by 4.5% at 9B.
item →