A Physical Response-and-Memory Model for Muon Optimization
TL;DR - This paper models Muon optimization as a physical medium with memory, interpreting semi-orthogonalized updates as maximally dissipative responses under a safety constraint. It derives a two-timescale Bi-Maxwell optimizer that reaches a target loss in fewer steps on a public LLM training benchmark.
- The model explains momentum as accumulated internal stress and its averaging window as a stress-relaxation timescale.
- Because real media can relax at multiple rates, Bi-Maxwell replaces Muon’s single-timescale memory kernel with fast and slow components.
- Measurements across eight independent training runs support the prediction that optimal memory length should increase as gradient directions change more slowly later in training.
- Changing only the memory kernel to the two-timescale form reduced the steps needed to reach the benchmark’s target loss.