REVIEW 7 cited by
A Note on the Convergence of Muon
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this note, we inspect the convergence of a new optimizer for pretraining LLMs, namely the Muon optimizer. Such an optimizer is closely related to a specialized steepest descent method where the update direction is the minimizer of the quadratic approximation of the objective function under spectral norm. We provide the convergence analysis on both versions of the optimizer and discuss its implications.
Forward citations
Cited by 7 Pith papers
-
A Continuous-Time Analysis of Smoothed Matrix-Polar Spectral Gradient Flows for Muon-Type Optimization
A smoothed matrix-polar spectral gradient flow for Muon-type optimization is globally convergent with O(1/T), O(1/t), and exponential rates, and its local advantage over the Frobenius direction is characterized by the...
-
Muse: Representation Geometry of Muon Beyond Normalized Momentum
The matrix shape given to Muon's orthonormalization step is a genuine optimizer axis: square-ish reshapes match native Muon, skinnier reshapes interpolate toward normalized SGD with momentum.
-
Reassessing Muon for Matrix Factorization
Muon's advantage over AdamW is problem-dependent: it loses or ties on plain low-rank factorization and completion but wins on nonnegative matrix factorization.
-
Muon in Associative Memory Learning: Training Dynamics and Scaling Laws
In a linear softmax memory model, Muon equalizes learning across frequency tiers and gives exponential (noiseless) or T^{-2} (noisy power-law) convergence, versus polynomial or T^{-(1-1/β)} for gradient descent.
-
DeMuon: A Decentralized Muon for Matrix Optimization over Graphs
A decentralized Muon optimizer with gradient tracking reaches a stochastic stationary point at the same iteration complexity as centralized heavy-tailed algorithms.
-
AdaGrad Meets Muon: Adaptive Stepsizes for Orthogonal Updates
The paper proposes AdaGO, a Muon variant with a scalar AdaGrad-Norm step size, and proves optimal nonconvex convergence rates while reporting empirical gains on regression and CIFAR-10.
-
Convergence Bound and Critical Batch Size of Muon Optimizer
Muon's four momentum and weight-decay variants are proven to converge with bounded gradient norms under weight decay, and a lower-bound formula links critical batch size to momentum and weight decay.
Discussion (0). Sign in to comment.