Pith. sign in

REVIEW 7 cited by

A Note on the Convergence of Muon

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.02900 v2 pith:7COUNJA2 submitted 2025-02-05 math.OC

classification math.OC
keywords optimizerconvergencemuonnoteanalysisapproximationcloselydescent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this note, we inspect the convergence of a new optimizer for pretraining LLMs, namely the Muon optimizer. Such an optimizer is closely related to a specialized steepest descent method where the update direction is the minimizer of the quadratic approximation of the objective function under spectral norm. We provide the convergence analysis on both versions of the optimizer and discuss its implications.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Continuous-Time Analysis of Smoothed Matrix-Polar Spectral Gradient Flows for Muon-Type Optimization

    math.OC 2026-08 accept novelty 6.0 of 10

    A smoothed matrix-polar spectral gradient flow for Muon-type optimization is globally convergent with O(1/T), O(1/t), and exponential rates, and its local advantage over the Frobenius direction is characterized by the...

  2. Muse: Representation Geometry of Muon Beyond Normalized Momentum

    cs.LG 2026-07 conditional novelty 6.0 of 10

    The matrix shape given to Muon's orthonormalization step is a genuine optimizer axis: square-ish reshapes match native Muon, skinnier reshapes interpolate toward normalized SGD with momentum.

  3. Reassessing Muon for Matrix Factorization

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Muon's advantage over AdamW is problem-dependent: it loses or ties on plain low-rank factorization and completion but wins on nonnegative matrix factorization.

  4. Muon in Associative Memory Learning: Training Dynamics and Scaling Laws

    cs.LG 2026-02 conditional novelty 6.0 of 10

    In a linear softmax memory model, Muon equalizes learning across frequency tiers and gives exponential (noiseless) or T^{-2} (noisy power-law) convergence, versus polynomial or T^{-(1-1/β)} for gradient descent.

  5. DeMuon: A Decentralized Muon for Matrix Optimization over Graphs

    math.OC 2025-10 conditional novelty 6.0 of 10

    A decentralized Muon optimizer with gradient tracking reaches a stochastic stationary point at the same iteration complexity as centralized heavy-tailed algorithms.

  6. AdaGrad Meets Muon: Adaptive Stepsizes for Orthogonal Updates

    cs.LG 2025-09 conditional novelty 5.0 of 10

    The paper proposes AdaGO, a Muon variant with a scalar AdaGrad-Norm step size, and proves optimal nonconvex convergence rates while reporting empirical gains on regression and CIFAR-10.

  7. Convergence Bound and Critical Batch Size of Muon Optimizer

    cs.LG 2025-07 conditional novelty 5.0 of 10

    Muon's four momentum and weight-decay variants are proven to converge with bounded gradient norms under weight decay, and a lower-bound formula links critical batch size to momentum and weight decay.

Pith tools