Southworth and Stephen Thomas , year =

Ben S · 2026 · arXiv 2603.17970

2 Pith papers cite this work. Polarity classification is still indexing.

2 Pith papers citing it

representative citing papers

Hierarchical Muon: Tiled Newton-Schulz Updates for Efficient Muon Optimization

math.NA · 2026-06-25 · unverdicted · novelty 7.0

HiMuon partitions momentum-gradient matrices into T x T tiles, runs independent Newton-Schulz iterations on each tile, and reassembles the results, reducing leading cost to O(H W T K) while defining a local rather than global matrix map.

Muon in Vision Transformers: Optimizer-Recipe Interactions and Gradient Spectra

cs.LG · 2026-05-23 · conditional · novelty 5.0

Muon optimizer outperforms AdamW in ViT training on two image datasets, with gains that depend on data augmentation strength and are linked to wider singular-value spread in QKV gradients and prevention of late-training mode collapse in MLP blocks.

citing papers explorer

Showing 1 of 1 citing paper after filters.

Muon in Vision Transformers: Optimizer-Recipe Interactions and Gradient Spectra cs.LG · 2026-05-23 · conditional · none · ref 32
Muon optimizer outperforms AdamW in ViT training on two image datasets, with gains that depend on data augmentation strength and are linked to wider singular-value spread in QKV gradients and prevention of late-training mode collapse in MLP blocks.

Southworth and Stephen Thomas , year =

fields

years

verdicts

representative citing papers

citing papers explorer