Pith. sign in

REVIEW 10 cited by

A New Perspective on Shampoo's Preconditioner

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.17748 v1 pith:B3XVLFOH submitted 2024-06-25 cs.LG math.OCstat.ML

classification cs.LGmath.OCstat.ML
keywords approximationshampookroneckerproducthessianoptimalpreconditioneralgorithm
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Shampoo, a second-order optimization algorithm which uses a Kronecker product preconditioner, has recently garnered increasing attention from the machine learning community. The preconditioner used by Shampoo can be viewed either as an approximation of the Gauss--Newton component of the Hessian or the covariance matrix of the gradients maintained by Adagrad. We provide an explicit and novel connection between the $\textit{optimal}$ Kronecker product approximation of these matrices and the approximation made by Shampoo. Our connection highlights a subtle but common misconception about Shampoo's approximation. In particular, the $\textit{square}$ of the approximation used by the Shampoo optimizer is equivalent to a single step of the power iteration algorithm for computing the aforementioned optimal Kronecker product approximation. Across a variety of datasets and architectures we empirically demonstrate that this is close to the optimal Kronecker product approximation. Additionally, for the Hessian approximation viewpoint, we empirically study the impact of various practical tricks to make Shampoo more computationally efficient (such as using the batch gradient and the empirical Fisher) on the quality of Hessian approximation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BaKron: Efficient Quantization with Kronecker-Factored Hessians

    cs.LG 2026-08 conditional novelty 7.0 of 10

    BaKron computes the same two-sided adaptive rounding as BoA/YAQA but in cubic total work and O(m+n) sequential steps, matching GPTQ's complexity while exploiting richer curvature.

  2. Optimization Geometrodynamics: Variational Reduction and Interaction Curvature

    math.OC 2026-07 conditional novelty 7.0 of 10

    Dynamic geometric complexity of reducing condition number under full SPD metric control equals the affine-invariant distance from the relative log-spectrum to a low-width set.

  3. SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales

    cs.LG 2026-07 conditional novelty 6.0 of 10

    SOAP and Muon, stabilized by per-step QR eigenbasis updates and KL-Shampoo covariance accumulation, beat AdamW on large-batch LLM pretraining up to 100M-token batches.

  4. Causal Optimizer Interaction Calculus: Hidden Geometric Relaxation and Identifiable Interventions

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Under fixed innovation coupling, finite-horizon optimizers admit minimal pathwise realizations and incidence-identifiable Möbius effects, with a five-term readout transfer from hidden relaxation and a closed reduced-v...

  5. Decoupling Variance and Scale-Invariant Updates in Adaptive Gradient Descent for Unified Vector and Matrix Optimization

    cs.LG 2026-02 conditional novelty 6.0 of 10

    DeVA_S8 reweights Muon's matrix-sign update in the matrix's eigenbasis with a singular-value signal-to-noise ratio, reaching target LLM validation perplexity with ~6.6% fewer tokens than Muon.

  6. Harnessing Optimization Dynamics for Curvature-Informed Model Merging

    cs.LG 2025-09 conditional novelty 6.0 of 10

    Optimization Trajectory Aware merging uses Adam second moments as a curvature proxy, first pruning task-vector edits with Fast Fisher Grafting, then reweighting survivors with a compressed curvature preconditioner.

  7. Low-rank Momentum Factorization for Memory Efficient Training

    cs.LG 2025-07 reject novelty 6.0 of 10

    MoFaSGD keeps a low-rank factored momentum and uses its singular vectors as the update direction, achieving LoRA-level memory with competitive fine-tuning performance, but its convergence proof is flawed.

  8. Lions and Muons: Optimization via Stochastic Frank-Wolfe under Heavy-Tailed Noise

    math.OC 2025-06 reject novelty 6.0 of 10

    Lion and Muon with weight decay are shown to be instances of one stochastic Frank-Wolfe algorithm, and clipped and variance-reduced variants get the first high-probability convergence rates for nonconvex Frank-Wolfe u...

  9. On the Convergence Analysis of Muon

    stat.ML 2025-05 unverdicted novelty 6.0 of 10

    Muon's convergence rate depends on an average Hessian curvature along its update directions, which can be much smaller than the worst-case Lipschitz constant when Hessians are low-rank.

  10. Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Optimizers such as Adam, Shampoo, and SOAP are unified as structured Fisher approximations, and two new derived optimizers, RACS and Alice, achieve faster LLaMA pre-training than Adam at lower memory.

Pith tools