Pith. sign in

REVIEW 5 cited by

On the Convergence of Adaptive Gradient Methods for Nonconvex Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1808.05671 v4 pith:3W73LM64 submitted 2018-08-16 cs.LG math.OCstat.ML

classification cs.LGmath.OCstat.ML
keywords adaptivegradientmethodsconvergencenonconvexadagradamsgradbeen
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Adaptive gradient methods are workhorses in deep learning. However, the convergence guarantees of adaptive gradient methods for nonconvex optimization have not been thoroughly studied. In this paper, we provide a fine-grained convergence analysis for a general class of adaptive gradient methods including AMSGrad, RMSProp and AdaGrad. For smooth nonconvex functions, we prove that adaptive gradient methods in expectation converge to a first-order stationary point. Our convergence rate is better than existing results for adaptive gradient methods in terms of dimension. In addition, we also prove high probability bounds on the convergence rates of AMSGrad, RMSProp as well as AdaGrad, which have not been established before. Our analyses shed light on better understanding the mechanism behind adaptive gradient methods in optimizing nonconvex objectives.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Dual Averaging Converges for Nonconvex Smooth Stochastic Optimization

    math.OC 2025-05 conditional novelty 7.0 of 10

    Stochastic dual averaging converges on nonconvex smooth stochastic optimization at rate O(1/T + σ log T/√T), matching SGD.

  2. FedACT: Federated Adaptive Coordinate Trust Modulation for Robust Transformer Training under Data Heterogeneity

    cs.LG 2026-07 conditional novelty 6.5 of 10

    Global-aware coordinate trust modulation after corrected AdamW updates improves federated Transformer and LLM training under data heterogeneity over strong adaptive baselines.

  3. Unified Scaling Laws for Compressed Representations

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A representation capacity derived from Gaussian fitting error predicts the training efficiency of sparse, quantized, and hybrid compressed models, and this capacity approximately multiplies across combined compression types.

  4. LightSAM: Parameter-Agnostic Sharpness-Aware Minimization

    cs.LG 2025-05 reject novelty 6.0 of 10

    An adaptive SAM variant using AdaGrad and Adam steps for both perturbation and update is claimed to converge at O(ln T / T^{1/4}) without tuning, but the Adam version still needs decaying hyperparameters and the proof...

  5. AdaGrad Meets Muon: Adaptive Stepsizes for Orthogonal Updates

    cs.LG 2025-09 conditional novelty 5.0 of 10

    The paper proposes AdaGO, a Muon variant with a scalar AdaGrad-Norm step size, and proves optimal nonconvex convergence rates while reporting empirical gains on regression and CIFAR-10.

Pith tools