Pith. sign in

REVIEW 4 cited by

Convergence of Adam for Non-convex Objectives: Relaxed Hyperparameters and Non-ergodic Case

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.11782 v1 pith:LCTR252X submitted 2023-07-20 math.OC cs.LGcs.NAmath.NA

classification math.OCcs.LGcs.NAmath.NA
keywords convergenceadamnon-ergodicergodicnon-convexoptimizationstochasticcondition
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

Adam is a commonly used stochastic optimization algorithm in machine learning. However, its convergence is still not fully understood, especially in the non-convex setting. This paper focuses on exploring hyperparameter settings for the convergence of vanilla Adam and tackling the challenges of non-ergodic convergence related to practical application. The primary contributions are summarized as follows: firstly, we introduce precise definitions of ergodic and non-ergodic convergence, which cover nearly all forms of convergence for stochastic optimization algorithms. Meanwhile, we emphasize the superiority of non-ergodic convergence over ergodic convergence. Secondly, we establish a weaker sufficient condition for the ergodic convergence guarantee of Adam, allowing a more relaxed choice of hyperparameters. On this basis, we achieve the almost sure ergodic convergence rate of Adam, which is arbitrarily close to $o(1/\sqrt{K})$. More importantly, we prove, for the first time, that the last iterate of Adam converges to a stationary point for non-convex objectives. Finally, we obtain the non-ergodic convergence rate of $O(1/K)$ for function values under the Polyak-Lojasiewicz (PL) condition. These findings build a solid theoretical foundation for Adam to solve non-convex stochastic optimization problems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Unified convergence analysis for gradient descent optimization methods in the training of deep neural networks

    math.OC 2026-07 accept novelty 7.0 of 10

    Bounded trajectories of a broad class of GD optimizers (Adam, RMSprop, NAG, Adan, etc.) converge with polynomial rates to critical points of KL objectives with locally Lipschitz gradients, covering analytic-activation...

  2. AdamS: Momentum Itself Can Be A Normalizer for LLM Pretraining and Post-training

    cs.LG 2025-05 conditional novelty 6.0 of 10

    AdamS replaces AdamW's second-moment storage with a momentum-and-gradient squared denominator, matching AdamW's loss curves with half the optimizer memory.

  3. Sharp higher order convergence rates for the Adam optimizer

    math.OC 2025-04 conditional novelty 6.0 of 10

    Adam can achieve the accelerated momentum convergence rate locally on smooth strongly convex problems when its momentum and step size are tuned to the condition number, while RMSprop is shown to converge at the slower...

  4. Anisotropic Gaussian Smoothing for Gradient-based Optimization

    math.OC 2024-11 reject novelty 4.0 of 10

    Anisotropic Gaussian smoothing with step-dependent covariance matrices is inserted into GD, SGD, and Adam, and convergence bounds are derived that generalize the isotropic case.

Pith tools