Pith. sign in

REVIEW 2 cited by

Momentum via Primal Averaging: Theoretical Insights and Learning Rate Schedules for Non-Convex Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.00406 v4 pith:2RZQWCED submitted 2020-10-01 cs.LG math.OCstat.ML

classification cs.LGmath.OCstat.ML
keywords momentumnon-convexanalysisaveraginginsightslearningprimalschedules
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Momentum methods are now used pervasively within the machine learning community for training non-convex models such as deep neural networks. Empirically, they out perform traditional stochastic gradient descent (SGD) approaches. In this work we develop a Lyapunov analysis of SGD with momentum (SGD+M), by utilizing a equivalent rewriting of the method known as the stochastic primal averaging (SPA) form. This analysis is much tighter than previous theory in the non-convex case, and due to this we are able to give precise insights into when SGD+M may out-perform SGD, and what hyper-parameter schedules will work and why.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Understanding Schedule-Free Methods in Nonconvex Optimization: Rate Guarantees and Escaping Saddles

    cs.LG 2026-07 accept novelty 6.5 of 10

    Standard Schedule-Free GD and SGD attain optimal nonconvex first-order rates via Lyapunov analysis of their continuous-time limit, and avoid strict saddles under arbitrarily small one-time noise.

  2. Analysis of Schedule-Free Nonconvex Optimization

    cs.LG 2025-08 conditional novelty 4.0 of 10

    A Lyapunov framework yields O(1/log T) and O(log T/T) gradient-norm rates for Schedule-Free on smooth nonconvex objectives, with the faster rate depending on an unproven assumption.

Pith tools