Pith. sign in

REVIEW 3 cited by

Convergence of Clipped SGD on Convex $(L_0,L_1)$-Smooth Functions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.16492 v2 pith:GGYIGKFX submitted 2025-02-23 math.OC

classification math.OC
keywords gradientclippingconvergenceconvexfunctionsratesmoothsmoothness
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

We study stochastic gradient descent (SGD) with gradient clipping on convex functions under a generalized smoothness assumption called $(L_0,L_1)$-smoothness. Using gradient clipping, we establish a high probability convergence rate that matches the SGD rate in the $L$ smooth case up to polylogarithmic factors and additive terms. We also propose a variation of adaptive SGD with gradient clipping, which achieves the same guarantee. We perform empirical experiments to examine our theory and algorithmic choices.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Decentralized Stochastic Nonconvex Optimization under the $(L_0,L_1)$-Smoothness

    math.OC 2025-09 conditional novelty 6.0 of 10

    DNSGD is a decentralized normalized stochastic gradient method for (L0,L1)-smooth nonconvex optimization, with complexity bounds that match standard smooth decentralized results when L1=0.

  2. Power of Generalized Smoothness in Stochastic Convex Optimization: First- and Zero-Order Algorithms

    math.OC 2025-01 conditional novelty 6.0 of 10

    For convex stochastic optimization under (L0,L1)-smoothness, clipped and normalized SGD (and zero-order variants) obtain linear-rate terms in their convergence bounds, and in the L0=0 regime NSGD achieves logarithmic ...

  3. Why Do We Need Warm-up? A Theoretical Perspective

    cs.LG 2025-10 conditional novelty 5.0 of 10

    Under the proposed (H0,H1)-smoothness condition, gradient descent with a warm-up-style adaptive step-size provably converges faster than with any fixed step-size.

Pith tools