REVIEW 2 cited by
Improved Convergence in High Probability of Clipped Gradient Methods with Heavy Tails
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
In this work, we study the convergence \emph{in high probability} of clipped gradient methods when the noise distribution has heavy tails, ie., with bounded $p$th moments, for some $1<p\le2$. Prior works in this setting follow the same recipe of using concentration inequalities and an inductive argument with union bound to bound the iterates across all iterations. This method results in an increase in the failure probability by a factor of $T$, where $T$ is the number of iterations. We instead propose a new analysis approach based on bounding the moment generating function of a well chosen supermartingale sequence. We improve the dependency on $T$ in the convergence guarantee for a wide range of algorithms with clipped gradients, including stochastic (accelerated) mirror descent for convex objectives and stochastic gradient descent for nonconvex objectives. This approach naturally allows the algorithms to use time-varying step sizes and clipping parameters when the time horizon is unknown, which appears impossible in prior works. We show that in the case of clipped stochastic mirror descent, problem constants, including the initial distance to the optimum, are not required when setting step sizes and clipping parameters.
Forward citations
Cited by 2 Pith papers
-
The Convergence Behavior of Adam under Heavy-Tailed Noise
Under heavy-tailed noise with bounded p-th moments, vector-form Adam converges to (ρ,ε)-stationary points at rate O(ε^{-(5p/(3p-4)+3/2)}) for p∈(4/3,2]; with known-radius clipping the rate is optimal O(ε^{-(p/(p-1)+3/2)}).
-
High Probability Convergence of Distributed Clipped Stochastic Gradient Descent with Heavy-tailed Noise
Distributed clipped stochastic gradient descent over time-varying directed graphs is proved to converge with high probability under heavy-tailed gradient noise.
Discussion (0). Sign in to comment.