Pith. sign in

REVIEW 4 cited by

On the Heavy-Tailed Theory of Stochastic Gradient Descent for Deep Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1912.00018 v1 pith:ZUUMESTN submitted 2019-11-29 stat.ML cs.LGmath.CA

classification stat.MLcs.LGmath.CA
keywords alphaemphassumptiondeepgradientminimastablestochastic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

The gradient noise (GN) in the stochastic gradient descent (SGD) algorithm is often considered to be Gaussian in the large data regime by assuming that the \emph{classical} central limit theorem (CLT) kicks in. This assumption is often made for mathematical convenience, since it enables SGD to be analyzed as a stochastic differential equation (SDE) driven by a Brownian motion. We argue that the Gaussianity assumption might fail to hold in deep learning settings and hence render the Brownian motion-based analyses inappropriate. Inspired by non-Gaussian natural phenomena, we consider the GN in a more general context and invoke the \emph{generalized} CLT, which suggests that the GN converges to a \emph{heavy-tailed} $\alpha$-stable random vector, where \emph{tail-index} $\alpha$ determines the heavy-tailedness of the distribution. Accordingly, we propose to analyze SGD as a discretization of an SDE driven by a L\'{e}vy motion. Such SDEs can incur `jumps', which force the SDE and its discretization \emph{transition} from narrow minima to wider minima, as proven by existing metastability theory and the extensions that we proved recently. In this study, under the $\alpha$-stable GN assumption, we further establish an explicit connection between the convergence rate of SGD to a local minimum and the tail-index $\alpha$. To validate the $\alpha$-stable assumption, we conduct experiments on common deep learning scenarios and show that in all settings, the GN is highly non-Gaussian and admits heavy-tails. We investigate the tail behavior in varying network architectures and sizes, loss functions, and datasets. Our results open up a different perspective and shed more light on the belief that SGD prefers wide minima.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Statistical Inference for Stochastic Gradient Descent: Beyond Finite Variance

    stat.ML 2026-05 unverdicted novelty 7.0 of 10

    Presents a self-normalized subsampling procedure for asymptotically valid confidence regions from SGD iterates under both finite and infinite variance assumptions.

  2. Nonconvex Decentralized Stochastic Bilevel Optimization under Heavy-Tailed Noise

    cs.LG 2025-09 conditional novelty 6.0 of 10

    The paper introduces D-NSVRGDA, a decentralized normalized variance-reduced method for nonconvex bilevel optimization, and proves the first convergence rate under heavy-tailed noise without gradient clipping.

  3. Lions and Muons: Optimization via Stochastic Frank-Wolfe under Heavy-Tailed Noise

    math.OC 2025-06 reject novelty 6.0 of 10

    Lion and Muon with weight decay are shown to be instances of one stochastic Frank-Wolfe algorithm, and clipped and variance-reduced variants get the first high-probability convergence rates for nonconvex Frank-Wolfe u...

  4. Improving Adaptive Moment Optimization via Preconditioner Diagonalization

    cs.LG 2025-02 conditional novelty 3.0 of 10

    Rotating gradients into their SVD coordinate system before Adam-style updates can roughly halve the number of steps LLaMA models need to reach a given perplexity.

Pith tools