Pith. sign in

REVIEW 3 cited by

An Improved Analysis of Training Over-parameterized Deep Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1906.04688 v1 pith:AEIGXQLW submitted 2019-06-11 cs.LG math.OCstat.ML

classification cs.LGmath.OCstat.ML
keywords neuraltrainingnetworksanalysisconditionconvergencedeepglobal
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

A recent line of research has shown that gradient-based algorithms with random initialization can converge to the global minima of the training loss for over-parameterized (i.e., sufficiently wide) deep neural networks. However, the condition on the width of the neural network to ensure the global convergence is very stringent, which is often a high-degree polynomial in the training sample size $n$ (e.g., $O(n^{24})$). In this paper, we provide an improved analysis of the global convergence of (stochastic) gradient descent for training deep neural networks, which only requires a milder over-parameterization condition than previous work in terms of the training sample size and other problem-dependent parameters. The main technical contributions of our analysis include (a) a tighter gradient lower bound that leads to a faster convergence of the algorithm, and (b) a sharper characterization of the trajectory length of the algorithm. By specializing our result to two-layer (i.e., one-hidden-layer) neural networks, it also provides a milder over-parameterization condition than the best-known result in prior work.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Stochastic AUC Maximization with Deep Neural Networks

    cs.LG 2019-08 conditional novelty 7.0 of 10

    Under the Polyak-Lojasiewicz condition, a proximal primal-dual algorithm and an AdaGrad-style variant maximize AUC with deep networks at O~(1/epsilon) sample complexity, with adaptive iteration complexity under slow c...

  2. A Geometric Characterization of the Stationary Plateau for Two-Layer Neural Networks

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    A geometric classification of stationary points on neuron-splitting plateaus in two-layer NN loss landscapes using the inner Hessian.

  3. Gradient Descent Finds Global Minima for Generalizable Deep Neural Networks of Practical Sizes

    stat.ML 2019-08 conditional novelty 6.0 of 10

    Gradient descent on a network with a final hidden layer of width O(n) can interpolate any n-point dataset and reach a global optimum, and this linear rate is optimal.

Pith tools