Pith. sign in

REVIEW 1 cited by

A Comparative Analysis of the Optimization and Generalization Property of Two-layer Neural Network and Random Feature Models Under Gradient Descent Dynamics

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1904.04326 v2 pith:SHOBHRPU submitted 2019-04-08 cs.LG math.OCstat.ML

classification cs.LGmath.OCstat.ML
keywords networktrainingdescentdynamicsgeneralgradientneuralanalysis
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A fairly comprehensive analysis is presented for the gradient descent dynamics for training two-layer neural network models in the situation when the parameters in both layers are updated. General initialization schemes as well as general regimes for the network width and training data size are considered. In the over-parametrized regime, it is shown that gradient descent dynamics can achieve zero training loss exponentially fast regardless of the quality of the labels. In addition, it is proved that throughout the training process the functions represented by the neural network model are uniformly close to that of a kernel method. For general values of the network width and training data size, sharp estimates of the generalization error is established for target functions in the appropriate reproducing kernel Hilbert space.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Growing Neural Networks: Dynamic Evolution through Gradient Descent

    cs.LG 2025-01 conditional novelty 4.0 of 10

    Making network size itself a trainable parameter via an auxiliary weight or a controller mask lets small networks grow during gradient descent and outperform equivalent fixed-size networks on toy tasks.

Pith tools