Pith. sign in

REVIEW 1 cited by

How many Neurons do we need? A refined Analysis for Shallow Networks trained with Gradient Descent

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.08044 v1 pith:7MWP4KRC submitted 2023-09-14 stat.ML cs.LG

classification stat.MLcs.LG
keywords descentgeneralizationgradientkernelnetworksneuralneuronsregression
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We analyze the generalization properties of two-layer neural networks in the neural tangent kernel (NTK) regime, trained with gradient descent (GD). For early stopped GD we derive fast rates of convergence that are known to be minimax optimal in the framework of non-parametric regression in reproducing kernel Hilbert spaces. On our way, we precisely keep track of the number of hidden neurons required for generalization and improve over existing results. We further show that the weights during training remain in a vicinity around initialization, the radius being dependent on structural assumptions such as degree of smoothness of the regression function and eigenvalue decay of the integral operator associated to the NTK.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Optimal Convergence Rates for Neural Operators

    stat.ML 2024-12 conditional novelty 5.0 of 10

    Two-layer neural operators trained with early-stopped gradient descent achieve the same minimax convergence rates as kernel methods in the neural tangent kernel regime.

Pith tools