Pith. sign in

REVIEW 3 cited by

A Theory of Non-Linear Feature Learning with One Gradient Step in Two-Layer Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.07891 v4 pith:GODG32OE submitted 2023-10-11 stat.ML cs.LG

classification stat.MLcs.LG
keywords learningfeaturenetworksneuralgradientnon-linearsteptraining
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Feature learning is thought to be one of the fundamental reasons for the success of deep neural networks. It is rigorously known that in two-layer fully-connected neural networks under certain conditions, one step of gradient descent on the first layer can lead to feature learning; characterized by the appearance of a separated rank-one component -- spike -- in the spectrum of the feature matrix. However, with a constant gradient descent step size, this spike only carries information from the linear component of the target function and therefore learning non-linear components is impossible. We show that with a learning rate that grows with the sample size, such training in fact introduces multiple rank-one components, each corresponding to a specific polynomial feature. We further prove that the limiting large-dimensional and large sample training and test errors of the updated neural networks are fully characterized by these spikes. By precisely analyzing the improvement in the training and test errors, we demonstrate that these non-linear features can enhance learning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning Regime

    cs.LG 2025-09 conditional novelty 6.0 of 10

    For diagonal and quadratic two-layer networks, training maps to LASSO and matrix compressed sensing, yielding a full phase diagram of excess-risk scaling exponents and a spectral characterization of the trained weights.

  2. Learning Hierarchical Polynomials of Multiple Nonlinear Features with Three-Layer Networks

    cs.LG 2024-11 conditional novelty 6.0 of 10

    A three-layer network with layerwise gradient descent provably recovers the span of multiple quadratic features in O~(d^4) samples and then learns any polynomial link in the features.

  3. Towards a Statistical Understanding of Neural Networks: Beyond the Neural Tangent Kernel Theories

    cs.LG 2024-12 conditional novelty 4.0 of 10

    The paper reviews fixed-kernel neural network theory and proposes an over-parameterized Gaussian sequence model as a prototype for feature learning.

Pith tools