REVIEW 3 cited by
On the Convergence of Gradient Descent Training for Two-layer ReLU-networks in the Mean Field Regime
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We describe a necessary and sufficient condition for the convergence to minimum Bayes risk when training two-layer ReLU-networks by gradient descent in the mean field regime with omni-directional initial parameter distribution. This article extends recent results of Chizat and Bach to ReLU-activated networks and to the situation in which there are no parameters which exactly achieve MBR. The condition does not depend on the initalization of parameters and concerns only the weak convergence of the realization of the neural network, not its parameter distribution.
Forward citations
Cited by 3 Pith papers
-
Elliptic Regularity Theory in Barron Spaces and Applications to the Deep Ritz Method
Harmonic functions with Barron Dirichlet data fail to be Lipschitz or H², yet admit Barron approximants of norm ~|log ε| with error ~ε on half-spaces and 2D rectangles, giving Deep Ritz a priori rates.
-
Singular perturbations and hierarchical learning in two-layer neural networks
Constant and linear Hermite components of a misspecified single-index target are recovered at the conjectured singular-perturbation timescales; quadratic learning remains coupled to them via an auxiliary constrained flow.
-
Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization
Weight-decay-regularized two-layer ReLU networks need width exponential in the number of samples for a benign loss landscape, and small initialization can still converge to spurious minima.
Discussion (0). Continue with ORCID to comment.