Pith. sign in

REVIEW 4 cited by

Generalization Ability of Wide Neural Networks on $\mathbb{R}$

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.05933 v1 pith:2QXST2LZ submitted 2023-02-12 stat.ML cs.LG

classification stat.MLcs.LG
keywords neuralnetworkmathbbwideabilitygeneralizationkernelminimax
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We perform a study on the generalization ability of the wide two-layer ReLU neural network on $\mathbb{R}$. We first establish some spectral properties of the neural tangent kernel (NTK): $a)$ $K_{d}$, the NTK defined on $\mathbb{R}^{d}$, is positive definite; $b)$ $\lambda_{i}(K_{1})$, the $i$-th largest eigenvalue of $K_{1}$, is proportional to $i^{-2}$. We then show that: $i)$ when the width $m\rightarrow\infty$, the neural network kernel (NNK) uniformly converges to the NTK; $ii)$ the minimax rate of regression over the RKHS associated to $K_{1}$ is $n^{-2/3}$; $iii)$ if one adopts the early stopping strategy in training a wide neural network, the resulting neural network achieves the minimax rate; $iv)$ if one trains the neural network till it overfits the data, the resulting neural network can not generalize well. Finally, we provide an explanation to reconcile our theory and the widely observed ``benign overfitting phenomenon''.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Cost of Discretization in Functional Linear Regression: Minimax Rates and Adaptation

    math.ST 2026-07 accept novelty 7.0 of 10

    Matching minimax prediction rates for discretely observed functional linear regression are n^{-ν/(ν+1)}+(nm)^{-ν/κ} under independent design, and those two terms plus m^{-ν}+m^{-4α} under common design.

  2. Breaking the Curse with BAND: Nonparametric Distribution Estimation in High Dimensions

    stat.ML 2026-07 conditional novelty 6.0 of 10

    Sparse Bayesian-network factorization plus sparsity-aware regression yields polynomial TV rates for high-dimensional mixed-type distribution estimation, beating classical histogram rates under sparsity.

  3. A ZeNN architecture to avoid the Gaussian trap

    cs.LG 2025-05 conditional novelty 6.0 of 10

    ZeNNs, which replace the equal-weight average of MLP neurons with an index-weighted sum of frequency-scaled neurons, provably converge pointwise, retain non-Gaussian limits, and learn high-frequency features in low-di...

  4. Uncertainty Quantification and Causal Considerations for Off-Policy Decision Making

    stat.ML 2025-02 conditional novelty 6.0 of 10

    Three methods for off-policy evaluation: marginal ratio variance reduction, conformal predictive intervals, and causal bounds that falsify digital twins under unmeasured confounding.

Pith tools