REVIEW 3 major objections 3 minor 103 references
A Classical View on Benign Overfitting: The Role of Sample Size
T0 review · 3 major / 3 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Benign overfitting can occur in the classical U-shaped regime when sample size and model complexity grow together, the paper proves for kernel ridge regression and trained two-layer ReLU networks.
desk verdict The proof machinery is real, but the ReLU NTK's infinite-dimensional kernel sinks the 'any bounded f*' theorem; major revision needed. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the approximation–estimation decomposition of the excess risk along the gradient-flow trajectory, with gradient flow itself treated as an implicit regularizer. The paper compares an $m$-neuron ReLU network $\hat f_t$ trained on the empirical risk with an oracle network $f_t$ of the same architecture trained on the population risk: the oracle gap $\|f_t-f_*\|_2$ is the approximation error, and the tracking gap $\|\hat f_t-f_t\|_2$ is the estimation error. The implicit regularizer is the stopping time $T_\varepsilon = 2\lambda_\varepsilon^{-1}\log(2/\sqrt\varepsilon)$, where $\lambda_\varepsilon$ is the $L_\varepsilon$-th eigenvalue of the analytical neural tangent kernel operator, chosen so that the component of $f_*$ in the remaining high-frequency eigenspace has norm at most $\sqrt\varepsilon/4$. Three supporting devices carry the proofs: a spectral-norm bound for Hadamard products extended from matrices to integral operators (used for the approximation error), new concentration inequalities for vector-valued U- and V-statistics that allow the estimation-error bound to be made at initialization and then iterated $U_\varepsilon$ times with the factorial $U_\varepsilon!$ absorbing the remainder, and a real-induction argument that keeps the weights near initialization so the minimum eigenvalue of the NTK Gram matrix stays bounded along the whole trajectory.
What would settle it
Evaluate the theorem's dimension–confidence trade-off at the dimensions its experiments use: with $d=7$ (Abalone), Assumption 2(i) requires $\delta\ge 12e^{-7}\approx 0.011$, so at $\delta=0.001$ the theorem's hypotheses are void. Running the Abalone-style experiment at $d=7$ with $\delta$ below that threshold — or checking whether a bounded target whose mass sits mostly on the low eigenvalues of the NTK operator still allows both risks below $\varepsilon$ at fixed $d$ — would show whether the claim survives outside the stated scaling regime.
Extended reading notes
Core claim
The central discovery, stated on the paper's own terms, is that whether overfitting is benign is decided by the joint scaling of sample size and model complexity, not by model complexity alone. The paper defines benign overfitting as achieving empirical risk at most $\varepsilon$ and excess risk $R(\hat f)-R(f_*) \le \varepsilon$ simultaneously, for every $\varepsilon,\delta>0$ with probability at least $1-\delta$ (Definition 1), deliberately allowing training error that is small but not exactly zero. For kernel ridge regression with the neural tangent kernel, Theorem 6 shows that when the regularization parameter $\gamma$ is small and the sample size $n$ is large relative to $\varepsilon$ and $\delta$ (Assumption 1), the regularized empirical risk minimizer achieves both bounds on a single high-probability event. For two-layer ReLU networks trained by full-batch gradient flow in the NTK regime — where each neuron moves so little that the network behaves like a kernel method — Theorem 11 shows that for any essentially bounded regression function $f_*$, if input dimension, width $m$, and sample size $n$ satisfy Assumptions 2 and 3, then at stopping time $T_\varepsilon = \frac{2}{\lambda_\varepsilon}\log\frac{2}{\sqrt\varepsilon}$ the trained network satisfies both bounds on the same event of probability at least $1-\delta$. The authors claim this as the first generalization result in this setting that assumes nothing about the regression function or the noise beyond boundedness.
Load-bearing premise
The load-bearing premise is a set of scaling relations: the input dimension $d$ must be at least logarithmic in the inverse failure probability ($e^{-d}\le\delta/12$), and the width $m$ and sample size $n$ must be large enough relative to $d$, $\varepsilon$, and the spectral gap $\lambda_\varepsilon$; if these fail, the high-probability minimum-eigenvalue bounds on which every theorem rests are unavailable.
Editorial extensions
If this is right
- Benign overfitting does not require leaving the classical regime: a practitioner who adds data and model capacity together can push both training error and test error below any desired level, without high-dimensional inputs or a well-specified model class.
- For two-layer ReLU networks in the NTK regime, generalization to the Bayes-optimal risk is achievable for every bounded regression function, with no RKHS-of-NTK assumption and no noise model beyond boundedness.
- The KRR result is not tied to the specific kernel: the same proof works for any bounded kernel whose Gram matrix has the required minimum-eigenvalue lower bound and whose RKHS is dense in $L^2(\rho_{d-1})$.
- Because the analysis deliberately avoids uniform convergence over the parameter space, the bounds do not deteriorate as the width $m$ grows; more parameters do not hurt the generalization guarantee.
- The experiments corroborate the predicted 'down and to the right' shift: on Abalone, Wine, and synthetic sphere data, the point where excess risk crosses below empirical risk moves later in training and lower in value as $n$ increases, and both risks at the crossing fall.
Reading between the lines
- I read the Section 2 narrative as a reconciliation claim the paper states but does not fully prove: earlier interpolation-regime results hold the model fixed on the upward slope of a U-curve drawn at a fixed sample size, whereas this paper moves along the moving trough; a head-to-head comparison of the two regimes under identical scaling is a natural next step.
- The vector-valued U-/V-statistic concentration bounds and the iterated-remainder trick are separable tools that should transfer to other settings needing high-probability control of two coupled trajectories without uniform convergence, such as stochastic gradient descent or feature-learning dynamics.
- A testable extension suggested by the assumptions: the paper notes that smooth activations would simplify the proofs, so if the mechanism is really the approximation–estimation trade-off rather than ReLU specifics, the benign-overfitting window should widen (smaller required $n$ and $m$) for smooth activations.
- The conclusion flags that only upper bounds are provided and matching lower bounds remain open; I draw the consequence that the 'low-dimensional inputs' claim should be read as conditioned on $d$ growing with $\log(1/\delta)$, given Assumption 2(i)'s requirement $e^{-d}\le\delta/12$.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes reinterpreting benign overfitting through 'almost benign overfitting,' where empirical risk and excess risk can both be made arbitrarily small by jointly increasing sample size and model complexity. The authors prove such guarantees for two case studies: kernel ridge regression with the ReLU neural tangent kernel, and least-squares regression with a two-layer ReLU network trained by gradient flow in the NTK regime. The main advertised contribution is that these results hold under essentially no assumptions on the regression function beyond boundedness, and in low input dimension, using a novel approximation/estimation decomposition of the excess risk that avoids uniform convergence. The paper also provides experiments on synthetic, Abalone, and Wine data that show risk curves consistent with the hypothesis.
Significance. If the main theorem were correct, the paper would represent a substantial conceptual shift: benign overfitting would no longer require high dimensionality, kernel eigenvalue decay conditions, or interpolation, and the neural-network result would be the first generalization guarantee of this type for arbitrary bounded regression functions. The proof machinery is genuinely elaborate and contains independently interesting tools, including a Hadamard-product bound for integral operators and concentration results for vector-valued U- and V-statistics. The paper is also honest about its limitations, such as only providing upper bounds. However, these strengths do not compensate for the central technical gaps described below, which affect the core claims of the paper.
major comments (3)
- [Appendix D.2.3 and Eq. (4.1)] The ReLU NTK on the sphere has an infinite-dimensional kernel, and this invalidates the claimed arbitrary-bounded-f* result. The paper's own spectral calculation in Appendix D.2.3 gives mu_h = 0 for every odd h >= 3, while positive eigenvalues exist for h = 1, h = 2, and infinitely many even h >= 4. Consequently, the operator H has both an infinite-dimensional positive eigenspace and an infinite-dimensional zero eigenspace, so the statement that the eigenfunctions form an orthonormal basis with lambda_1 >= lambda_2 >= ... and lambda_l -> 0 'from above' is not correct as an enumeration with nonzero eigenvalues only. More importantly, for a bounded f* with a nonzero component on an odd spherical harmonic of degree at least 3, for example the normalized Y_3, Eq. (4.1) cannot be satisfied with lambda_epsilon > 0: if Y_3 is left in the tail, the tail norm is 1 for every finite L, and if Y_3 is placed among the top L eigenfunctions, the corresponding eigenvalue is 0, so lambda_epsilon = 0 and T_epsilon in Eq. (4.2) is undefined. Hence Theorem 8, and therefore Theorems 10 and 11, do not hold for arbitrary bounded f* as claimed in the abstract and in Section 4.2.
- [Section 3, paragraph 'By the denseness of H in L2'] The KRR section relies on the assertion that the RKHS of the ReLU NTK is dense in L2(rho_{d-1}), writing 'By the denseness of H in L2(rho), there is an f_epsilon in H such that ||f* - f_epsilon||_2^2 <= epsilon/8.' This assertion is false for the ReLU NTK: as the spectral calculation in Appendix D.2.3 shows, the kernel misses all odd spherical harmonics of degree at least 3, so functions such as Y_3 are orthogonal to the RKHS. Therefore the approximation argument behind Theorem 3 and the benign-overfitting conclusion in Theorem 6 only hold for regression functions lying in a proper subspace, not for arbitrary bounded f*. This is a load-bearing error in the paper's central claim of no assumptions on f*.
- [Definition 1 vs. Assumptions 1(i) and 2(i)] Definition 1 requires that for every epsilon, delta > 0 there exists a sample size n such that, with probability at least 1 - delta, both risks are at most epsilon. For a fixed problem, the input dimension d is fixed, but Assumptions 1(i) and 2(i) require e^{-d} <= delta/4 and e^{-d} <= delta/12, respectively, i.e., d >= Omega(log(1/delta)). Thus, for any fixed low-dimensional problem, the stated theorems cannot cover arbitrarily small delta, and the results do not achieve Definition 1. The paper's claim that the results 'hold on low-dimensional inputs' is therefore only true in the limited sense that d is not required to grow with n or m, but d must still grow without bound as delta tends to zero. This is a mismatch between the theorem statements and the formal definition of the object they claim to establish.
minor comments (3)
- [Title of Section 4 and Figure captions] There is a typo in the Section 4 heading: 'Bengin Overfitting' should read 'Benign Overfitting.' The same typo appears in the table of contents.
- [Appendix D.2.3, paragraph after spectral decomposition] The phrase 'lambda_l -> 0 from above' is misleading and should be replaced by a correct treatment of the zero eigenspace; the current wording suggests that all eigenvalues are positive, which contradicts the formula mu_h = 0 for odd h >= 3 appearing a few lines later.
- [Section 4.3 and Appendix D.8] The experiments use gradient descent with learning rate 0.1, while the theory is developed for gradient flow; the paper should state explicitly that discretization is not covered by the proofs and that the experiments are only heuristic support.
Circularity Check
No significant circularity: the derivation is self-contained and the parameters are chosen to satisfy stated inequalities rather than fitted to the target risks.
full rationale
The paper's central claims are derived from stated assumptions without the target bound being used as an input. In the KRR section, f_gamma and hat f_gamma are explicit regularized minimizers; epsilon, gamma, and n are chosen sequentially so that the approximation and estimation bounds each contribute at most sqrt(epsilon)/2. No constant is fitted to the empirical or excess risk. In the neural-network section, L_epsilon, lambda_epsilon, T_epsilon, and U_epsilon are analysis quantities: L_epsilon truncates f* by tail norm, T_epsilon is obtained by solving the independently derived exponential-decay inequality exp(-lambda_epsilon t/2) <= sqrt(epsilon)/2, and U_epsilon is a factorial remainder device. These choices are not predictions; they are proof parameters. The high-probability event E3 is built by union bounds from concentration lemmas, and Theorems 7, 8, and 9 hold on that same event by deterministic arguments. The citations to Park and Muandet (2020, 2023) are auxiliary standard closed-form and concentration results; they are not load-bearing self-citations and do not presuppose benign overfitting. The manuscript's own limitation statement (upper bounds only) and the possible spectral obstruction from the zero eigenvalues for odd h >= 3 in Appendix D.2.3 concern correctness or assumption-satisfiability, not circularity: the proof would still derive the stated bound whenever the assumptions hold. Therefore no step reduces by construction to its inputs.
Assumptions & free parameters
free parameters (6)
- gamma (KRR regularization)
- L_epsilon (spectral truncation level)
- lambda_epsilon = lambda_{L_epsilon}
- T_epsilon (gradient flow time) =
2/lambda_epsilon log(2/sqrt(eps))
- U_epsilon =
smallest U with (1/U!)(8T_epsilon/d)^U <= sqrt(eps)/14
- Width m and sample size n
assumptions (10)
- standard math Real induction (Hathaway 2011, Clark 2019)
- standard math Spectral theory for compact self-adjoint operators
- standard math Concentration inequalities: Hoeffding, McDiarmid, Matrix Chernoff, vector-valued Hoeffding (Pinelis)
- domain assumption x is uniform on the sphere S^{d-1}
- domain assumption |y| <= 1 almost surely, hence |f*| <= 1 and ||f*||_2 <= 1
- domain assumption The NTK kernel RKHS is dense in L2(rho)
- ad hoc to paper Antisymmetric initialization makes the network output exactly zero at initialization
- ad hoc to paper Output layer weights are fixed random +/-1
- ad hoc to paper NTK (lazy training) regime with m sufficiently large
- ad hoc to paper Relaxation to epsilon-tolerance in Definition 1 ('almost benign overfitting')
Cite this review
Pith. "Pith review of A Classical View on Benign Overfitting: The Role of Sample Size." pith.science (2026). https://pith.science/paper/AON5HGFA
@misc{pith2026250511621,
author = {Pith},
title = {Pith review of: A Classical View on Benign Overfitting: The Role of Sample Size},
year = {2026},
howpublished = {\url{https://pith.science/paper/AON5HGFA}},
note = {Machine review of arXiv:2505.11621}
}
read the original abstract
Benign overfitting is a phenomenon in machine learning where a model perfectly fits (interpolates) the training data, including noisy examples, yet still generalizes well to unseen data. Understanding this phenomenon has attracted considerable attention in recent years. In this work, we introduce a conceptual shift, by focusing on almost benign overfitting, where models simultaneously achieve both arbitrarily small training and test errors. This behavior is characteristic of neural networks, which often achieve low (but non-zero) training error while still generalizing well. We hypothesize that this almost benign overfitting can emerge even in classical regimes, by analyzing how the interaction between sample size and model complexity enables larger models to achieve both good training fit but still approach Bayes-optimal generalization. We substantiate this hypothesis with theoretical evidence from two case studies: (i) kernel ridge regression, and (ii) least-squares regression using a two-layer fully connected ReLU neural network trained via gradient flow. In both cases, we overcome the strong assumptions often required in prior work on benign overfitting. Our results on neural networks also provide the first generalization result in this setting that does not rely on any assumptions about the underlying regression function or noise, beyond boundedness. Our analysis introduces a novel proof technique based on decomposing the excess risk into estimation and approximation errors, interpreting gradient flow as an implicit regularizer, that helps avoid uniform convergence traps. This analysis idea could be of independent interest.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
The N eural T angent K ernel in H igh D imensions: T riple D escent and a M ulti- S cale T heory of G eneralization
Ben Adlam and Jeffrey Pennington. The N eural T angent K ernel in H igh D imensions: T riple D escent and a M ulti- S cale T heory of G eneralization. In International Conference on Machine Learning, pages 74--84. PMLR, 2020
2020
-
[2]
Stefan Aeberhard and M. Forina. Wine . UCI Machine Learning Repository, 1992. DOI : https://doi.org/10.24432/C5PC7J
doi:10.24432/c5pc7j 1992
-
[3]
Learning and G eneralization in O verparameterized N eural N etworks, G oing B eyond T wo L ayers
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang. Learning and G eneralization in O verparameterized N eural N etworks, G oing B eyond T wo L ayers. Advances in neural information processing systems, 32, 2019 a
2019
-
[4]
A C onvergence T heory for D eep L earning via O ver- P arameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song. A C onvergence T heory for D eep L earning via O ver- P arameterization. In International conference on machine learning, pages 242--252. PMLR, 2019 b
2019
-
[5]
Fine- G rained A nalysis of O ptimization and G eneralization for O verparameterized T wo- L ayer N eural N etworks
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li, and Ruosong Wang. Fine- G rained A nalysis of O ptimization and G eneralization for O verparameterized T wo- L ayer N eural N etworks. In International Conference on Machine Learning, pages 322--332. PMLR, 2019
2019
-
[6]
Sharp E stimates for E igenvalues of I ntegral O perators G enerated by D ot P roduct K ernels on the S phere
Douglas Azevedo and Valdir Antonio Menegatto. Sharp E stimates for E igenvalues of I ntegral O perators G enerated by D ot P roduct K ernels on the S phere. Journal of Approximation Theory, 177: 0 57--68, 2014
2014
-
[7]
On the I mplicit B ias of I nitialization S hape: B eyond I nfinitesimal M irror D escent
Shahar Azulay, Edward Moroshko, Mor Shpigel Nacson, Blake E Woodworth, Nathan Srebro, Amir Globerson, and Daniel Soudry. On the I mplicit B ias of I nitialization S hape: B eyond I nfinitesimal M irror D escent. In International Conference on Machine Learning, pages 468--477. PMLR, 2021
2021
-
[8]
Benign O verfitting in L inear R egression
Peter L Bartlett, Philip M Long, G \'a bor Lugosi, and Alexander Tsigler. Benign O verfitting in L inear R egression. Proceedings of the National Academy of Sciences, 117 0 (48): 0 30063--30070, 2020
2020
Show all 103 references
-
[9]
Deep L earning: A S tatistical V iewpoint
Peter L Bartlett, Andrea Montanari, and Alexander Rakhlin. Deep L earning: A S tatistical V iewpoint. Acta numerica, 30: 0 87--201, 2021
2021
-
[10]
Generalization in K ernel R egression U nder R ealistic A ssumptions
Daniel Barzilai and Ohad Shamir. Generalization in K ernel R egression U nder R ealistic A ssumptions. In Forty-first International Conference on Machine Learning, 2024
2024
-
[11]
On the I nconsistency of K ernel R idgeless R egression in F ixed D imensions
Daniel Beaglehole, Mikhail Belkin, and Parthe Pandit. On the I nconsistency of K ernel R idgeless R egression in F ixed D imensions. SIAM Journal on Mathematics of Data Science, 5 0 (4): 0 854--872, 2023
2023
-
[12]
Reconciling M odern M achine- L earning P ractice and the C lassical B ias-- V ariance T rade- O ff
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal. Reconciling M odern M achine- L earning P ractice and the C lassical B ias-- V ariance T rade- O ff. Proceedings of the National Academy of Sciences, 116 0 (32): 0 15849--15854, 2019
2019
-
[13]
Reproducing K ernel H ilbert S paces in P robability and S tatistics
Alain Berlinet and Christine Thomas-Agnan. Reproducing K ernel H ilbert S paces in P robability and S tatistics . Springer Science & Business Media, 2004
2004
-
[14]
On the I nductive B ias of N eural T angent K ernels
Alberto Bietti and Julien Mairal. On the I nductive B ias of N eural T angent K ernels. Advances in Neural Information Processing Systems, 32, 2019
2019
-
[15]
Implicit B ias of MSE G radient O ptimization in U nderparameterized N eural N etworks
Benjamin Bowman and Guido Montufar. Implicit B ias of MSE G radient O ptimization in U nderparameterized N eural N etworks. In International Conference on Learning Representations, 2021
2021
-
[16]
Spectral B ias O utside T he T raining S et for D eep N etworks in the K ernel R egime
Benjamin Bowman and Guido F Montufar. Spectral B ias O utside T he T raining S et for D eep N etworks in the K ernel R egime. Advances in Neural Information Processing Systems, 35: 0 30362--30377, 2022
2022
-
[17]
Kernel I nterpolation in S obolev S paces is not C onsistent in L ow D imensions
Simon Buchholz. Kernel I nterpolation in S obolev S paces is not C onsistent in L ow D imensions. In Conference on Learning Theory, pages 3410--3440. PMLR, 2022
2022
-
[18]
Towards understanding the spectral bias of deep learning
Yuan Cao, Zhiying Fang, Yue Wu, Ding-Xuan Zhou, and Quanquan Gu. Towards understanding the spectral bias of deep learning. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, pages 2205--2211. International Joint Conferences on Artificial...
2021
-
[19]
Benign O verfitting in T wo- L ayer C onvolutional N eural N etworks
Yuan Cao, Zixiang Chen, Misha Belkin, and Quanquan Gu. Benign O verfitting in T wo- L ayer C onvolutional N eural N etworks. Advances in neural information processing systems, 35: 0 25237--25250, 2022
2022
-
[20]
Optimal R ates for the R egularized L east- S quares A lgorithm
Andrea Caponnetto and Ernesto De Vito. Optimal R ates for the R egularized L east- S quares A lgorithm. Foundations of Computational Mathematics, 7: 0 331--368, 2007
2007
-
[21]
Characterizing O verfitting in K ernel R idgeless R egression T hrough the E igenspectrum
Tin Sum Cheng, Aurelien Lucchi, Anastasis Kratsios, and David Belius. Characterizing O verfitting in K ernel R idgeless R egression T hrough the E igenspectrum. arXiv preprint arXiv:2402.01297, 2024
2024 arXiv
-
[22]
On the R obustness of the M inimim l2 I nterpolator
Geoffrey Chinot and Matthieu Lerasle. On the R obustness of the M inimim l2 I nterpolator. Bernoulli, 2022
2022
-
[23]
On the G lobal C onvergence of G radient D escent for O ver- P arameterized M odels using O ptimal T ransport
Lenaic Chizat and Francis Bach. On the G lobal C onvergence of G radient D escent for O ver- P arameterized M odels using O ptimal T ransport. Advances in neural information processing systems, 31, 2018
2018
-
[24]
The I nstructor’s G uide to R eal I nduction
Pete L Clark. The I nstructor’s G uide to R eal I nduction. Mathematics Magazine, 92 0 (2): 0 136--150, 2019
2019
-
[25]
A U - T urn on D ouble D escent: R ethinking P arameter C ounting in S tatistical L earning
Alicia Curth, Alan Jeffares, and Mihaela van der Schaar. A U - T urn on D ouble D escent: R ethinking P arameter C ounting in S tatistical L earning. In Advances in Neural Information Processing Systems, volume 36, 2023
2023
-
[26]
Gradient D escent F inds G lobal M inima of D eep N eural N etworks
Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai. Gradient D escent F inds G lobal M inima of D eep N eural N etworks. In International conference on machine learning, pages 1675--1685. PMLR, 2019 a
2019
-
[27]
Gradient D escent P rovably O ptimizes O ver- P arameterized N eural N etworks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh. Gradient D escent P rovably O ptimizes O ver- P arameterized N eural N etworks. In International Conference on Learning Representations, 2019 b
2019
-
[28]
A C omparative A nalysis of O ptimization and G eneralization P roperties of T wo- L ayer N eural N etwork and R andom F eature M odels under G radient D escent D ynamics
Weinan E, Chao Ma, and Lei Wu. A C omparative A nalysis of O ptimization and G eneralization P roperties of T wo- L ayer N eural N etwork and R andom F eature M odels under G radient D escent D ynamics. Sci. China Math, 2019
2019
-
[29]
Benign O verfitting W ithout L inearity: N eural N etwork C lassifiers T rained by G radient D escent for N oisy L inear D ata
Spencer Frei, Niladri S Chatterji, and Peter Bartlett. Benign O verfitting W ithout L inearity: N eural N etwork C lassifiers T rained by G radient D escent for N oisy L inear D ata. In Conference on Learning Theory, pages 2668--2703. PMLR, 2022
2022
-
[30]
Benign O verfitting in L inear C lassifiers and L eaky R e LU N etworks from KKT C onditions for M argin M aximization
Spencer Frei, Gal Vardi, Peter Bartlett, and Nathan Srebro. Benign O verfitting in L inear C lassifiers and L eaky R e LU N etworks from KKT C onditions for M argin M aximization. In The Thirty Sixth Annual Conference on Learning Theory, pages 3173--3228. PMLR, 2023
2023
-
[31]
When do N eural N etworks O utperform K ernel M ethods? Advances in Neural Information Processing Systems, 33: 0 14820--14830, 2020
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari. When do N eural N etworks O utperform K ernel M ethods? Advances in Neural Information Processing Systems, 33: 0 14820--14830, 2020
2020
-
[32]
Linearized T wo- L ayers N eural N etworks in H igh D imension
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari. Linearized T wo- L ayers N eural N etworks in H igh D imension. The Annals of Statistics, 49 0 (2): 0 1029--1054, 2021
2021
-
[33]
A D istribution- F ree T heory of N onparametric R egression
L \'a szl \'o Gy \"o rfi, Michael Kohler, Adam Krzyzak, and Harro Walk. A D istribution- F ree T heory of N onparametric R egression . Springer Science & Business Media, 2006
2006
-
[34]
Mind the S pikes: B enign O verfitting of K ernels and N eural N etworks in F ixed D imension
Moritz Haas, David Holzm \"u ller, Ulrike von Luxburg, and Ingo Steinwart. Mind the S pikes: B enign O verfitting of K ernels and N eural N etworks in F ixed D imension. arXiv preprint arXiv:2305.14077, 2023
2023 arXiv
-
[35]
Provable T empered O verfitting of M inimal N ets and T ypical N ets
Itamar Harel, William M Hoza, Gal Vardi, Itay Evron, Nathan Srebro, and Daniel Soudry. Provable T empered O verfitting of M inimal N ets and T ypical N ets. arXiv preprint arXiv:2410.19092, 2024
2024 arXiv
-
[36]
The E lements of S tatistical L earning: D ata M ining, I nference, and P rediction , volume 2
Trevor Hastie, Robert Tibshirani, Jerome H Friedman, and Jerome H Friedman. The E lements of S tatistical L earning: D ata M ining, I nference, and P rediction , volume 2. Springer, 2009
2009
-
[37]
Surprises in H igh- D imensional R idgeless L east S quares I nterpolation
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J Tibshirani. Surprises in H igh- D imensional R idgeless L east S quares I nterpolation. Annals of statistics, 50 0 (2): 0 949, 2022
2022
-
[38]
Using C ontinuity I nduction
Dan Hathaway. Using C ontinuity I nduction. The College Mathematics Journal, 42 0 (3): 0 229--231, 2011
2011
-
[39]
Matrix A nalysis
Roger A Horn and Charles R Johnson. Matrix A nalysis . Cambridge university press, 2013
2013
-
[40]
Neural T angent K ernel: C onvergence and G eneralization in N eural N etworks
Arthur Jacot, Franck Gabriel, and Cl \'e ment Hongler. Neural T angent K ernel: C onvergence and G eneralization in N eural N etworks. Advances in neural information processing systems, 31, 2018
2018
-
[41]
Directional C onvergence and A lignment in D eep L earning
Ziwei Ji and Matus Telgarsky. Directional C onvergence and A lignment in D eep L earning. Advances in Neural Information Processing Systems, 33: 0 17176--17186, 2020
2020
-
[42]
Implicit B ias of G radient D escent for M ean S quared E rror R egression with T wo- L ayer W ide N eural N etworks
Hui Jin and Guido Mont \'u far. Implicit B ias of G radient D escent for M ean S quared E rror R egression with T wo- L ayer W ide N eural N etworks. Journal of Machine Learning Research, 24 0 (137): 0 1--97, 2023
2023
-
[43]
Noisy I nterpolation L earning with S hallow U nivariate R e LU N etworks
Nirmit Joshi, Gal Vardi, and Nathan Srebro. Noisy I nterpolation L earning with S hallow U nivariate R e LU N etworks. In The Twelfth International Conference on Learning Representations, 2024
2024
-
[44]
On the G eneralization P ower of O verfitted T wo- L ayer N eural T angent K ernel M odels
Peizhong Ju, Xiaojun Lin, and Ness Shroff. On the G eneralization P ower of O verfitted T wo- L ayer N eural T angent K ernel M odels. In International Conference on Machine Learning, pages 5137--5147. PMLR, 2021
2021
-
[45]
On the G eneralization P ower of the O verfitted T hree- L ayer N eural T angent K ernel M odel
Peizhong Ju, Xiaojun Lin, and Ness Shroff. On the G eneralization P ower of the O verfitted T hree- L ayer N eural T angent K ernel M odel. Advances in Neural Information Processing Systems, 35: 0 26135--26146, 2022
2022
-
[46]
Uniform C onvergence of I nterpolators: G aussian W idth, N orm B ounds and B enign O verfitting
Frederic Koehler, Lijia Zhou, Danica J Sutherland, and Nathan Srebro. Uniform C onvergence of I nterpolators: G aussian W idth, N orm B ounds and B enign O verfitting. Advances in Neural Information Processing Systems, 34: 0 20657--20668, 2021
2021
-
[47]
From T empered to B enign O verfitting in R e LU N eural N etworks
Guy Kornowski, Gilad Yehudai, and Ohad Shamir. From T empered to B enign O verfitting in R e LU N eural N etworks. arXiv preprint arXiv:2305.15141, 2023
2023 arXiv
-
[48]
Benign O verfitting for T wo- L ayer R e LU N etworks
Yiwen Kou, Zixiang Chen, Yuanzhou Chen, and Quanquan Gu. Benign O verfitting for T wo- L ayer R e LU N etworks. arXiv preprint arXiv:2303.04145, 2023
2023 arXiv
-
[49]
Generalization A bility of W ide N eural N etworks on R
Jianfa Lai, Manyun Xu, Rui Chen, and Qian Lin. Generalization A bility of W ide N eural N etworks on R . arXiv preprint arXiv:2302.05933, 2023
2023 arXiv
-
[50]
Real and F unctional A nalysis , volume 142
Serge Lang. Real and F unctional A nalysis , volume 142. Springer Science & Business Media, 1993
1993
-
[51]
Adaptive E stimation of a Q uadratic F unctional by M odel S election
Beatrice Laurent and Pascal Massart. Adaptive E stimation of a Q uadratic F unctional by M odel S election. Annals of statistics, pages 1302--1338, 2000
2000
-
[52]
A. J. Lee. U- S tatistics: T heory and P ractice , volume 110. CRC Press, Taylor & Francis Group, 1990
1990
-
[53]
Stability and G eneralization A nalysis of G radient M ethods for S hallow N eural N etworks
Yunwen Lei, Rong Jin, and Yiming Ying. Stability and G eneralization A nalysis of G radient M ethods for S hallow N eural N etworks. Advances in Neural Information Processing Systems, 35: 0 38557--38570, 2022
2022
-
[54]
Kernel I nterpolation G eneralizes P oorly
Yicheng Li, Haobo Zhang, and Qian Lin. Kernel I nterpolation G eneralizes P oorly. Biometrika, 111 0 (2): 0 715--722, 2024
2024
-
[55]
Towards an U nderstanding of B enign O verfitting in N eural N etworks
Zhu Li, Zhi-Hua Zhou, and Arthur Gretton. Towards an U nderstanding of B enign O verfitting in N eural N etworks. arXiv preprint arXiv:2106.03212, 2021
2021 arXiv
-
[56]
R idgeless
Tengyuan Liang and Alexander Rakhlin. Just I nterpolate: K ernel “ R idgeless” R egression can G eneralize. The Annals of Statistics, 48 0 (3): 0 1329--1347, 2020
2020
-
[57]
On the M ultiple D escent of M inimum- N orm I nterpolants and R estricted L ower I sometry of K ernels
Tengyuan Liang, Alexander Rakhlin, and Xiyu Zhai. On the M ultiple D escent of M inimum- N orm I nterpolants and R estricted L ower I sometry of K ernels. In Conference on Learning Theory, pages 2683--2711. PMLR, 2020
2020
-
[58]
Benign, T empered, or C atastrophic: T oward a R efined T axonomy of O verfitting
Neil Mallinar, James Simon, Amirhesam Abedsoltan, Parthe Pandit, Misha Belkin, and Preetum Nakkiran. Benign, T empered, or C atastrophic: T oward a R efined T axonomy of O verfitting. Advances in Neural Information Processing Systems, 35: 0 1182--1195, 2022
2022
-
[59]
Overfitting B ehaviour of G aussian K ernel R idgeless R egression: V arying B andwidth or D imensionality
Marko Medvedev, Gal Vardi, and Nathan Srebro. Overfitting B ehaviour of G aussian K ernel R idgeless R egression: V arying B andwidth or D imensionality. arXiv preprint arXiv:2409.03891, 2024
2024 arXiv
-
[60]
The G eneralization E rror of R andom F eatures R egression: P recise A symptotics and the D ouble D escent C urve
Song Mei and Andrea Montanari. The G eneralization E rror of R andom F eatures R egression: P recise A symptotics and the D ouble D escent C urve. Communications on Pure and Applied Mathematics, 75 0 (4): 0 667--766, 2022
2022
-
[61]
A M ean F ield V iew of the L andscape of T wo- L ayers N eural N etworks
Song Mei, Andrea Montanari, and P Nguyen. A M ean F ield V iew of the L andscape of T wo- L ayers N eural N etworks. Proceedings of the National Academy of Sciences, 115 0 (33): 0 E7665--E7671, 2018
2018
-
[62]
Mean-field T heory of T wo- L ayers N eural N etworks: D imension- F ree B ounds and K ernel L imit
Song Mei, Theodor Misiakiewicz, and Andrea Montanari. Mean-field T heory of T wo- L ayers N eural N etworks: D imension- F ree B ounds and K ernel L imit. In Conference on Learning Theory, pages 2388--2464. PMLR, 2019
2019
-
[63]
Universal K ernels
Charles A Micchelli, Yuesheng Xu, and Haizhang Zhang. Universal K ernels. Journal of Machine Learning Research, 7 0 (12), 2006
2006
-
[64]
Foundations of M achine L earning
Mehryar Mohri, Afshin Rostamizadeh, and Ameet Talwalkar. Foundations of M achine L earning . MIT press, 2012
2012
-
[65]
The I nterpolation P hase T ransition in N eural N etworks: M emorization and G eneralization under L azy T raining
Andrea Montanari and Yiqiao Zhong. The I nterpolation P hase T ransition in N eural N etworks: M emorization and G eneralization under L azy T raining. The Annals of Statistics, 50 0 (5): 0 2816--2847, 2022
2022
-
[66]
An elementary analysis of ridge regression with random design
Jaouad Mourtada and Lorenzo Rosasco. An elementary analysis of ridge regression with random design. Comptes Rendus. Math \'e matique , 360 0 (G9): 0 1055--1063, 2022
2022
-
[67]
Analysis of S pherical S ymmetries in E uclidean S paces , volume 129
Claus M \"u ller. Analysis of S pherical S ymmetries in E uclidean S paces , volume 129. Springer Science & Business Media, 1998
1998
-
[68]
Harmless I nterpolation of N oisy D ata in R egression
Vidya Muthukumar, Kailas Vodrahalli, Vignesh Subramanian, and Anant Sahai. Harmless I nterpolation of N oisy D ata in R egression. IEEE Journal on Selected Areas in Information Theory, 1 0 (1): 0 67--83, 2020
2020
-
[69]
Uniform C onvergence may be U nable to E xplain G eneralization in D eep L earning
Vaishnavh Nagarajan and J Zico Kolter. Uniform C onvergence may be U nable to E xplain G eneralization in D eep L earning. Advances in Neural Information Processing Systems, 32, 2019
2019
-
[70]
Deep D ouble D escent: W here B igger M odels and M ore D ata H urt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever. Deep D ouble D escent: W here B igger M odels and M ore D ata H urt. Journal of Statistical Mechanics: Theory and Experiment, 2021 0 (12): 0 124003, 2021
2021
-
[71]
Warwick Nash, Tracy Sellers, Simon Talbot, Andrew Cawthorn, and Wes Ford. Abalone . UCI Machine Learning Repository, 1994. DOI : https://doi.org/10.24432/C55C7W
1994 doi
-
[72]
On the P roof of G lobal C onvergence of G radient D escent for D eep R e LU N etworks with L inear W idths
Quynh Nguyen. On the P roof of G lobal C onvergence of G radient D escent for D eep R e LU N etworks with L inear W idths. In International Conference on Machine Learning, pages 8056--8062. PMLR, 2021
2021
-
[73]
Toward M oderate O verparameterization: G lobal C onvergence G uarantees for T raining S hallow N eural N etworks
Samet Oymak and Mahdi Soltanolkotabi. Toward M oderate O verparameterization: G lobal C onvergence G uarantees for T raining S hallow N eural N etworks. IEEE Journal on Selected Areas in Information Theory, 1 0 (1): 0 84--105, 2020
2020
-
[74]
Regularised L east- S quares R egression with I nfinite- D imensional O utput S pace
Junhyung Park and Krikamol Muandet. Regularised L east- S quares R egression with I nfinite- D imensional O utput S pace. arXiv preprint arXiv:2010.10973, 2020
2010 arXiv
-
[75]
Towards E mpirical P rocess T heory for V ector- V alued F unctions: M etric E ntropy of S mooth F unction C lasses
Junhyung Park and Krikamol Muandet. Towards E mpirical P rocess T heory for V ector- V alued F unctions: M etric E ntropy of S mooth F unction C lasses. In International Conference on Algorithmic Learning Theory, pages 1216--1260. PMLR, 2023
2023
-
[76]
An A pproach to I nequalities for the D istributions of I nfinite- D imensional M artingales
Iosif Pinelis. An A pproach to I nequalities for the D istributions of I nfinite- D imensional M artingales. In Probability in Banach Spaces, 8: Proceedings of the Eighth International Conference, pages 128--134. Springer, 1992
1992
-
[77]
Methods in N onlinear I ntegral E quations
Radu Precup. Methods in N onlinear I ntegral E quations . Springer Science & Business Media, 2002
2002
-
[78]
Consistency of I nterpolation with L aplace K ernels is a H igh- D imensional P henomenon
Alexander Rakhlin and Xiyu Zhai. Consistency of I nterpolation with L aplace K ernels is a H igh- D imensional P henomenon. In Conference on Learning Theory, pages 2595--2623. PMLR, 2019
2019
-
[79]
Matrix A lgebra and its A pplications to S tatistics and E conometrics
Calyampudi Radhakrishna Rao and Mareppalli Bhaskara Rao. Matrix A lgebra and its A pplications to S tatistics and E conometrics . World Scientific, 1998
1998
-
[80]
Improved C onvergence G uarantees for S hallow N eural N etworks
Alexander Razborov. Improved C onvergence G uarantees for S hallow N eural N etworks. arXiv preprint arXiv:2212.02323, 2022
2022 arXiv
-
[81]
Stability & G eneralisation of G radient D escent for S hallow N eural N etworks without the N eural T angent K ernel
Dominic Richards and Ilja Kuzborskij. Stability & G eneralisation of G radient D escent for S hallow N eural N etworks without the N eural T angent K ernel. Advances in Neural Information Processing Systems, 34: 0 8609--8621, 2021
2021
-
[82]
On L earning with I ntegral O perators
Lorenzo Rosasco, Mikhail Belkin, and Ernesto De Vito. On L earning with I ntegral O perators. Journal of Machine Learning Research, 11 0 (2), 2010
2010
-
[83]
Generalization properties of learning with random features
Alessandro Rudi and Lorenzo Rosasco. Generalization properties of learning with random features. Advances in neural information processing systems, 30, 2017
2017
-
[84]
Approximation T heorems of M athematical S tatistics
Robert J Serfling. Approximation T heorems of M athematical S tatistics. Wiley Series in Probability and Statistics, 1980
1980
-
[85]
Understanding M achine L earning: F rom T heory to A lgorithms
Shai Shalev-Shwartz and Shai Ben-David. Understanding M achine L earning: F rom T heory to A lgorithms . Cambridge university press, 2014
2014
-
[86]
Support V ector M achines
Ingo Steinwart and Andreas Christmann. Support V ector M achines . Springer Science & Business Media, 2008
2008
-
[87]
A N on- P arametric R egression V iewpoint: G eneralization of O verparametrized D eep R e LU N etwork under N oisy O bservations
Namjoon Suh, Hyunouk Ko, and Xiaoming Huo. A N on- P arametric R egression V iewpoint: G eneralization of O verparametrized D eep R e LU N etwork under N oisy O bservations. In International Conference on Learning Representations, 2021
2021
-
[88]
User- F riendly T ail B ounds for S ums of R andom M atrices
Joel A Tropp. User- F riendly T ail B ounds for S ums of R andom M atrices. Foundations of computational mathematics, 12: 0 389--434, 2012
2012
-
[89]
Empirical P rocesses in M - E stimation , volume 6
Sara A van de Geer. Empirical P rocesses in M - E stimation , volume 6. Cambridge university press, 2000
2000
-
[90]
On the I mplicit B ias in D eep- L earning A lgorithms
Gal Vardi. On the I mplicit B ias in D eep- L earning A lgorithms. Communications of the ACM, 66 0 (6): 0 86--93, 2023
2023
-
[91]
High- D imensional P robability: A n I ntroduction with A pplications in D ata S cience , volume 47
Roman Vershynin. High- D imensional P robability: A n I ntroduction with A pplications in D ata S cience , volume 47. Cambridge university press, 2018
2018
-
[92]
Benign overfitting in adversarial training of neural networks
Yunjuan Wang, Kaibo Zhang, and Raman Arora. Benign overfitting in adversarial training of neural networks. In Forty-first International Conference on Machine Learning, 2024
2024
-
[93]
Linear O perators in H ilbert S paces , volume 68
Joachim Weidmann. Linear O perators in H ilbert S paces , volume 68. Springer New York, 1980
1980
-
[94]
Precise L earning C urves and H igher- O rder S caling L imits for D ot P roduct K ernel R egression
Lechao Xiao, Hong Hu, Theodor Misiakiewicz, Yue M Lu, and Jeffrey Pennington. Precise L earning C urves and H igher- O rder S caling L imits for D ot P roduct K ernel R egression. In Thirty-sixth Conference on Neural Information Processing Systems (NeurIPS), 2022
2022
-
[95]
Rethinking benign overfitting in two-layer neural networks
Ruichen Xu and Kexin Chen. Rethinking benign overfitting in two-layer neural networks. arXiv preprint arXiv:2502.11893, 2025
2025 arXiv
-
[96]
Benign O verfitting of N on- S mooth N eural N etworks B eyond L azy T raining
Xingyu Xu and Yuantao Gu. Benign O verfitting of N on- S mooth N eural N etworks B eyond L azy T raining. In International Conference on Artificial Intelligence and Statistics, pages 11094--11117. PMLR, 2023
2023
-
[97]
Feature L earning in I nfinite- W idth N eural N etworks
Greg Yang and Edward J Hu. Feature L earning in I nfinite- W idth N eural N etworks. arXiv preprint arXiv:2011.14522, 2020
2011 arXiv
-
[98]
Sobolev norm inconsistency of kernel interpolation
Yunfei Yang. Sobolev norm inconsistency of kernel interpolation. arXiv preprint arXiv:2504.20617, 2025
2025
-
[99]
A U nifying V iew on I mplicit B ias in T raining L inear N eural N etworks
Chulhee Yun, Shankar Krishnan, and Hossein Mobahi. A U nifying V iew on I mplicit B ias in T raining L inear N eural N etworks. arXiv preprint arXiv:2010.02501, 2020
2010 arXiv
-
[100]
A T ype of G eneralization E rror I nduced by I nitialization in D eep N eural N etworks
Yaoyu Zhang, Zhi-Qin John Xu, Tao Luo, and Zheng Ma. A T ype of G eneralization E rror I nduced by I nitialization in D eep N eural N etworks. In Mathematical and Scientific Machine Learning, pages 144--164. PMLR, 2020
2020
-
[101]
An A gnostic V iew on the C ost of O verfitting in ( K ernel) R idge R egression
Lijia Zhou, James B Simon, Gal Vardi, and Nathan Srebro. An A gnostic V iew on the C ost of O verfitting in ( K ernel) R idge R egression. In International Conference on Learning Representations, 2024
2024
-
[102]
Benign O verfitting in D eep N eural N etworks under L azy T raining
Zhenyu Zhu, Fanghui Liu, Grigorios Chrysos, Francesco Locatello, and Volkan Cevher. Benign O verfitting in D eep N eural N etworks under L azy T raining. In International Conference on Machine Learning, pages 43105--43128. PMLR, 2023
2023
-
[103]
Benign O verfitting of C onstant- S tepsize SGD for L inear R egression
Difan Zou, Jingfeng Wu, Vladimir Braverman, Quanquan Gu, and Sham Kakade. Benign O verfitting of C onstant- S tepsize SGD for L inear R egression. In Conference on Learning Theory, pages 4633--4635. PMLR, 2021
2021
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.