Pith. sign in

REVIEW 3 major objections 3 minor 103 references

A Classical View on Benign Overfitting: The Role of Sample Size

T0 review · 3 major / 3 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Benign overfitting can occur in the classical U-shaped regime when sample size and model complexity grow together, the paper proves for kernel ridge regression and trained two-layer ReLU networks.

desk verdict The proof machinery is real, but the ReLU NTK's infinite-dimensional kernel sinks the 'any bounded f*' theorem; major revision needed. read the letter →

arxiv 2505.11621 v1 pith:AON5HGFA submitted 2025-05-16 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords benignoverfittingalmostkernelridgeregressionneuraltangenttwo-layerReLUnetworksgradientflowimplicitregularizationgeneralizationbound
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that 'almost benign overfitting' — a model that fits noisy training data to arbitrarily small error while also generalizing to arbitrarily small excess risk — can occur in the classical U-shaped regime of the risk-versus-complexity curve, provided sample size and model complexity are increased together. Prior work placed benign overfitting outside the classical regime, in the interpolation region, and typically required high-dimensional inputs or structural assumptions on the regression function. The paper proves the claim in two settings: kernel ridge regression, and two-layer ReLU networks trained by gradient flow in the neural-tangent-kernel regime, with no assumptions on the regression function or the noise beyond boundedness. The proof decomposes excess risk into approximation and estimation errors and reads gradient flow as an implicit regularizer, which lets the analysis avoid uniform convergence. The central picture is that the trough of the U-curve moves 'down and to the right' as data accumulate, so the best model at large sample size is both larger and better fitting.

What carries the argument

The load-bearing mechanism is the approximation–estimation decomposition of the excess risk along the gradient-flow trajectory, with gradient flow itself treated as an implicit regularizer. The paper compares an $m$-neuron ReLU network $\hat f_t$ trained on the empirical risk with an oracle network $f_t$ of the same architecture trained on the population risk: the oracle gap $\|f_t-f_*\|_2$ is the approximation error, and the tracking gap $\|\hat f_t-f_t\|_2$ is the estimation error. The implicit regularizer is the stopping time $T_\varepsilon = 2\lambda_\varepsilon^{-1}\log(2/\sqrt\varepsilon)$, where $\lambda_\varepsilon$ is the $L_\varepsilon$-th eigenvalue of the analytical neural tangent kernel operator, chosen so that the component of $f_*$ in the remaining high-frequency eigenspace has norm at most $\sqrt\varepsilon/4$. Three supporting devices carry the proofs: a spectral-norm bound for Hadamard products extended from matrices to integral operators (used for the approximation error), new concentration inequalities for vector-valued U- and V-statistics that allow the estimation-error bound to be made at initialization and then iterated $U_\varepsilon$ times with the factorial $U_\varepsilon!$ absorbing the remainder, and a real-induction argument that keeps the weights near initialization so the minimum eigenvalue of the NTK Gram matrix stays bounded along the whole trajectory.

What would settle it

Evaluate the theorem's dimension–confidence trade-off at the dimensions its experiments use: with $d=7$ (Abalone), Assumption 2(i) requires $\delta\ge 12e^{-7}\approx 0.011$, so at $\delta=0.001$ the theorem's hypotheses are void. Running the Abalone-style experiment at $d=7$ with $\delta$ below that threshold — or checking whether a bounded target whose mass sits mostly on the low eigenvalues of the NTK operator still allows both risks below $\varepsilon$ at fixed $d$ — would show whether the claim survives outside the stated scaling regime.

Watch

Extended reading notes

Core claim

The central discovery, stated on the paper's own terms, is that whether overfitting is benign is decided by the joint scaling of sample size and model complexity, not by model complexity alone. The paper defines benign overfitting as achieving empirical risk at most $\varepsilon$ and excess risk $R(\hat f)-R(f_*) \le \varepsilon$ simultaneously, for every $\varepsilon,\delta>0$ with probability at least $1-\delta$ (Definition 1), deliberately allowing training error that is small but not exactly zero. For kernel ridge regression with the neural tangent kernel, Theorem 6 shows that when the regularization parameter $\gamma$ is small and the sample size $n$ is large relative to $\varepsilon$ and $\delta$ (Assumption 1), the regularized empirical risk minimizer achieves both bounds on a single high-probability event. For two-layer ReLU networks trained by full-batch gradient flow in the NTK regime — where each neuron moves so little that the network behaves like a kernel method — Theorem 11 shows that for any essentially bounded regression function $f_*$, if input dimension, width $m$, and sample size $n$ satisfy Assumptions 2 and 3, then at stopping time $T_\varepsilon = \frac{2}{\lambda_\varepsilon}\log\frac{2}{\sqrt\varepsilon}$ the trained network satisfies both bounds on the same event of probability at least $1-\delta$. The authors claim this as the first generalization result in this setting that assumes nothing about the regression function or the noise beyond boundedness.

Load-bearing premise

The load-bearing premise is a set of scaling relations: the input dimension $d$ must be at least logarithmic in the inverse failure probability ($e^{-d}\le\delta/12$), and the width $m$ and sample size $n$ must be large enough relative to $d$, $\varepsilon$, and the spectral gap $\lambda_\varepsilon$; if these fail, the high-probability minimum-eigenvalue bounds on which every theorem rests are unavailable.

Editorial extensions

If this is right

  • Benign overfitting does not require leaving the classical regime: a practitioner who adds data and model capacity together can push both training error and test error below any desired level, without high-dimensional inputs or a well-specified model class.
  • For two-layer ReLU networks in the NTK regime, generalization to the Bayes-optimal risk is achievable for every bounded regression function, with no RKHS-of-NTK assumption and no noise model beyond boundedness.
  • The KRR result is not tied to the specific kernel: the same proof works for any bounded kernel whose Gram matrix has the required minimum-eigenvalue lower bound and whose RKHS is dense in $L^2(\rho_{d-1})$.
  • Because the analysis deliberately avoids uniform convergence over the parameter space, the bounds do not deteriorate as the width $m$ grows; more parameters do not hurt the generalization guarantee.
  • The experiments corroborate the predicted 'down and to the right' shift: on Abalone, Wine, and synthetic sphere data, the point where excess risk crosses below empirical risk moves later in training and lower in value as $n$ increases, and both risks at the crossing fall.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • I read the Section 2 narrative as a reconciliation claim the paper states but does not fully prove: earlier interpolation-regime results hold the model fixed on the upward slope of a U-curve drawn at a fixed sample size, whereas this paper moves along the moving trough; a head-to-head comparison of the two regimes under identical scaling is a natural next step.
  • The vector-valued U-/V-statistic concentration bounds and the iterated-remainder trick are separable tools that should transfer to other settings needing high-probability control of two coupled trajectories without uniform convergence, such as stochastic gradient descent or feature-learning dynamics.
  • A testable extension suggested by the assumptions: the paper notes that smooth activations would simplify the proofs, so if the mechanism is really the approximation–estimation trade-off rather than ReLU specifics, the benign-overfitting window should widen (smaller required $n$ and $m$) for smooth activations.
  • The conclusion flags that only upper bounds are provided and matching lower bounds remain open; I draw the consequence that the 'low-dimensional inputs' claim should be read as conditioned on $d$ growing with $\log(1/\delta)$, given Assumption 2(i)'s requirement $e^{-d}\le\delta/12$.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper proposes reinterpreting benign overfitting through 'almost benign overfitting,' where empirical risk and excess risk can both be made arbitrarily small by jointly increasing sample size and model complexity. The authors prove such guarantees for two case studies: kernel ridge regression with the ReLU neural tangent kernel, and least-squares regression with a two-layer ReLU network trained by gradient flow in the NTK regime. The main advertised contribution is that these results hold under essentially no assumptions on the regression function beyond boundedness, and in low input dimension, using a novel approximation/estimation decomposition of the excess risk that avoids uniform convergence. The paper also provides experiments on synthetic, Abalone, and Wine data that show risk curves consistent with the hypothesis.

Significance. If the main theorem were correct, the paper would represent a substantial conceptual shift: benign overfitting would no longer require high dimensionality, kernel eigenvalue decay conditions, or interpolation, and the neural-network result would be the first generalization guarantee of this type for arbitrary bounded regression functions. The proof machinery is genuinely elaborate and contains independently interesting tools, including a Hadamard-product bound for integral operators and concentration results for vector-valued U- and V-statistics. The paper is also honest about its limitations, such as only providing upper bounds. However, these strengths do not compensate for the central technical gaps described below, which affect the core claims of the paper.

major comments (3)
  1. [Appendix D.2.3 and Eq. (4.1)] The ReLU NTK on the sphere has an infinite-dimensional kernel, and this invalidates the claimed arbitrary-bounded-f* result. The paper's own spectral calculation in Appendix D.2.3 gives mu_h = 0 for every odd h >= 3, while positive eigenvalues exist for h = 1, h = 2, and infinitely many even h >= 4. Consequently, the operator H has both an infinite-dimensional positive eigenspace and an infinite-dimensional zero eigenspace, so the statement that the eigenfunctions form an orthonormal basis with lambda_1 >= lambda_2 >= ... and lambda_l -> 0 'from above' is not correct as an enumeration with nonzero eigenvalues only. More importantly, for a bounded f* with a nonzero component on an odd spherical harmonic of degree at least 3, for example the normalized Y_3, Eq. (4.1) cannot be satisfied with lambda_epsilon > 0: if Y_3 is left in the tail, the tail norm is 1 for every finite L, and if Y_3 is placed among the top L eigenfunctions, the corresponding eigenvalue is 0, so lambda_epsilon = 0 and T_epsilon in Eq. (4.2) is undefined. Hence Theorem 8, and therefore Theorems 10 and 11, do not hold for arbitrary bounded f* as claimed in the abstract and in Section 4.2.
  2. [Section 3, paragraph 'By the denseness of H in L2'] The KRR section relies on the assertion that the RKHS of the ReLU NTK is dense in L2(rho_{d-1}), writing 'By the denseness of H in L2(rho), there is an f_epsilon in H such that ||f* - f_epsilon||_2^2 <= epsilon/8.' This assertion is false for the ReLU NTK: as the spectral calculation in Appendix D.2.3 shows, the kernel misses all odd spherical harmonics of degree at least 3, so functions such as Y_3 are orthogonal to the RKHS. Therefore the approximation argument behind Theorem 3 and the benign-overfitting conclusion in Theorem 6 only hold for regression functions lying in a proper subspace, not for arbitrary bounded f*. This is a load-bearing error in the paper's central claim of no assumptions on f*.
  3. [Definition 1 vs. Assumptions 1(i) and 2(i)] Definition 1 requires that for every epsilon, delta > 0 there exists a sample size n such that, with probability at least 1 - delta, both risks are at most epsilon. For a fixed problem, the input dimension d is fixed, but Assumptions 1(i) and 2(i) require e^{-d} <= delta/4 and e^{-d} <= delta/12, respectively, i.e., d >= Omega(log(1/delta)). Thus, for any fixed low-dimensional problem, the stated theorems cannot cover arbitrarily small delta, and the results do not achieve Definition 1. The paper's claim that the results 'hold on low-dimensional inputs' is therefore only true in the limited sense that d is not required to grow with n or m, but d must still grow without bound as delta tends to zero. This is a mismatch between the theorem statements and the formal definition of the object they claim to establish.
minor comments (3)
  1. [Title of Section 4 and Figure captions] There is a typo in the Section 4 heading: 'Bengin Overfitting' should read 'Benign Overfitting.' The same typo appears in the table of contents.
  2. [Appendix D.2.3, paragraph after spectral decomposition] The phrase 'lambda_l -> 0 from above' is misleading and should be replaced by a correct treatment of the zero eigenspace; the current wording suggests that all eigenvalues are positive, which contradicts the formula mu_h = 0 for odd h >= 3 appearing a few lines later.
  3. [Section 4.3 and Appendix D.8] The experiments use gradient descent with learning rate 0.1, while the theory is developed for gradient flow; the paper should state explicitly that discretization is not covered by the proofs and that the experiments are only heuristic support.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the derivation is self-contained and the parameters are chosen to satisfy stated inequalities rather than fitted to the target risks.

full rationale

The paper's central claims are derived from stated assumptions without the target bound being used as an input. In the KRR section, f_gamma and hat f_gamma are explicit regularized minimizers; epsilon, gamma, and n are chosen sequentially so that the approximation and estimation bounds each contribute at most sqrt(epsilon)/2. No constant is fitted to the empirical or excess risk. In the neural-network section, L_epsilon, lambda_epsilon, T_epsilon, and U_epsilon are analysis quantities: L_epsilon truncates f* by tail norm, T_epsilon is obtained by solving the independently derived exponential-decay inequality exp(-lambda_epsilon t/2) <= sqrt(epsilon)/2, and U_epsilon is a factorial remainder device. These choices are not predictions; they are proof parameters. The high-probability event E3 is built by union bounds from concentration lemmas, and Theorems 7, 8, and 9 hold on that same event by deterministic arguments. The citations to Park and Muandet (2020, 2023) are auxiliary standard closed-form and concentration results; they are not load-bearing self-citations and do not presuppose benign overfitting. The manuscript's own limitation statement (upper bounds only) and the possible spectral obstruction from the zero eigenvalues for odd h >= 3 in Appendix D.2.3 concern correctness or assumption-satisfiability, not circularity: the proof would still derive the stated bound whenever the assumptions hold. Therefore no step reduces by construction to its inputs.

Assumptions & free parameters 6 free parameters · 10 assumptions · 0 invented entities

The central claim rests on standard analysis tools (real induction, spectral theory, concentration inequalities), on the domain assumptions of spherical inputs and bounded labels, and on several modeling choices specific to this paper: antisymmetric initialization, fixed output layer, the NTK (lazy training) regime, and the relaxation from exact interpolation to epsilon-tolerance. No new physical entities are introduced.

free parameters (6)
  • gamma (KRR regularization)
    Chosen to satisfy Assumption 1: roughly gamma <= O(sqrt(eps)/d) for overfitting and gamma <= eps/(8||f_eps||_H^2) for approximation; smaller gamma increases estimation complexity n.
  • L_epsilon (spectral truncation level)
    Smallest index such that the tail norm of f* beyond L_epsilon is <= sqrt(eps)/4; depends on the unknown target f*.
  • lambda_epsilon = lambda_{L_epsilon}
    The L_epsilon-th eigenvalue of the NTK operator; determines training time T_epsilon.
  • T_epsilon (gradient flow time) = 2/lambda_epsilon log(2/sqrt(eps))
    Chosen so approximation error decays below sqrt(eps)/2; scales inversely with lambda_epsilon.
  • U_epsilon = smallest U with (1/U!)(8T_epsilon/d)^U <= sqrt(eps)/14
    Number of iterated integrations in the estimation error bound; controls the size of the remaining integral.
  • Width m and sample size n
    Chosen to satisfy Assumptions 2 and 3, which require m >= poly(d, 1/lambda_epsilon, log(1/delta)) and n >= poly(1/eps, log(1/delta)).
assumptions (10)
  • standard math Real induction (Hathaway 2011, Clark 2019)
    Used to turn local bounds on the gradient flow into global bounds for all t in [0,T].
  • standard math Spectral theory for compact self-adjoint operators
    Used to decompose the NTK operator and define the eigenbasis for approximation error.
  • standard math Concentration inequalities: Hoeffding, McDiarmid, Matrix Chernoff, vector-valued Hoeffding (Pinelis)
    Basis for the high-probability events E1, E2, E3.
  • domain assumption x is uniform on the sphere S^{d-1}
    Simplifies the NTK spectrum and isotropy; experiments use real data that violate this.
  • domain assumption |y| <= 1 almost surely, hence |f*| <= 1 and ||f*||_2 <= 1
    Used in every risk bound; no other assumption on the regression function or noise.
  • domain assumption The NTK kernel RKHS is dense in L2(rho)
    Needed to ensure the approximation error can be made small by some f_eps in H (KRR section).
  • ad hoc to paper Antisymmetric initialization makes the network output exactly zero at initialization
    Gives a clean starting point for population and empirical gradient flows.
  • ad hoc to paper Output layer weights are fixed random +/-1
    Standard in NTK analysis; keeps the gradient in a tractable form.
  • ad hoc to paper NTK (lazy training) regime with m sufficiently large
    Assumptions 3(i)-(iv) enforce that weights move much less than the scale at which the kernel changes.
  • ad hoc to paper Relaxation to epsilon-tolerance in Definition 1 ('almost benign overfitting')
    The paper does not prove exact interpolation with near-optimal excess risk; it proves epsilon-tolerance, which makes the result easier.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Classical View on Benign Overfitting: The Role of Sample Size." pith.science (2026). https://pith.science/paper/AON5HGFA

@misc{pith2026250511621,
  author       = {Pith},
  title        = {Pith review of: A Classical View on Benign Overfitting: The Role of Sample Size},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/AON5HGFA}},
  note         = {Machine review of arXiv:2505.11621}
}
read the original abstract

Benign overfitting is a phenomenon in machine learning where a model perfectly fits (interpolates) the training data, including noisy examples, yet still generalizes well to unseen data. Understanding this phenomenon has attracted considerable attention in recent years. In this work, we introduce a conceptual shift, by focusing on almost benign overfitting, where models simultaneously achieve both arbitrarily small training and test errors. This behavior is characteristic of neural networks, which often achieve low (but non-zero) training error while still generalizing well. We hypothesize that this almost benign overfitting can emerge even in classical regimes, by analyzing how the interaction between sample size and model complexity enables larger models to achieve both good training fit but still approach Bayes-optimal generalization. We substantiate this hypothesis with theoretical evidence from two case studies: (i) kernel ridge regression, and (ii) least-squares regression using a two-layer fully connected ReLU neural network trained via gradient flow. In both cases, we overcome the strong assumptions often required in prior work on benign overfitting. Our results on neural networks also provide the first generalization result in this setting that does not rely on any assumptions about the underlying regression function or noise, beyond boundedness. Our analysis introduces a novel proof technique based on decomposing the excess risk into estimation and approximation errors, interpreting gradient flow as an implicit regularizer, that helps avoid uniform convergence traps. This analysis idea could be of independent interest.

Figures

Figures reproduced from arXiv: 2505.11621 by the authors.

Figure 1
Figure 1. Dashed and solid lines show empirical and excess risk, respectively. On plots (b), (c) and (d), [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Risk vs. model complexity plot for Abalone dataset with Gaussian noise (mean-zero, std. dev [PITH_FULL_IMAGE:figures/full_fig_p013_2.png] view at source ↗
Figure 3
Figure 3. In the third picture, the shaded region represents [PITH_FULL_IMAGE:figures/full_fig_p045_3.png] view at source ↗
Figures from the paper (4 more)
Figure 5
Figure 5. Figure 5: Synthetic Data Experiment: The av￾erage iteration at which the excess risk crosses and stays over the empirical evaluated over 10 runs with different random initializations to the neural network. The bars indicate the standard deviation on the iteration number. Note th…
Figure 6
Figure 6. Figure 6: Abalone Data Experiment: Results with varying noise levels. The figure (b) is duplicated from [PITH_FULL_IMAGE:figures/full_fig_p070_6.png]
Figure 7
Figure 7. Figure 7: Abalone Data Experiment: The average iteration at which the excess risk crosses and stays [PITH_FULL_IMAGE:figures/full_fig_p071_7.png]
Figure 8
Figure 8. Figure 8: Wine Data Experiment: Risk vs. model complexity plot with varying sample size [PITH_FULL_IMAGE:figures/full_fig_p071_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

103 extracted references · 69 canonical work pages

  1. [1]

    The N eural T angent K ernel in H igh D imensions: T riple D escent and a M ulti- S cale T heory of G eneralization

    Ben Adlam and Jeffrey Pennington. The N eural T angent K ernel in H igh D imensions: T riple D escent and a M ulti- S cale T heory of G eneralization. In International Conference on Machine Learning, pages 74--84. PMLR, 2020

  2. [2]

    Stefan Aeberhard and M. Forina. Wine . UCI Machine Learning Repository, 1992. DOI : https://doi.org/10.24432/C5PC7J

  3. [3]

    Learning and G eneralization in O verparameterized N eural N etworks, G oing B eyond T wo L ayers

    Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang. Learning and G eneralization in O verparameterized N eural N etworks, G oing B eyond T wo L ayers. Advances in neural information processing systems, 32, 2019 a

  4. [4]

    A C onvergence T heory for D eep L earning via O ver- P arameterization

    Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song. A C onvergence T heory for D eep L earning via O ver- P arameterization. In International conference on machine learning, pages 242--252. PMLR, 2019 b

  5. [5]

    Fine- G rained A nalysis of O ptimization and G eneralization for O verparameterized T wo- L ayer N eural N etworks

    Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li, and Ruosong Wang. Fine- G rained A nalysis of O ptimization and G eneralization for O verparameterized T wo- L ayer N eural N etworks. In International Conference on Machine Learning, pages 322--332. PMLR, 2019

  6. [6]

    Sharp E stimates for E igenvalues of I ntegral O perators G enerated by D ot P roduct K ernels on the S phere

    Douglas Azevedo and Valdir Antonio Menegatto. Sharp E stimates for E igenvalues of I ntegral O perators G enerated by D ot P roduct K ernels on the S phere. Journal of Approximation Theory, 177: 0 57--68, 2014

  7. [7]

    On the I mplicit B ias of I nitialization S hape: B eyond I nfinitesimal M irror D escent

    Shahar Azulay, Edward Moroshko, Mor Shpigel Nacson, Blake E Woodworth, Nathan Srebro, Amir Globerson, and Daniel Soudry. On the I mplicit B ias of I nitialization S hape: B eyond I nfinitesimal M irror D escent. In International Conference on Machine Learning, pages 468--477. PMLR, 2021

  8. [8]

    Benign O verfitting in L inear R egression

    Peter L Bartlett, Philip M Long, G \'a bor Lugosi, and Alexander Tsigler. Benign O verfitting in L inear R egression. Proceedings of the National Academy of Sciences, 117 0 (48): 0 30063--30070, 2020

Show all 103 references
  1. [9]

    Deep L earning: A S tatistical V iewpoint

    Peter L Bartlett, Andrea Montanari, and Alexander Rakhlin. Deep L earning: A S tatistical V iewpoint. Acta numerica, 30: 0 87--201, 2021

  2. [10]

    Generalization in K ernel R egression U nder R ealistic A ssumptions

    Daniel Barzilai and Ohad Shamir. Generalization in K ernel R egression U nder R ealistic A ssumptions. In Forty-first International Conference on Machine Learning, 2024

  3. [11]

    On the I nconsistency of K ernel R idgeless R egression in F ixed D imensions

    Daniel Beaglehole, Mikhail Belkin, and Parthe Pandit. On the I nconsistency of K ernel R idgeless R egression in F ixed D imensions. SIAM Journal on Mathematics of Data Science, 5 0 (4): 0 854--872, 2023

  4. [12]

    Reconciling M odern M achine- L earning P ractice and the C lassical B ias-- V ariance T rade- O ff

    Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal. Reconciling M odern M achine- L earning P ractice and the C lassical B ias-- V ariance T rade- O ff. Proceedings of the National Academy of Sciences, 116 0 (32): 0 15849--15854, 2019

  5. [13]

    Reproducing K ernel H ilbert S paces in P robability and S tatistics

    Alain Berlinet and Christine Thomas-Agnan. Reproducing K ernel H ilbert S paces in P robability and S tatistics . Springer Science & Business Media, 2004

  6. [14]

    On the I nductive B ias of N eural T angent K ernels

    Alberto Bietti and Julien Mairal. On the I nductive B ias of N eural T angent K ernels. Advances in Neural Information Processing Systems, 32, 2019

  7. [15]

    Implicit B ias of MSE G radient O ptimization in U nderparameterized N eural N etworks

    Benjamin Bowman and Guido Montufar. Implicit B ias of MSE G radient O ptimization in U nderparameterized N eural N etworks. In International Conference on Learning Representations, 2021

  8. [16]

    Spectral B ias O utside T he T raining S et for D eep N etworks in the K ernel R egime

    Benjamin Bowman and Guido F Montufar. Spectral B ias O utside T he T raining S et for D eep N etworks in the K ernel R egime. Advances in Neural Information Processing Systems, 35: 0 30362--30377, 2022

  9. [17]

    Kernel I nterpolation in S obolev S paces is not C onsistent in L ow D imensions

    Simon Buchholz. Kernel I nterpolation in S obolev S paces is not C onsistent in L ow D imensions. In Conference on Learning Theory, pages 3410--3440. PMLR, 2022

  10. [18]

    Towards understanding the spectral bias of deep learning

    Yuan Cao, Zhiying Fang, Yue Wu, Ding-Xuan Zhou, and Quanquan Gu. Towards understanding the spectral bias of deep learning. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, pages 2205--2211. International Joint Conferences on Artificial...

  11. [19]

    Benign O verfitting in T wo- L ayer C onvolutional N eural N etworks

    Yuan Cao, Zixiang Chen, Misha Belkin, and Quanquan Gu. Benign O verfitting in T wo- L ayer C onvolutional N eural N etworks. Advances in neural information processing systems, 35: 0 25237--25250, 2022

  12. [20]

    Optimal R ates for the R egularized L east- S quares A lgorithm

    Andrea Caponnetto and Ernesto De Vito. Optimal R ates for the R egularized L east- S quares A lgorithm. Foundations of Computational Mathematics, 7: 0 331--368, 2007

  13. [21]

    Characterizing O verfitting in K ernel R idgeless R egression T hrough the E igenspectrum

    Tin Sum Cheng, Aurelien Lucchi, Anastasis Kratsios, and David Belius. Characterizing O verfitting in K ernel R idgeless R egression T hrough the E igenspectrum. arXiv preprint arXiv:2402.01297, 2024

  14. [22]

    On the R obustness of the M inimim l2 I nterpolator

    Geoffrey Chinot and Matthieu Lerasle. On the R obustness of the M inimim l2 I nterpolator. Bernoulli, 2022

  15. [23]

    On the G lobal C onvergence of G radient D escent for O ver- P arameterized M odels using O ptimal T ransport

    Lenaic Chizat and Francis Bach. On the G lobal C onvergence of G radient D escent for O ver- P arameterized M odels using O ptimal T ransport. Advances in neural information processing systems, 31, 2018

  16. [24]

    The I nstructor’s G uide to R eal I nduction

    Pete L Clark. The I nstructor’s G uide to R eal I nduction. Mathematics Magazine, 92 0 (2): 0 136--150, 2019

  17. [25]

    A U - T urn on D ouble D escent: R ethinking P arameter C ounting in S tatistical L earning

    Alicia Curth, Alan Jeffares, and Mihaela van der Schaar. A U - T urn on D ouble D escent: R ethinking P arameter C ounting in S tatistical L earning. In Advances in Neural Information Processing Systems, volume 36, 2023

  18. [26]

    Gradient D escent F inds G lobal M inima of D eep N eural N etworks

    Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai. Gradient D escent F inds G lobal M inima of D eep N eural N etworks. In International conference on machine learning, pages 1675--1685. PMLR, 2019 a

  19. [27]

    Gradient D escent P rovably O ptimizes O ver- P arameterized N eural N etworks

    Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh. Gradient D escent P rovably O ptimizes O ver- P arameterized N eural N etworks. In International Conference on Learning Representations, 2019 b

  20. [28]

    A C omparative A nalysis of O ptimization and G eneralization P roperties of T wo- L ayer N eural N etwork and R andom F eature M odels under G radient D escent D ynamics

    Weinan E, Chao Ma, and Lei Wu. A C omparative A nalysis of O ptimization and G eneralization P roperties of T wo- L ayer N eural N etwork and R andom F eature M odels under G radient D escent D ynamics. Sci. China Math, 2019

  21. [29]

    Benign O verfitting W ithout L inearity: N eural N etwork C lassifiers T rained by G radient D escent for N oisy L inear D ata

    Spencer Frei, Niladri S Chatterji, and Peter Bartlett. Benign O verfitting W ithout L inearity: N eural N etwork C lassifiers T rained by G radient D escent for N oisy L inear D ata. In Conference on Learning Theory, pages 2668--2703. PMLR, 2022

  22. [30]

    Benign O verfitting in L inear C lassifiers and L eaky R e LU N etworks from KKT C onditions for M argin M aximization

    Spencer Frei, Gal Vardi, Peter Bartlett, and Nathan Srebro. Benign O verfitting in L inear C lassifiers and L eaky R e LU N etworks from KKT C onditions for M argin M aximization. In The Thirty Sixth Annual Conference on Learning Theory, pages 3173--3228. PMLR, 2023

  23. [31]

    When do N eural N etworks O utperform K ernel M ethods? Advances in Neural Information Processing Systems, 33: 0 14820--14830, 2020

    Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari. When do N eural N etworks O utperform K ernel M ethods? Advances in Neural Information Processing Systems, 33: 0 14820--14830, 2020

  24. [32]

    Linearized T wo- L ayers N eural N etworks in H igh D imension

    Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari. Linearized T wo- L ayers N eural N etworks in H igh D imension. The Annals of Statistics, 49 0 (2): 0 1029--1054, 2021

  25. [33]

    A D istribution- F ree T heory of N onparametric R egression

    L \'a szl \'o Gy \"o rfi, Michael Kohler, Adam Krzyzak, and Harro Walk. A D istribution- F ree T heory of N onparametric R egression . Springer Science & Business Media, 2006

  26. [34]

    Mind the S pikes: B enign O verfitting of K ernels and N eural N etworks in F ixed D imension

    Moritz Haas, David Holzm \"u ller, Ulrike von Luxburg, and Ingo Steinwart. Mind the S pikes: B enign O verfitting of K ernels and N eural N etworks in F ixed D imension. arXiv preprint arXiv:2305.14077, 2023

  27. [35]

    Provable T empered O verfitting of M inimal N ets and T ypical N ets

    Itamar Harel, William M Hoza, Gal Vardi, Itay Evron, Nathan Srebro, and Daniel Soudry. Provable T empered O verfitting of M inimal N ets and T ypical N ets. arXiv preprint arXiv:2410.19092, 2024

  28. [36]

    The E lements of S tatistical L earning: D ata M ining, I nference, and P rediction , volume 2

    Trevor Hastie, Robert Tibshirani, Jerome H Friedman, and Jerome H Friedman. The E lements of S tatistical L earning: D ata M ining, I nference, and P rediction , volume 2. Springer, 2009

  29. [37]

    Surprises in H igh- D imensional R idgeless L east S quares I nterpolation

    Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J Tibshirani. Surprises in H igh- D imensional R idgeless L east S quares I nterpolation. Annals of statistics, 50 0 (2): 0 949, 2022

  30. [38]

    Using C ontinuity I nduction

    Dan Hathaway. Using C ontinuity I nduction. The College Mathematics Journal, 42 0 (3): 0 229--231, 2011

  31. [39]

    Matrix A nalysis

    Roger A Horn and Charles R Johnson. Matrix A nalysis . Cambridge university press, 2013

  32. [40]

    Neural T angent K ernel: C onvergence and G eneralization in N eural N etworks

    Arthur Jacot, Franck Gabriel, and Cl \'e ment Hongler. Neural T angent K ernel: C onvergence and G eneralization in N eural N etworks. Advances in neural information processing systems, 31, 2018

  33. [41]

    Directional C onvergence and A lignment in D eep L earning

    Ziwei Ji and Matus Telgarsky. Directional C onvergence and A lignment in D eep L earning. Advances in Neural Information Processing Systems, 33: 0 17176--17186, 2020

  34. [42]

    Implicit B ias of G radient D escent for M ean S quared E rror R egression with T wo- L ayer W ide N eural N etworks

    Hui Jin and Guido Mont \'u far. Implicit B ias of G radient D escent for M ean S quared E rror R egression with T wo- L ayer W ide N eural N etworks. Journal of Machine Learning Research, 24 0 (137): 0 1--97, 2023

  35. [43]

    Noisy I nterpolation L earning with S hallow U nivariate R e LU N etworks

    Nirmit Joshi, Gal Vardi, and Nathan Srebro. Noisy I nterpolation L earning with S hallow U nivariate R e LU N etworks. In The Twelfth International Conference on Learning Representations, 2024

  36. [44]

    On the G eneralization P ower of O verfitted T wo- L ayer N eural T angent K ernel M odels

    Peizhong Ju, Xiaojun Lin, and Ness Shroff. On the G eneralization P ower of O verfitted T wo- L ayer N eural T angent K ernel M odels. In International Conference on Machine Learning, pages 5137--5147. PMLR, 2021

  37. [45]

    On the G eneralization P ower of the O verfitted T hree- L ayer N eural T angent K ernel M odel

    Peizhong Ju, Xiaojun Lin, and Ness Shroff. On the G eneralization P ower of the O verfitted T hree- L ayer N eural T angent K ernel M odel. Advances in Neural Information Processing Systems, 35: 0 26135--26146, 2022

  38. [46]

    Uniform C onvergence of I nterpolators: G aussian W idth, N orm B ounds and B enign O verfitting

    Frederic Koehler, Lijia Zhou, Danica J Sutherland, and Nathan Srebro. Uniform C onvergence of I nterpolators: G aussian W idth, N orm B ounds and B enign O verfitting. Advances in Neural Information Processing Systems, 34: 0 20657--20668, 2021

  39. [47]

    From T empered to B enign O verfitting in R e LU N eural N etworks

    Guy Kornowski, Gilad Yehudai, and Ohad Shamir. From T empered to B enign O verfitting in R e LU N eural N etworks. arXiv preprint arXiv:2305.15141, 2023

  40. [48]

    Benign O verfitting for T wo- L ayer R e LU N etworks

    Yiwen Kou, Zixiang Chen, Yuanzhou Chen, and Quanquan Gu. Benign O verfitting for T wo- L ayer R e LU N etworks. arXiv preprint arXiv:2303.04145, 2023

  41. [49]

    Generalization A bility of W ide N eural N etworks on R

    Jianfa Lai, Manyun Xu, Rui Chen, and Qian Lin. Generalization A bility of W ide N eural N etworks on R . arXiv preprint arXiv:2302.05933, 2023

  42. [50]

    Real and F unctional A nalysis , volume 142

    Serge Lang. Real and F unctional A nalysis , volume 142. Springer Science & Business Media, 1993

  43. [51]

    Adaptive E stimation of a Q uadratic F unctional by M odel S election

    Beatrice Laurent and Pascal Massart. Adaptive E stimation of a Q uadratic F unctional by M odel S election. Annals of statistics, pages 1302--1338, 2000

  44. [52]

    A. J. Lee. U- S tatistics: T heory and P ractice , volume 110. CRC Press, Taylor & Francis Group, 1990

  45. [53]

    Stability and G eneralization A nalysis of G radient M ethods for S hallow N eural N etworks

    Yunwen Lei, Rong Jin, and Yiming Ying. Stability and G eneralization A nalysis of G radient M ethods for S hallow N eural N etworks. Advances in Neural Information Processing Systems, 35: 0 38557--38570, 2022

  46. [54]

    Kernel I nterpolation G eneralizes P oorly

    Yicheng Li, Haobo Zhang, and Qian Lin. Kernel I nterpolation G eneralizes P oorly. Biometrika, 111 0 (2): 0 715--722, 2024

  47. [55]

    Towards an U nderstanding of B enign O verfitting in N eural N etworks

    Zhu Li, Zhi-Hua Zhou, and Arthur Gretton. Towards an U nderstanding of B enign O verfitting in N eural N etworks. arXiv preprint arXiv:2106.03212, 2021

  48. [56]

    R idgeless

    Tengyuan Liang and Alexander Rakhlin. Just I nterpolate: K ernel “ R idgeless” R egression can G eneralize. The Annals of Statistics, 48 0 (3): 0 1329--1347, 2020

  49. [57]

    On the M ultiple D escent of M inimum- N orm I nterpolants and R estricted L ower I sometry of K ernels

    Tengyuan Liang, Alexander Rakhlin, and Xiyu Zhai. On the M ultiple D escent of M inimum- N orm I nterpolants and R estricted L ower I sometry of K ernels. In Conference on Learning Theory, pages 2683--2711. PMLR, 2020

  50. [58]

    Benign, T empered, or C atastrophic: T oward a R efined T axonomy of O verfitting

    Neil Mallinar, James Simon, Amirhesam Abedsoltan, Parthe Pandit, Misha Belkin, and Preetum Nakkiran. Benign, T empered, or C atastrophic: T oward a R efined T axonomy of O verfitting. Advances in Neural Information Processing Systems, 35: 0 1182--1195, 2022

  51. [59]

    Overfitting B ehaviour of G aussian K ernel R idgeless R egression: V arying B andwidth or D imensionality

    Marko Medvedev, Gal Vardi, and Nathan Srebro. Overfitting B ehaviour of G aussian K ernel R idgeless R egression: V arying B andwidth or D imensionality. arXiv preprint arXiv:2409.03891, 2024

  52. [60]

    The G eneralization E rror of R andom F eatures R egression: P recise A symptotics and the D ouble D escent C urve

    Song Mei and Andrea Montanari. The G eneralization E rror of R andom F eatures R egression: P recise A symptotics and the D ouble D escent C urve. Communications on Pure and Applied Mathematics, 75 0 (4): 0 667--766, 2022

  53. [61]

    A M ean F ield V iew of the L andscape of T wo- L ayers N eural N etworks

    Song Mei, Andrea Montanari, and P Nguyen. A M ean F ield V iew of the L andscape of T wo- L ayers N eural N etworks. Proceedings of the National Academy of Sciences, 115 0 (33): 0 E7665--E7671, 2018

  54. [62]

    Mean-field T heory of T wo- L ayers N eural N etworks: D imension- F ree B ounds and K ernel L imit

    Song Mei, Theodor Misiakiewicz, and Andrea Montanari. Mean-field T heory of T wo- L ayers N eural N etworks: D imension- F ree B ounds and K ernel L imit. In Conference on Learning Theory, pages 2388--2464. PMLR, 2019

  55. [63]

    Universal K ernels

    Charles A Micchelli, Yuesheng Xu, and Haizhang Zhang. Universal K ernels. Journal of Machine Learning Research, 7 0 (12), 2006

  56. [64]

    Foundations of M achine L earning

    Mehryar Mohri, Afshin Rostamizadeh, and Ameet Talwalkar. Foundations of M achine L earning . MIT press, 2012

  57. [65]

    The I nterpolation P hase T ransition in N eural N etworks: M emorization and G eneralization under L azy T raining

    Andrea Montanari and Yiqiao Zhong. The I nterpolation P hase T ransition in N eural N etworks: M emorization and G eneralization under L azy T raining. The Annals of Statistics, 50 0 (5): 0 2816--2847, 2022

  58. [66]

    An elementary analysis of ridge regression with random design

    Jaouad Mourtada and Lorenzo Rosasco. An elementary analysis of ridge regression with random design. Comptes Rendus. Math \'e matique , 360 0 (G9): 0 1055--1063, 2022

  59. [67]

    Analysis of S pherical S ymmetries in E uclidean S paces , volume 129

    Claus M \"u ller. Analysis of S pherical S ymmetries in E uclidean S paces , volume 129. Springer Science & Business Media, 1998

  60. [68]

    Harmless I nterpolation of N oisy D ata in R egression

    Vidya Muthukumar, Kailas Vodrahalli, Vignesh Subramanian, and Anant Sahai. Harmless I nterpolation of N oisy D ata in R egression. IEEE Journal on Selected Areas in Information Theory, 1 0 (1): 0 67--83, 2020

  61. [69]

    Uniform C onvergence may be U nable to E xplain G eneralization in D eep L earning

    Vaishnavh Nagarajan and J Zico Kolter. Uniform C onvergence may be U nable to E xplain G eneralization in D eep L earning. Advances in Neural Information Processing Systems, 32, 2019

  62. [70]

    Deep D ouble D escent: W here B igger M odels and M ore D ata H urt

    Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever. Deep D ouble D escent: W here B igger M odels and M ore D ata H urt. Journal of Statistical Mechanics: Theory and Experiment, 2021 0 (12): 0 124003, 2021

  63. [71]

    Warwick Nash, Tracy Sellers, Simon Talbot, Andrew Cawthorn, and Wes Ford. Abalone . UCI Machine Learning Repository, 1994. DOI : https://doi.org/10.24432/C55C7W

  64. [72]

    On the P roof of G lobal C onvergence of G radient D escent for D eep R e LU N etworks with L inear W idths

    Quynh Nguyen. On the P roof of G lobal C onvergence of G radient D escent for D eep R e LU N etworks with L inear W idths. In International Conference on Machine Learning, pages 8056--8062. PMLR, 2021

  65. [73]

    Toward M oderate O verparameterization: G lobal C onvergence G uarantees for T raining S hallow N eural N etworks

    Samet Oymak and Mahdi Soltanolkotabi. Toward M oderate O verparameterization: G lobal C onvergence G uarantees for T raining S hallow N eural N etworks. IEEE Journal on Selected Areas in Information Theory, 1 0 (1): 0 84--105, 2020

  66. [74]

    Regularised L east- S quares R egression with I nfinite- D imensional O utput S pace

    Junhyung Park and Krikamol Muandet. Regularised L east- S quares R egression with I nfinite- D imensional O utput S pace. arXiv preprint arXiv:2010.10973, 2020

  67. [75]

    Towards E mpirical P rocess T heory for V ector- V alued F unctions: M etric E ntropy of S mooth F unction C lasses

    Junhyung Park and Krikamol Muandet. Towards E mpirical P rocess T heory for V ector- V alued F unctions: M etric E ntropy of S mooth F unction C lasses. In International Conference on Algorithmic Learning Theory, pages 1216--1260. PMLR, 2023

  68. [76]

    An A pproach to I nequalities for the D istributions of I nfinite- D imensional M artingales

    Iosif Pinelis. An A pproach to I nequalities for the D istributions of I nfinite- D imensional M artingales. In Probability in Banach Spaces, 8: Proceedings of the Eighth International Conference, pages 128--134. Springer, 1992

  69. [77]

    Methods in N onlinear I ntegral E quations

    Radu Precup. Methods in N onlinear I ntegral E quations . Springer Science & Business Media, 2002

  70. [78]

    Consistency of I nterpolation with L aplace K ernels is a H igh- D imensional P henomenon

    Alexander Rakhlin and Xiyu Zhai. Consistency of I nterpolation with L aplace K ernels is a H igh- D imensional P henomenon. In Conference on Learning Theory, pages 2595--2623. PMLR, 2019

  71. [79]

    Matrix A lgebra and its A pplications to S tatistics and E conometrics

    Calyampudi Radhakrishna Rao and Mareppalli Bhaskara Rao. Matrix A lgebra and its A pplications to S tatistics and E conometrics . World Scientific, 1998

  72. [80]

    Improved C onvergence G uarantees for S hallow N eural N etworks

    Alexander Razborov. Improved C onvergence G uarantees for S hallow N eural N etworks. arXiv preprint arXiv:2212.02323, 2022

  73. [81]

    Stability & G eneralisation of G radient D escent for S hallow N eural N etworks without the N eural T angent K ernel

    Dominic Richards and Ilja Kuzborskij. Stability & G eneralisation of G radient D escent for S hallow N eural N etworks without the N eural T angent K ernel. Advances in Neural Information Processing Systems, 34: 0 8609--8621, 2021

  74. [82]

    On L earning with I ntegral O perators

    Lorenzo Rosasco, Mikhail Belkin, and Ernesto De Vito. On L earning with I ntegral O perators. Journal of Machine Learning Research, 11 0 (2), 2010

  75. [83]

    Generalization properties of learning with random features

    Alessandro Rudi and Lorenzo Rosasco. Generalization properties of learning with random features. Advances in neural information processing systems, 30, 2017

  76. [84]

    Approximation T heorems of M athematical S tatistics

    Robert J Serfling. Approximation T heorems of M athematical S tatistics. Wiley Series in Probability and Statistics, 1980

  77. [85]

    Understanding M achine L earning: F rom T heory to A lgorithms

    Shai Shalev-Shwartz and Shai Ben-David. Understanding M achine L earning: F rom T heory to A lgorithms . Cambridge university press, 2014

  78. [86]

    Support V ector M achines

    Ingo Steinwart and Andreas Christmann. Support V ector M achines . Springer Science & Business Media, 2008

  79. [87]

    A N on- P arametric R egression V iewpoint: G eneralization of O verparametrized D eep R e LU N etwork under N oisy O bservations

    Namjoon Suh, Hyunouk Ko, and Xiaoming Huo. A N on- P arametric R egression V iewpoint: G eneralization of O verparametrized D eep R e LU N etwork under N oisy O bservations. In International Conference on Learning Representations, 2021

  80. [88]

    User- F riendly T ail B ounds for S ums of R andom M atrices

    Joel A Tropp. User- F riendly T ail B ounds for S ums of R andom M atrices. Foundations of computational mathematics, 12: 0 389--434, 2012

  81. [89]

    Empirical P rocesses in M - E stimation , volume 6

    Sara A van de Geer. Empirical P rocesses in M - E stimation , volume 6. Cambridge university press, 2000

  82. [90]

    On the I mplicit B ias in D eep- L earning A lgorithms

    Gal Vardi. On the I mplicit B ias in D eep- L earning A lgorithms. Communications of the ACM, 66 0 (6): 0 86--93, 2023

  83. [91]

    High- D imensional P robability: A n I ntroduction with A pplications in D ata S cience , volume 47

    Roman Vershynin. High- D imensional P robability: A n I ntroduction with A pplications in D ata S cience , volume 47. Cambridge university press, 2018

  84. [92]

    Benign overfitting in adversarial training of neural networks

    Yunjuan Wang, Kaibo Zhang, and Raman Arora. Benign overfitting in adversarial training of neural networks. In Forty-first International Conference on Machine Learning, 2024

  85. [93]

    Linear O perators in H ilbert S paces , volume 68

    Joachim Weidmann. Linear O perators in H ilbert S paces , volume 68. Springer New York, 1980

  86. [94]

    Precise L earning C urves and H igher- O rder S caling L imits for D ot P roduct K ernel R egression

    Lechao Xiao, Hong Hu, Theodor Misiakiewicz, Yue M Lu, and Jeffrey Pennington. Precise L earning C urves and H igher- O rder S caling L imits for D ot P roduct K ernel R egression. In Thirty-sixth Conference on Neural Information Processing Systems (NeurIPS), 2022

  87. [95]

    Rethinking benign overfitting in two-layer neural networks

    Ruichen Xu and Kexin Chen. Rethinking benign overfitting in two-layer neural networks. arXiv preprint arXiv:2502.11893, 2025

  88. [96]

    Benign O verfitting of N on- S mooth N eural N etworks B eyond L azy T raining

    Xingyu Xu and Yuantao Gu. Benign O verfitting of N on- S mooth N eural N etworks B eyond L azy T raining. In International Conference on Artificial Intelligence and Statistics, pages 11094--11117. PMLR, 2023

  89. [97]

    Feature L earning in I nfinite- W idth N eural N etworks

    Greg Yang and Edward J Hu. Feature L earning in I nfinite- W idth N eural N etworks. arXiv preprint arXiv:2011.14522, 2020

  90. [98]

    Sobolev norm inconsistency of kernel interpolation

    Yunfei Yang. Sobolev norm inconsistency of kernel interpolation. arXiv preprint arXiv:2504.20617, 2025

  91. [99]

    A U nifying V iew on I mplicit B ias in T raining L inear N eural N etworks

    Chulhee Yun, Shankar Krishnan, and Hossein Mobahi. A U nifying V iew on I mplicit B ias in T raining L inear N eural N etworks. arXiv preprint arXiv:2010.02501, 2020

  92. [100]

    A T ype of G eneralization E rror I nduced by I nitialization in D eep N eural N etworks

    Yaoyu Zhang, Zhi-Qin John Xu, Tao Luo, and Zheng Ma. A T ype of G eneralization E rror I nduced by I nitialization in D eep N eural N etworks. In Mathematical and Scientific Machine Learning, pages 144--164. PMLR, 2020

  93. [101]

    An A gnostic V iew on the C ost of O verfitting in ( K ernel) R idge R egression

    Lijia Zhou, James B Simon, Gal Vardi, and Nathan Srebro. An A gnostic V iew on the C ost of O verfitting in ( K ernel) R idge R egression. In International Conference on Learning Representations, 2024

  94. [102]

    Benign O verfitting in D eep N eural N etworks under L azy T raining

    Zhenyu Zhu, Fanghui Liu, Grigorios Chrysos, Francesco Locatello, and Volkan Cevher. Benign O verfitting in D eep N eural N etworks under L azy T raining. In International Conference on Machine Learning, pages 43105--43128. PMLR, 2023

  95. [103]

    Benign O verfitting of C onstant- S tepsize SGD for L inear R egression

    Difan Zou, Jingfeng Wu, Vladimir Braverman, Quanquan Gu, and Sham Kakade. Benign O verfitting of C onstant- S tepsize SGD for L inear R egression. In Conference on Learning Theory, pages 4633--4635. PMLR, 2021

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.