Pith. sign in

REVIEW 2 major objections 5 minor 160 references

The Cost of Discretization in Functional Linear Regression: Minimax Rates and Adaptation

T0 review · 2 major / 5 minor · reviewed 2026-07-13 · grok-4.5

Pith's one-line read When functional predictors are seen only through noisy points, minimax prediction risk splits into a dense-FLR term plus design-specific discretization costs.

desk verdict Sharp joint (n,m) minimax rates for discretely observed FLR, with a genuine four-term common-design rate and matching adaptive estimators. read the letter →

arxiv 2607.09350 v1 pith:44HC6J6H submitted 2026-07-10 math.ST stat.TH

classification math.STstat.TH MSC 62G0862M2062R10
keywords functionallinearregressionminimaxratesdiscretenoisyobservationsindependentdesigncommonadaptiveestimationphasetransitioneigenvalueidentification
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Functional linear regression is usually studied as if each predictor curve is fully observed. In practice one sees only m noisy point samples on each of n curves. This paper gives exact minimax rates for prediction (and estimation) as joint functions of n and m under two standard designs. Under independent random grids the rate is the familiar fully observed benchmark plus a second term that reflects noisy sampling after amplification by the inverse covariance; the phase transition occurs when m scales like n to a power set by eigenvalue decay and slope smoothness. Under a shared equispaced grid two further unavoidable terms appear: fixed-grid approximation error and the cost of identifying unknown eigenvalues from the common lattice. Adaptive estimators that screen covariance scale and threshold blockwise prediction energy attain the rates without knowing the spectrum or the smoothness indices. The message is that discretization is not a mere technical nuisance: sampling geometry itself changes the statistical experiment and forces new rate terms that cannot be averaged away by collecting more subjects.

What carries the argument

Matching upper and lower bounds obtained by combining a plug-in Fourier estimator (with eigenvalue screening and blockwise prediction-energy thresholding) against van Trees lower bounds that control Fisher information for noisy point evaluations, plus two indistinguishability constructions that isolate fixed-grid aliasing and unknown-eigenvalue identification on a common lattice.

What would settle it

If, for independent design, the excess prediction risk fails to obey the two-term rate (or the stated phase transition at m ~ n^{2α/(2α+2s+1)}) under the paper's Gaussian trigonometric model, or if common-design risk remains free of the m^{-4α} term when eigenvalues are unknown, the claimed minimax characterization is false.

Watch

Extended reading notes

Core claim

Under independent design the minimax prediction risk is of exact order n to the power -(2α+2s)/(2α+2s+1) plus (nm) to the power -(2α+2s)/(4α+2s+1). Under common design with unknown eigenvalues the same two terms remain and are joined by m to the power -(2α+2s) and m to the power -4α, with matching adaptive upper bounds. The first term recovers the fully observed functional-linear-regression benchmark; the second is the statistical price of noisy point evaluations after inverse-covariance amplification; the last two are geometric obstructions created by a shared grid.

Load-bearing premise

The covariance operator is assumed to be diagonalized by a fixed, known trigonometric basis, so the theory does not yet pay for estimating an unknown eigenbasis or for misalignment between covariance geometry and slope smoothness.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper derives matching minimax rates for prediction (and sketches estimation) in scalar-on-function linear regression when each trajectory is observed only at m noisy points. Working in a fixed trigonometric eigenbasis with eigenvalue decay α and slope Sobolev smoothness s, it treats independent random design and common equispaced design separately. Under independent design the minimax prediction risk is of order n^{-(2α+2s)/(2α+2s+1)}+(nm)^{-(2α+2s)/(4α+2s+1)}; under common design with unknown eigenvalues the rate adds the grid terms m^{-(2α+2s)}+m^{-4α}. Adaptive estimators that screen covariance scale and threshold blockwise prediction energy attain these rates without knowledge of (λ_r), α, or s. Phase diagrams, simulations, and a wheat-spectra example illustrate the theory.

Significance. The contribution is substantial for functional data analysis and nonparametric inverse problems. Fully observed FLR rates and sparse-to-dense transitions for mean/covariance estimation were known; a complete (n,m)-minimax theory for FLR that separates noisy-point cost from fixed-grid aliasing and eigenvalue identification was not. Matching lower bounds (van Trees with refined Fisher control; two distinct common-grid indistinguishability constructions) and adaptive upper bounds make the phase transitions sharp rather than merely sufficient. The fixed trigonometric eigenbasis is an explicit modeling choice that isolates discretization cost; within that model the results are definitive and will serve as a benchmark for subsequent work on unknown eigenbases and alignment.

major comments (2)
  1. Section 6 states L2 estimation rates (independent: n^{-2s/(2α+2s+1)}+(nm)^{-2s/(4α+2s+1)}; common: plus m^{-2s}+m^{-4α}) as following by the same arguments after reweighting by λ_r^2 and adjusting the energy threshold. The sketch is plausible but incomplete: the van Trees block priors, Fisher bounds, and eligible-block risk decompositions all change with the unweighted loss, and the plug-in remainder B^λ_ℓ must be re-controlled. Either supply the full parallel proofs (or a self-contained appendix lemma) or present the L2 rates as conjectured extensions rather than established corollaries.
  2. The fixed trigonometric eigenbasis (Section 2 after (4)–(5)) is load-bearing for every rate. Section 6 correctly flags eigenbasis estimation and alignment as open, but the abstract and introduction still present the rates as the cost of discretization for FLR without always restating the basis restriction. A short, prominent caveat in the abstract and at the start of Corollaries 3.4 and 4.4 would prevent misapplication when the covariance eigenbasis is unknown or misaligned with the Sobolev scale of β.
minor comments (5)
  1. Table 2 (wheat data): CD beats IND and the published benchmarks, contrary to the asymptotic ranking. The finite-sample explanation is reasonable; a one-sentence note that the common subgrid is deterministic while IND and the Zhou et al. numbers use random subsampling would make the comparison protocol fully transparent.
  2. Figure 1 phase diagrams are clear; adding the critical exponents (e.g. ζ_ind = 2α/(ν+1)) as axis annotations would help readers match the figure to Subsection 4.4.1 without flipping back to the text.
  3. Notation: ν := 2α+2s and κ := 4α+2s+1 are introduced late (Table 1 / §4.4). Defining them once in Section 2 or at the first rate display would reduce repetition.
  4. Adaptive constants (M0, CV, CL, cL) are fixed but their practical choice is not discussed. A brief remark in Section 5 on the values used in the simulations would aid reproducibility.
  5. Typos/style: “What left open” (p. 3) → “What is left open”; occasional double spaces and “them −4α” line-break artifacts in the PDF.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: minimax rates are derived from stated model assumptions via Fisher/van Trees lower bounds and constructive oracle/adaptive upper bounds, not by fitting or redefining the target.

full rationale

The paper is a standard minimax theory contribution. Under a fixed trigonometric eigenbasis that diagonalizes the covariance (Section 2), prediction risk is the weighted ℓ2 error ∑ λ_r |θ̂_r − θ_r|². Oracle upper bounds (Theorems 3.1, 4.1) balance truncation bias against variance of cross-covariance estimators; matching lower bounds use van Trees on Gaussian block submodels with Fisher information controlled for noisy point evaluations (Appendix D) and, under common design, two distinct indistinguishability constructions for fixed-grid aliasing and unknown-eigenvalue identification (Appendix G). Adaptive procedures screen eigenvalues and threshold blockwise prediction energy with fixed universal constants (M0, CV, CL), attaining the same rates without knowledge of (λ_r), α, or s (Theorems 3.3, 4.3). Self-citations to Cai–Yuan supply the fully-observed FLR benchmark term and mean/covariance phase-transition context; they do not redefine or force the new (nm) and m-only discretization terms. Simulations and the wheat example illustrate the theory rather than fit the rates. No step reduces the claimed rates to their inputs by construction.

Assumptions & free parameters 2 free parameters · 5 assumptions · 0 invented entities

The central rates rest on a standard Gaussian FLR model plus a strong but explicit geometric assumption that a fixed trigonometric basis diagonalizes the covariance. Smoothness and eigenvalue-decay classes are classical. Adaptive estimators introduce fixed numerical tuning constants that affect constants, not rates. No new physical entities are postulated.

free parameters (2)
  • Adaptive pilot/threshold constants (M0, CV, CL, cL, R)
    Fixed large/small numerical constants in eigenvalue floors and block energy thresholds; chosen sufficiently large/small for theory, not fitted to the wheat or simulation risk surfaces.
  • Noise levels σε, σδ and eigenvalue envelope constants cλ, Cλ, R0
    Model parameters treated as fixed known-order constants in rate statements; simulations pick concrete values (e.g. σε=0.5, σδ=0.1) for illustration only.
assumptions (5)
  • domain assumption Predictor trajectories are centered Gaussian processes with continuous paths; response and measurement errors are independent Gaussian.
    Section 2 model (1)–(2); used for Fisher information and concentration throughout lower and upper bounds.
  • domain assumption Covariance operator is diagonalized by the fixed trigonometric basis with eigenvalues in Lα(cλ,Cλ), α>1/2.
    Section 2 after (4)–(9); isolates discretization cost from eigenbasis estimation.
  • domain assumption Slope Fourier coefficients lie in Sobolev ball Θs(R0), s≥0.
    Equation (10); defines the parameter space for minimax rates.
  • domain assumption Independent design: tij iid Unif[0,1]; common design: equal grid tj=(j-1)/m.
    Assumptions 1–2; define the two statistical experiments.
  • standard math Van Trees inequality and standard sub-Weibull/Bernstein concentration tools.
    Lemma D.1 and Appendix H; used for lower and adaptive upper bounds.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Cost of Discretization in Functional Linear Regression: Minimax Rates and Adaptation." pith.science (2026). https://pith.science/paper/44HC6J6H

@misc{pith2026260709350,
  author       = {Pith},
  title        = {Pith review of: The Cost of Discretization in Functional Linear Regression: Minimax Rates and Adaptation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/44HC6J6H}},
  note         = {Machine review of arXiv:2607.09350}
}
abstract

We study scalar-on-function linear regression when each covariate curve is observed only through finitely many noisy point evaluations. Our goal is to characterize the minimax estimation and prediction risks as joint functions of the number of trajectories $n$ and the within-trajectory resolution $m$. Working in a fixed trigonometric eigenbasis, with covariance eigenvalues decaying at rate $\alpha$ and slope function of Sobolev smoothness $s$, we derive matching minimax upper and lower bounds under two canonical sampling schemes. Under an independent random design, the minimax prediction rate is $n^{-\frac{2\alpha+2s}{2\alpha+2s+1}} + (nm)^{-\frac{2\alpha+2s}{4\alpha+2s+1}}$. The first term is the fully observed functional linear regression benchmark, while the second term captures the cost of noisy point evaluations after amplification by the inverse covariance operator. Under a common design on an equally spaced grid, the shared sampling geometry introduces additional obstructions, and the minimax prediction rate becomes $n^{-\frac{2\alpha+2s}{2\alpha+2s+1}} + (nm)^{-\frac{2\alpha+2s}{4\alpha+2s+1}} + m^{-(2\alpha+2s)} + m^{-4\alpha}$. Here the third term represents discretization error induced by the fixed grid, whereas the fourth reflects the cost of identifying unknown eigenvalues from observations on a common grid. We further construct data-driven adaptive estimators that screen the covariance scale and threshold blockwise prediction energy, attaining these rates without prior knowledge of the eigenvalue sequence or the smoothness indices. The results reveal a sharp phase transition that depends on the sampling resolution under independent design and a richer phase diagram under common design. Numerical simulations and a real data example illustrate the theoretical findings.

Figures

Figures reproduced from arXiv: 2607.09350 by the authors.

Figure 1
Figure 1. Phase diagram for the minimax rate under common design. The dominant term in ( [PITH_FULL_IMAGE:figures/full_fig_p015_1.png] view at source ↗
Figure 2
Figure 2. Simulation heatmaps under independent design (top) and common design (bottom). [PITH_FULL_IMAGE:figures/full_fig_p019_2.png] view at source ↗
Figure 3
Figure 3. Matched comparison between common and independent design for ( [PITH_FULL_IMAGE:figures/full_fig_p019_3.png] view at source ↗
Figures from the paper (9 more)
Figure 4
Figure 4. Figure 4: Common design rate comparisons as m varies (top) and along m = n γ (bottom). Each panel uses 50 repetitions and plots the median excess prediction risk with interquartile error bars. The dashed reference lines show the low-resolution term predicted by (35) in the top p…
Figure 5
Figure 5. Figure 5: Known versus estimated eigenvalues under common design for ( [PITH_FULL_IMAGE:figures/full_fig_p021_5.png]
Figure 6
Figure 6. Figure 6: Additional simulation heatmaps under common design. [PITH_FULL_IMAGE:figures/full_fig_p090_6.png]
Figure 7
Figure 7. Figure 7: Additional simulation heatmaps under independent design. [PITH_FULL_IMAGE:figures/full_fig_p091_7.png]
Figure 8
Figure 8. Figure 8: Additional common-design rate comparisons. [PITH_FULL_IMAGE:figures/full_fig_p092_8.png]
Figure 9
Figure 9. Figure 9: Additional independent-design rate comparisons. [PITH_FULL_IMAGE:figures/full_fig_p093_9.png]
Figure 10
Figure 10. Figure 10: Common-design rate comparisons with known and unknown eigenvalues. The left [PITH_FULL_IMAGE:figures/full_fig_p094_10.png]
Figure 11
Figure 11. Figure 11: Measurement-noise sensitivity under independent design (top) and common design [PITH_FULL_IMAGE:figures/full_fig_p096_11.png]
Figure 12
Figure 12. Figure 12: Observation-noise sensitivity under independent design (top) and common design (bot [PITH_FULL_IMAGE:figures/full_fig_p097_12.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

160 extracted references · 27 canonical work pages

  1. [1]

    2003 , publisher =

    Sobolev Spaces , author =. 2003 , publisher =

  2. [2]

    2008 , journal =

    A Tail Inequality for Suprema of Unbounded Empirical Processes with Applications to Markov Chains , author =. 2008 , journal =

  3. [3]

    Bousquet, Olivier , year =. A. Comptes Rendus. Math. doi:10.1016/S1631-073X(02)02292-6 , url =

  4. [4]

    2013 , publisher =

    Concentration Inequalities: A Nonasymptotic Theory of Independence , author =. 2013 , publisher =. doi:10.1093/acprof:oso/9780199535255.001.0001 , isbn =

  5. [5]

    2019 , journal =

    On the Convergence Rate of Training Recurrent Neural Networks , author =. 2019 , journal =

  6. [6]

    2019 , month = jun, eprint =

    A Convergence Theory for Deep Learning via Over-Parameterization , author =. 2019 , month = jun, eprint =

  7. [7]

    2021 , journal =

    Concentration of Kernel Matrices with Application to Kernel Spectral Clustering , author =. 2021 , journal =

  8. [8]

    2008 , series =

    Support Vector Machines , author =. 2008 , series =. doi:10.1007/978-0-387-77242-4 , isbn =

Show all 160 references
  1. [9]

    and Hu, Wei and Li, Zhiyuan and Salakhutdinov, Russ R and Wang, Ruosong , year =

    Arora, Sanjeev and Du, Simon S. and Hu, Wei and Li, Zhiyuan and Salakhutdinov, Russ R and Wang, Ruosong , year =. On Exact Computation with an Infinitely Wide Neural Net , booktitle =

  2. [10]

    Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks , booktitle =

    Arora, Sanjeev and Du, Simon and Hu, Wei and Li, Zhiyuan and Wang, Ruosong , year =. Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks , booktitle =

  3. [11]

    2014 , month = jan, journal =

    Sharp Estimates for Eigenvalues of Integral Operators Generated by Dot Product Kernels on the Sphere , author =. 2014 , month = jan, journal =. doi:10.1016/j.jat.2013.10.002 , abstract =

  4. [12]

    2017 , journal =

    Breaking the Curse of Dimensionality with Convex Neural Networks , author =. 2017 , journal =

  5. [13]

    Characterization of

    Barcel. Characterization of. 2019 , month =. doi:10.48550/arXiv.1907.01571 , abstract =. arxiv , keywords =:arXiv:1907.01571 , publisher =

  6. [14]

    2022 , month = nov, number =

    A Kernel Perspective of Skip Connections in Convolutional Networks , author =. 2022 , month = nov, number =. doi:10.48550/arXiv.2211.14810 , abstract =. arxiv , keywords =:arXiv:2211.14810 , publisher =

  7. [15]

    2007 , journal =

    On Regularization Algorithms in Learning Theory , author =. 2007 , journal =. doi:10.1016/j.jco.2006.07.001 , abstract =

  8. [16]

    On Deep Learning as a Remedy for the Curse of Dimensionality in Nonparametric Regression , author =

  9. [17]

    2022 , month = jun, number =

    Kernel Ridgeless Regression Is Inconsistent in Low Dimensions , author =. 2022 , month = jun, number =. doi:10.48550/arXiv.2205.13525 , abstract =. arXiv:2205.13525 , publisher =

  10. [18]

    Beatson, R. K. and zu Castell, W. and Xu, Y. , year =. A. doi:10.48550/arXiv.1110.2437 , abstract =. arxiv , keywords =:arXiv:1110.2437 , publisher =

  11. [19]

    To Understand Deep Learning We Need to Understand Kernel Learning , booktitle =

    Belkin, Mikhail and Ma, Siyuan and Mandal, Soumik , year =. To Understand Deep Learning We Need to Understand Kernel Learning , booktitle =

  12. [20]

    On the Inductive Bias of Neural Tangent Kernels , booktitle =

    Bietti, Alberto and Mairal, Julien , year =. On the Inductive Bias of Neural Tangent Kernels , booktitle =

  13. [21]

    2020 , journal =

    Deep Equals Shallow for Relu Networks in Kernel Regimes , author =. 2020 , journal =. 2009.14397 , archiveprefix =

  14. [22]

    , year =

    Bingham, Nicholas H. , year =. Positive Definite Functions on Spheres , booktitle =

  15. [23]

    2018 , journal =

    Optimal Rates for Regularization of Statistical Inverse Learning Problems , author =. 2018 , journal =. doi:10.1007/s10208-017-9359-7 , abstract =

  16. [24]

    2003 , journal =

    Spectral Properties of Distance Matrices , author =. 2003 , journal =

  17. [25]

    Spectrum Dependent Learning Curves in Kernel Regression and Wide Neural Networks , booktitle =

    Bordelon, Blake and Canatar, Abdulkadir and Pehlevan, Cengiz , year =. Spectrum Dependent Learning Curves in Kernel Regression and Wide Neural Networks , booktitle =

  18. [26]

    and Dick, Josef , year =

    Brauchart, Johann S. and Dick, Josef , year =. A Characterization of. Constructive Approximation , volume =. doi:10.1007/s00365-013-9217-z , abstract =

  19. [27]

    2005 , abstract =

    Spectral Properties of the Kernel Matrix and Their Relation to Kernel Methods in Machine Learning , author =. 2005 , abstract =

  20. [28]

    2022 , month =

    What Can Be Learnt with Wide Convolutional Neural Networks? , author =. 2022 , month =. doi:10.48550/arXiv.2208.01003 , abstract =. arxiv , keywords =:arXiv:2208.01003 , publisher =

  21. [29]

    2021 , month = may, journal =

    Spectral Bias and Task-Model Alignment Explain Generalization in Kernel Regression and Infinitely Wide Neural Networks , author =. 2021 , month = may, journal =. doi:10.1038/s41467-021-23103-1 , abstract =

  22. [30]

    Generalization Error Bounds of Gradient Descent for Learning Over-Parameterized Deep

    Cao, Yuan and Gu, Quanquan , year =. Generalization Error Bounds of Gradient Descent for Learning Over-Parameterized Deep. Proceedings of the

  23. [31]

    2007 , journal =

    Optimal Rates for the Regularized Least-Squares Algorithm , author =. 2007 , journal =. doi:10.1007/s10208-006-0196-8 , abstract =

  24. [32]

    2010 , journal =

    Cross-Validation Based Adaptation for Regularization Operators in Learning Theory , author =. 2010 , journal =. doi:10.1142/S0219530510001564 , abstract =

  25. [33]

    2012 , journal =

    Eigenvalue Decay of Positive Integral Operators on the Sphere , author =. 2012 , journal =

  26. [34]

    Deep Neural Tangent Kernel and Laplace Kernel Have the Same

    Chen, Lin and Xu, Sheng , year =. Deep Neural Tangent Kernel and Laplace Kernel Have the Same. arXiv preprint arXiv:2009.10683 , eprint =

  27. [35]

    Kernel Methods for Deep Learning , booktitle =

    Cho, Youngmin and Saul, Lawrence , editor =. Kernel Methods for Deep Learning , booktitle =. 2009 , volume =

  28. [36]

    The Loss Surfaces of Multilayer Networks , booktitle =

    Choromanska, Anna and Henaff, Mikael and Mathieu, Michael and Arous, G. The Loss Surfaces of Multilayer Networks , booktitle =. 2015 , pages =

  29. [37]

    Conference on Learning Theory , author =

    On the Expressive Power of Deep Learning:. Conference on Learning Theory , author =. 2016 , pages =

  30. [38]

    2013 , series =

    Approximation Theory and Harmonic Analysis on Spheres and Balls , author =. 2013 , series =. doi:10.1007/978-1-4614-6660-4 , isbn =

  31. [39]

    doi:10.48550/arXiv.1810.04805 , abstract =

    Devlin, Jacob and Chang, Ming-Wei and Lee, Kenton and Toutanova, Kristina , year =. doi:10.48550/arXiv.1810.04805 , abstract =. arxiv , keywords =:arXiv:1810.04805 , publisher =

  32. [40]

    and Zhai, Xiyu and Poczos, Barnabas and Singh, Aarti , year =

    Du, Simon S. and Zhai, Xiyu and Poczos, Barnabas and Singh, Aarti , year =. Gradient Descent Provably Optimizes Over-Parameterized Neural Networks , booktitle =

  33. [41]

    and Lee, Jason and Li, Haochuan and Wang, Liwei and Zhai, Xiyu , year =

    Du, Simon S. and Lee, Jason and Li, Haochuan and Wang, Liwei and Zhai, Xiyu , year =. Gradient Descent Finds Global Minima of Deep Neural Networks , booktitle =

  34. [42]

    Deep Neural Networks as Point Estimates for Deep

    Dutordoir, Vincent and Hensman, James and. Deep Neural Networks as Point Estimates for Deep. Advances in. 2021 , volume =

  35. [43]

    2020 , journal =

    Spectra of the Conjugate Kernel and Neural Tangent Kernel for Linear-Width Neural Networks , author =. 2020 , journal =

  36. [44]

    2008 , month = oct, journal =

    Integral Operators on the Sphere Generated by Positive Definite Smooth Kernels , author =. 2008 , month = oct, journal =. doi:10.1016/j.jco.2008.04.001 , abstract =

  37. [45]

    2009 , month = may, journal =

    Eigenvalues of Integral Operators Defined by Smooth Positive Definite Kernels , author =. 2009 , month = may, journal =. doi:10.1007/s00020-009-1680-3 , abstract =

  38. [46]

    2013 , month = dec, journal =

    Eigenvalue Decay Rates for Positive Integral Operators , author =. 2013 , month = dec, journal =. doi:10.1007/s10231-012-0256-z , abstract =

  39. [47]

    2020 , journal =

    Sobolev Norm Learning Rates for Regularized Least-Squares Algorithms , author =. 2020 , journal =

  40. [48]

    and Bartlett, Peter , year =

    Frei, Spencer and Chatterji, Niladri S. and Bartlett, Peter , year =. Benign Overfitting without Linearity:. Proceedings of

  41. [49]

    Notes on Spherical Harmonics and Linear Representations of Lie Groups , author =

  42. [50]

    2018 , journal =

    Deep Convolutional Networks as Shallow Gaussian Processes , author =. 2018 , journal =. 1808.05587 , archiveprefix =

  43. [51]

    On the Similarity between the

    Geifman, Amnon and Yadav, Abhay and Kasten, Yoni and Galun, Meirav and Jacobs, David and Ronen, Basri , year =. On the Similarity between the. Advances in

  44. [52]

    2008 , journal =

    Spectral Algorithms for Supervised Learning , author =. 2008 , journal =. doi:10.1162/neco.2008.05-07-517 , abstract =

  45. [53]

    2020 , journal =

    When Do Neural Networks Outperform Kernel Methods? , author =. 2020 , journal =

  46. [54]

    Linearized Two-Layers Neural Networks in High Dimension , author =

  47. [55]

    2013 , journal =

    Strictly and Non-Strictly Positive Definite Functions on Spheres , author =. 2013 , journal =

  48. [56]

    2002 , series =

    A Distribution-Free Theory of Nonparametric Regression , editor =. 2002 , series =

  49. [57]

    Approximating Continuous Functions by

    Hanin, Boris and Sellke, Mark , year =. Approximating Continuous Functions by. arXiv preprint arXiv:1710.11278 , eprint =

  50. [58]

    Random Neural Networks in the Infinite Width Limit as

    Hanin, Boris , year =. Random Neural Networks in the Infinite Width Limit as. arXiv preprint arXiv:2107.01562 , eprint =

  51. [59]

    2020 , journal =

    On the Minimax Optimality and Superiority of Deep Neural Network Learning over Sparse Parameter Spaces , author =. 2020 , journal =

  52. [60]

    The Curse of Depth in Kernel Regime , booktitle =

    Hayou, Soufiane and Doucet, Arnaud and Rousseau, Judith , year =. The Curse of Depth in Kernel Regime , booktitle =

  53. [61]

    Deep Residual Learning for Image Recognition , booktitle =

    He, Kaiming and Zhang, Xiangyu and Ren, Shaoqing and Sun, Jian , year =. Deep Residual Learning for Image Recognition , booktitle =

  54. [62]

    2017 , edition =

    Matrix Analysis , author =. 2017 , edition =

  55. [63]

    1989 , journal =

    Multilayer Feedforward Networks Are Universal Approximators , author =. 1989 , journal =

  56. [64]

    1939 , journal =

    Tubes and Spheres in N-Spaces, and a Class of Statistical Problems , author =. 1939 , journal =

  57. [65]

    Regularization Matters:

    Hu, Tianyang and Wang, Wenjia and Lin, Cong and Cheng, Guang , year =. Regularization Matters:. International

  58. [66]

    2022 , month = may, number =

    Sharp Asymptotics of Kernel Ridge Regression beyond the Linear Regime , author =. 2022 , month = may, number =. arxiv , langid =:arXiv:2205.06798 , publisher =

  59. [67]

    Dynamics of Deep Neural Networks and Neural Tangent Hierarchy , booktitle =

    Huang, Jiaoyang and Yau, Horng-Tzer , year =. Dynamics of Deep Neural Networks and Neural Tangent Hierarchy , booktitle =

  60. [68]

    Advances in Neural Information Processing Systems , author =

    Neural Tangent Kernel:. Advances in Neural Information Processing Systems , author =. 2018 , volume =

  61. [69]

    2022 , abstract =

    Generalization of Wide Neural Networks from the Perspective of Linearization and Kernel Learning , author =. 2022 , abstract =

  62. [70]

    , author =

    Distribution of Eigenvalues of Certain Integral Operators. , author =. 1955 , journal =

  63. [71]

    , year =

    Kanagawa, Motonobu and Hennig, Philipp and Sejdinovic, Dino and Sriperumbudur, Bharath K. , year =. Gaussian Processes and Kernel Methods:. arXiv preprint arXiv:1807.02582 , eprint =

  64. [72]

    A Style-Based Generator Architecture for Generative Adversarial Networks , booktitle =

    Karras, Tero and Laine, Samuli and Aila, Timo , year =. A Style-Based Generator Architecture for Generative Adversarial Networks , booktitle =

  65. [73]

    1991 , month = feb, journal =

    Asymptotic Properties of Eigenvalues of Integral Equations , author =. 1991 , month = feb, journal =. doi:10.1137/0151013 , abstract =

  66. [74]

    2000 , journal =

    Random Matrix Approximation of Spectra of Integral Operators , author =. 2000 , journal =. doi:10.2307/3318636 , abstract =

  67. [75]

    2017 , journal =

    Imagenet Classification with Deep Convolutional Neural Networks , author =. 2017 , journal =

  68. [76]

    2022 , journal =

    Moving Beyond Sub-Gaussianity in High-Dimensional Statistics: Applications in Covariance Estimation and Linear Regression , author =. 2022 , journal =

  69. [77]

    Nonparametric Regression with Shallow Overparameterized Neural Networks Trained by

    Kuzborskij, Ilja and Szepesv. Nonparametric Regression with Shallow Overparameterized Neural Networks Trained by. Conference on. 2021 , pages =

  70. [78]

    Generalization Ability of Wide Neural Networks on

    Lai, Jianfa and Xu, Manyun and Chen, Rui and Lin, Qian , year =. Generalization Ability of Wide Neural Networks on. doi:10.48550/arXiv.2302.05933 , abstract =. arxiv , keywords =:arXiv:2302.05933 , publisher =

  71. [79]

    2017 , journal =

    Deep Neural Networks as Gaussian Processes , author =. 2017 , journal =. 1711.00165 , archiveprefix =

  72. [80]

    Wide Neural Networks of Any Depth Evolve as Linear Models under Gradient Descent , booktitle =

    Lee, Jaehoon and Xiao, Lechao and Schoenholz, Samuel and Bahri, Yasaman and Novak, Roman and. Wide Neural Networks of Any Depth Evolve as Linear Models under Gradient Descent , booktitle =. 2019 , volume =

  73. [81]

    2022 , month = sep, number =

    Stability and Generalization Analysis of Gradient Methods for Shallow Neural Networks , author =. 2022 , month = sep, number =. doi:10.48550/arXiv.2209.09298 , abstract =. arxiv , keywords =:arXiv:2209.09298 , publisher =

  74. [82]

    2018 , journal =

    Learning Overparameterized Neural Networks via Stochastic Gradient Descent on Structured Data , author =. 2018 , journal =

  75. [83]

    On the Saturation Effect of Kernel Ridge Regression , booktitle =

    Li, Yicheng and Zhang, Haobo and Lin, Qian , year =. On the Saturation Effect of Kernel Ridge Regression , booktitle =

  76. [84]

    Just Interpolate:

    Liang, Tengyuan and Rakhlin, Alexander , year =. Just Interpolate:. The Annals of Statistics , volume =. doi:10.1214/19-AOS1849 , abstract =. arxiv , keywords =:1808.00387 , primaryclass =

  77. [85]

    and Cevher, V

    Lin, Junhong and Rudi, Alessandro and Rosasco, L. and Cevher, V. , year =. Optimal Rates for Spectral Algorithms with Least-Squares Regression over. Applied and Computational Harmonic Analysis , volume =. doi:10.1016/j.acha.2018.09.009 , abstract =

  78. [86]

    Kernel Regression in High Dimensions:

    Liu, Fanghui and Liao, Zhenyu and Suykens, Johan , year =. Kernel Regression in High Dimensions:. International

  79. [87]

    The Expressive Power of Neural Networks:

    Lu, Zhou and Pu, Hongming and Wang, Feicheng and Hu, Zhiqiang and Wang, Liwei , year =. The Expressive Power of Neural Networks:. Advances in neural information processing systems , volume =

  80. [88]

    2018 , journal =

    Gaussian Process Behaviour in Wide Deep Neural Networks , author =. 2018 , journal =. 1804.11271 , archiveprefix =

  81. [89]

    2010 , month = feb, journal =

    Regularization in Kernel Learning , author =. 2010 , month = feb, journal =. doi:10.1214/09-AOS728 , abstract =

  82. [90]

    The Interpolation Phase Transition in Neural Networks:

    Montanari, Andrea and Zhong, Yiqiao , year =. The Interpolation Phase Transition in Neural Networks:. The Annals of Statistics , volume =

  83. [91]

    Characterizing the Spectrum of the

    Murray, Michael and Jin, Hui and Bowman, Benjamin and Montufar, Guido , year =. Characterizing the Spectrum of the. arXiv preprint arXiv:2211.07844 , eprint =

  84. [92]

    Deep Double Descent:

    Nakkiran, Preetum and Kaplun, Gal and Bansal, Yamini and Yang, Tristan and Barak, Boaz and Sutskever, Ilya , year =. Deep Double Descent:. International

  85. [93]

    , year =

    Nguyen, Quynh and Mondelli, Marco and Montufar, Guido F. , year =. Tight Bounds on the Smallest Eigenvalue of the Neural Tangent Kernel for Deep. International

  86. [94]

    Toward Moderate Overparameterization: Global Convergence Guarantees for Training Shallow Neural Networks , shorttitle =

    Oymak, Samet and Soltanolkotabi, Mahdi , year =. Toward Moderate Overparameterization: Global Convergence Guarantees for Training Shallow Neural Networks , shorttitle =. IEEE Journal on Selected Areas in Information Theory , volume =. doi:10.1109/JSAIT.2020.2991332 , abstract =

  87. [95]

    Optimal Approximation of Piecewise Smooth Functions Using Deep

    Petersen, Philipp and Voigtlaender, Felix , year =. Optimal Approximation of Piecewise Smooth Functions Using Deep. Neural Networks , volume =

  88. [96]

    Eigenvalues of Integral Operators

    Pietsch, Albrecht , year =. Eigenvalues of Integral Operators. Mathematische Annalen , volume =. doi:10.1007/BF01456014 , langid =

  89. [97]

    Consistency of Interpolation with

    Rakhlin, Alexander and Zhai, Xiyu , year =. Consistency of Interpolation with. arxiv , keywords =:arXiv:1812.11167 , publisher =

  90. [98]

    2017 , journal =

    Optimal Rates for the Regularized Learning Algorithms under General Source Condition , author =. 2017 , journal =. doi:10.3389/fams.2017.00003 , abstract =

  91. [99]

    2021 , journal =

    Stability & Generalisation of Gradient Descent for Shallow Neural Networks without the Neural Tangent Kernel , author =. 2021 , journal =

  92. [100]

    2019 , journal =

    The Convergence Rate of Neural Networks for Learned Functions of Different Frequencies , author =. 2019 , journal =

  93. [101]

    1963 , journal =

    Some Results on the Asymptotic Behavior of Eigenvalues for a Class of Integral Equations with Translation Kernels , author =. 1963 , journal =

  94. [102]

    Theory of

    Sawano, Yoshihiro , year =. Theory of. doi:10.1007/978-981-13-0836-9 , isbn =

  95. [103]

    A Spectral Analysis of Dot-Product Kernels , booktitle =

    Scetbon, Meyer and Harchaoui, Zaid , year =. A Spectral Analysis of Dot-Product Kernels , booktitle =

  96. [104]

    2015 , month = nov, publisher =

    Operator Theory , author =. 2015 , month = nov, publisher =. doi:10.1090/simon/004 , isbn =

  97. [105]

    2021 , journal =

    Neural Tangent Kernel Eigenvalues Accurately Predict Generalization , author =. 2021 , journal =. 2110.03922 , archiveprefix =

  98. [106]

    2000 , journal =

    Regularization with Dot-Product Kernels , author =. 2000 , journal =

  99. [107]

    and Tomas, Peter A

    Stanton, Robert J. and Tomas, Peter A. , year =. Polyhedral Summability of. American Journal of Mathematics , volume =. 2373834 , eprinttype =

  100. [108]

    and Scovel, C

    Steinwart, Ingo and Hush, D. and Scovel, C. , year =. Optimal Rates for Regularized Least Squares Regression , booktitle =

  101. [109]

    , year =

    Steinwart, Ingo and Scovel, C. , year =. Mercer's Theorem on General Domains:. Constructive Approximation , volume =. doi:10.1007/S00365-012-9153-3 , abstract =

  102. [110]

    Convergence Types and Rates in Generic

    Steinwart, Ingo , year =. Convergence Types and Rates in Generic. Potential Analysis , volume =

  103. [111]

    A Non-Parametric Regression Viewpoint:

    Suh, Namjoon and Ko, Hyunouk and Huo, Xiaoming , year =. A Non-Parametric Regression Viewpoint:. International

  104. [112]

    Adaptivity of Deep

    Suzuki, Taiji , year =. Adaptivity of Deep. doi:10.48550/arXiv.1810.08033 , abstract =. arxiv , keywords =:arXiv:1810.08033 , publisher =

  105. [113]

    2015 , journal =

    Representation Benefits of Deep Feedforward Networks , author =. 2015 , journal =. 1509.08101 , archiveprefix =

  106. [114]

    Generalization of

    Than, Khoat and Vu, Nghia , year =. Generalization of. arXiv preprint arXiv:2104.02388 , eprint =

  107. [115]

    Kernel-Based Smoothness Analysis of Residual Networks , booktitle =

    Tirer, Tom and Bruna, Joan and Giryes, Raja , year =. Kernel-Based Smoothness Analysis of Residual Networks , booktitle =

  108. [116]

    2009 , series =

    Introduction to Nonparametric Estimation , author =. 2009 , series =

  109. [117]

    1999 , publisher =

    The Nature of Statistical Learning Theory , author =. 1999 , publisher =

  110. [118]

    2010 , journal =

    Introduction to the Non-Asymptotic Analysis of Random Matrices , author =. 2010 , journal =. 1011.3027 , archiveprefix =

  111. [119]

    Limitations of the

    Vyas, Nikhil and Bansal, Yamini and Nakkiran, Preetum , year =. Limitations of the. arXiv preprint arXiv:2206.10012 , eprint =

  112. [120]

    2004 , series =

    Scattered Data Approximation , author =. 2004 , series =. doi:10.1017/CBO9780511617539 , abstract =

  113. [121]

    1939 , journal =

    On the Volume of Tubes , author =. 1939 , journal =

  114. [122]

    1963 , journal =

    Asymptotic Behavior of the Eigenvalues of Certain Integral Equations , author =. 1963 , journal =. 1993907 , eprinttype =

  115. [123]

    Eigenspace Restructuring:

    Xiao, Lechao , year =. Eigenspace Restructuring:. Proceedings of

  116. [124]

    Tensor Programs i: Wide Feedforward or Recurrent Neural Networks of Any Architecture Are Gaussian Processes , shorttitle =

    Yang, Greg , year =. Tensor Programs i: Wide Feedforward or Recurrent Neural Networks of Any Architecture Are Gaussian Processes , shorttitle =. doi:10.48550/arXiv.1910.12478 , abstract =. arxiv , keywords =:arXiv:1910.12478 , publisher =

  117. [125]

    2007 , month = aug, journal =

    On Early Stopping in Gradient Descent Learning , author =. 2007 , month = aug, journal =. doi:10.1007/s00365-006-0663-2 , abstract =

  118. [126]

    Error Bounds for Approximations with Deep

    Yarotsky, Dmitry , year =. Error Bounds for Approximations with Deep. Neural Networks , volume =

  119. [127]

    2019 , journal =

    On the Power and Limitations of Random Features for Understanding Neural Networks , author =. 2019 , journal =

  120. [128]

    Boosting with Early Stopping:

    Zhang, Tong and Yu, Bin , year =. Boosting with Early Stopping:. The Annals of Statistics , volume =

  121. [129]

    2017 , month = feb, number =

    Understanding Deep Learning Requires Rethinking Generalization , author =. 2017 , month = feb, number =. arxiv , langid =:arXiv:1611.03530 , publisher =

  122. [130]

    Gradient Descent Optimizes Over-Parameterized Deep

    Zou, Difan and Cao, Yuan and Zhou, Dongruo and Gu, Quanquan , year =. Gradient Descent Optimizes Over-Parameterized Deep. Machine Learning , volume =. doi:10.1007/s10994-019-05839-6 , abstract =

  123. [131]

    Journal of the American Statistical Association , volume =

    Minimax and Adaptive Prediction for Functional Linear Regression , author =. Journal of the American Statistical Association , volume =. 2012 , month =. doi:10.1080/01621459.2012.716337 , url =

  124. [132]

    2010 , institution =

    Nonparametric Covariance Function Estimation for Functional and Longitudinal Data , author =. 2010 , institution =

  125. [133]

    The Annals of Statistics , volume =

    Optimal Estimation of the Mean Function Based on Discretely Sampled Functional Data: Phase Transition , author =. The Annals of Statistics , volume =. 2011 , month =. doi:10.1214/11-AOS898 , url =

  126. [134]

    The Annals of Statistics , volume =

    From Sparse to Dense Functional Data and Beyond , author =. The Annals of Statistics , volume =. 2016 , month =. doi:10.1214/16-AOS1446 , url =

  127. [135]

    Communications on Pure and Applied Analysis , volume =

    Spectral Algorithms for Functional Linear Regression , author =. Communications on Pure and Applied Analysis , volume =. 2024 , month =. doi:10.3934/cpaa.2024039 , url =

  128. [136]

    Applied and Computational Harmonic Analysis , volume =

    Optimal Rates for Functional Linear Regression with General Regularization , author =. Applied and Computational Harmonic Analysis , volume =. 2025 , month =. doi:10.1016/j.acha.2024.101745 , url =

  129. [137]

    The Annals of Statistics , volume =

    Generalized Functional Linear Models , author =. The Annals of Statistics , volume =. 2005 , month =. doi:10.1214/009053604000001156 , url =

  130. [138]

    2006 , journal =

    Properties of Principal Component Methods for Functional and Longitudinal Data Analysis , author =. 2006 , journal =. doi:10.1214/009053606000000272 , url =

  131. [139]

    The Annals of Statistics , volume =

    Methodology and Convergence Rates for Functional Linear Regression , author =. The Annals of Statistics , volume =. 2007 , month =. doi:10.1214/009053606000000957 , url =

  132. [140]

    2015 , series =

    Theoretical Foundations of Functional Data Analysis, with an Introduction to Linear Operators , author =. 2015 , series =. doi:10.1002/9781118762547 , url =

  133. [141]

    Annual Review of Statistics and Its Application , volume =

    Functional Data Analysis , author =. Annual Review of Statistics and Its Application , volume =. 2016 , month =. doi:10.1146/annurev-statistics-041715-033624 , url =

  134. [142]

    2005 , edition =

    Functional Data Analysis , author =. 2005 , edition =

  135. [143]

    Journal of the American Statistical Association , volume =

    Functional Data Analysis for Sparse Longitudinal Data , author =. Journal of the American Statistical Association , volume =. 2005 , month =. doi:10.1198/016214504000001745 , url =

  136. [144]

    The Annals of Statistics , volume =

    A Reproducing Kernel Hilbert Space Approach to Functional Linear Regression , author =. The Annals of Statistics , volume =. 2010 , month =. doi:10.1214/09-AOS772 , url =

  137. [145]

    2018 , journal =

    Adaptive Functional Linear Regression via Functional Principal Component Analysis and Block Thresholding , author =. 2018 , journal =. doi:10.5705/ss.202017.0099 , url =

  138. [146]

    Functional Linear and Single-Index Models:

    Balasubramanian, Krishnakumar and M. Functional Linear and Single-Index Models:. Bernoulli , volume =. 2025 , month =. doi:10.3150/24-BEJ1755 , url =

  139. [147]

    Journal of Multivariate Analysis , volume =

    Mean and Covariance Estimation for Discretely Observed High-Dimensional Functional Data: Rates of Convergence and Division of Observational Regimes , author =. Journal of Multivariate Analysis , volume =. 2024 , month =. doi:10.1016/j.jmva.2024.105355 , url =

  140. [148]

    2025 , journal =

    From Sparse to Dense Functional Data in High Dimensions: Revisiting Phase Transitions from a Non-Asymptotic Perspective , author =. 2025 , journal =

  141. [149]

    2021 , eprint =

    Predictive Distributions and the Transition from Sparse to Dense Functional Data , author =. 2021 , eprint =. doi:10.48550/arXiv.2109.02236 , url =

  142. [150]

    2023 , journal =

    Functional Linear Regression for Discretely Observed Data: From Ideal to Reality , author =. 2023 , journal =. doi:10.1093/biomet/asac053 , url =

  143. [151]

    2024 , month =

    Distributed Learning with Discretely Observed Functional Data , author =. 2024 , month =. 2410.02376 , archiveprefix =

  144. [152]

    The Annals of Statistics , volume =

    Prediction in Functional Linear Regression , author =. The Annals of Statistics , volume =. 2006 , month =. doi:10.1214/009053606000000830 , url =

  145. [153]

    The Annals of Statistics , volume =

    Asymptotic Equivalence of Functional Linear Regression and a White Noise Inverse Problem , author =. The Annals of Statistics , volume =. 2011 , month =. doi:10.1214/10-AOS872 , url =

  146. [154]

    2005 , journal =

    Functional Linear Regression Analysis for Longitudinal Data , author =. 2005 , journal =. doi:10.1214/009053605000000660 , url =

  147. [155]

    Journal of Multivariate Analysis , volume =

    On Rates of Convergence in Functional Linear Regression , author =. Journal of Multivariate Analysis , volume =. 2007 , month =. doi:10.1016/j.jmva.2006.10.004 , url =

  148. [156]

    Computational Statistics & Data Analysis , volume =

    Smoothing Splines Estimators in Functional Linear Regression with Errors-in-Variables , author =. Computational Statistics & Data Analysis , volume =. 2007 , month =. doi:10.1016/j.csda.2006.07.029 , url =

  149. [157]

    2012 , journal =

    Presmoothing in Functional Linear Regression , author =. 2012 , journal =. doi:10.5705/ss.2010.085 , url =

  150. [158]

    Computational Statistics & Data Analysis , volume =

    Prediction in Functional Regression with Discretely Observed and Noisy Covariates , author =. Computational Statistics & Data Analysis , volume =. 2023 , month =. doi:10.1016/j.csda.2022.107600 , url =

  151. [159]

    2007 , series =

    Random Fields and Geometry , author =. 2007 , series =

  152. [160]

    2014 , edition =

    Classical Fourier Analysis , author =. 2014 , edition =

Pith tools

Reviewed July 13, 2026 · model on record in the stance chip above.