REVIEW 2 major objections 6 minor 3 cited by
Integral Representations of Sobolev Spaces via ReLU$^k$ Activation Function and Optimal Error Estimates for Linearized Networks
T0 review · 2 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read A fixed-parameter linearized ReLU^k network attains the same optimal approximation rate in Sobolev spaces as a fully trained nonlinear shallow network.
desk verdict The main approximation-rate theorem is believable and likely correct; the alleged parity-extension contradiction in the stress-test note is misfired, but the proof still needs to spell out the extension and compactness arguments. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machine is the spherical harmonic expansion of the truncated power function $\sigma_k(t)=\max(t,0)^k$. Its Legendre coefficients $\hat\sigma_k(m)$ vanish except when $m\in\{0,\ldots,k\}$ or $m\ge k+1$ with $m-k$ odd, and satisfy $m^{d+2k+1}\hat\sigma_k(m)^2\simeq1$; this makes the integral operator $\psi\mapsto\int_{S^d}\sigma_k(\theta\cdot\eta)\psi(\theta)\,d\theta$ an isomorphism between the even/odd $L^2$ space and the even/odd Sobolev space $H^{(d+2k+1)/2}(S^d)$. To move from the sphere to the ball, the proof uses the pair of maps $S_k g(x)=|\tilde x|^k g(\tilde x/|\tilde x|)$ on the cap $G=\{\eta\in S^d:\eta_{d+1}\ge1/\sqrt2\}$ and its inverse $T_k$, which send ReLU$^k$ ridge functions to ReLU$^k$ ridge functions and are norm equivalences for $H^r$. Finite approximation on the sphere uses a scattered-point quadrature rule with weights $\tau_j\lesssim h^d$ to match all spherical-harmonic coefficients of degree up to $J\simeq h^{-1}$, and a summation-by-parts estimate on Cesàro sums of Legendre matrices to control the high-frequency tail. The tail control, combined with the coefficient estimate $\|a\|_2^2\lesssim h^{2r-2k-1}\|f\|_{H^r}^2$, produces the final rates.
What would settle it
Test the extension step numerically in the simplest case $d=1$, $k=0$, $r=1$: for a family of smooth functions on $\Omega=(-1,1)$, compute $g$ as the odd extension of $T_0f_E$ from the cap and measure $\|g\|_{H^1(S^1)}/\|f\|_{H^1(\Omega)}$; if this ratio is unbounded as $f$ varies, Theorem 2.2 fails at $k=0$. A second check for the converse direction of Theorem 2.4 is to run the least-squares fit on a fixed quasi-uniform family and test whether the normalized coefficient vectors produce piecewise-constant densities with a convergent subsequence in $L^2(S^d)$; failure of compactness would break the characterization.
Extended reading notes
Core claim
At the center is a characterization theorem: $H^{(d+2k+1)/2}(\Omega)$ coincides with the set of functions $\int_{S^d}\sigma_k(\theta\cdot\tilde x)\psi(\theta)\,d\theta$ with $\psi\in L^2(S^d)$, and the two natural norms are equivalent. The same machinery yields approximation by the finite neuron space $L^k_{n,M}$ spanned by $\sigma_k(\theta_j^*\cdot\tilde x)$ for a well-distributed point set: for $f\in H^{(d+2k+1)/2}(\Omega)$ there are coefficients $a_j$ with $\sqrt{n}\|a\|_2\lesssim\|f\|$ such that $\|f-\sum_j a_j\sigma_k(\theta_j^*\cdot\tilde x)\|_{L^2(\Omega)} \lesssim n^{-1/2-(2k+1)/(2d)}\|f\|$. More generally, $H^r$-to-$H^s$ approximation holds at rate $h^{r-s}$ for mesh size $h$, hence $n^{-(r-s)/d}$ for well-distributed points. The paper further proves the converse: a function is approximable by such fixed-parameter linear spaces, with uniformly bounded coefficients and quasi-uniform parameters, if and only if it lies in $H^{(d+2k+1)/2}(\Omega)$. Consequently the Sobolev space is the RKHS of the kernel $\int_{S^d}\sigma_k(\theta\cdot x)\sigma_k(\theta\cdot y)\,d\theta$, random feature sampling reaches the same rate up to a log factor, and an $O(m^{-1/2})$ generalization bound in $H^1$ follows for an elliptic PDE loss.
Load-bearing premise
The proof of Theorem 2.2 assumes an unproved extension step: after lifting the Sobolev function to the cap $G$ of the sphere, a function $g$ must exist on the whole sphere with the parity $g(\eta)=(-1)^{k+1}g(-\eta)$ and with $\|g\|_{H^r(S^d)}\le C\|f\|_{H^r(\Omega)}$; if such an extension fails for some Lipschitz domain or fractional $r$, the main approximation theorem collapses.
Editorial extensions
If this is right
- For every $f\in H^{(d+2k+1)/2}(\Omega)$ and any well-distributed parameter mesh, the finite neuron space reaches $O(n^{-1/2-(2k+1)/(2d)})$ in $L^2$, the same rate known for fully nonlinear shallow networks up to logarithmic factors.
- Because only output coefficients are trained, the optimal rate is achieved by solving a convex least-squares problem rather than a nonconvex parameter optimization.
- Random i.i.d. parameters on $S^d$ recover the same rate up to a log factor with probability at least $1-\delta$, which is why random feature methods work; deterministic well-distributed points remove the log factor and the failure probability.
- The coefficient bound $\sqrt{n}\|a\|_2\lesssim\|f\|$ yields an expected $H^1$ generalization error of $O(m^{-1/2})$ for the empirical risk minimizer of an elliptic PDE loss.
- The characterization also implies $H^{(d+2k+1)/2}(\Omega)\hookrightarrow B_k(\Omega)$ and that the unit balls of the two spaces have metric entropies of the same order.
Reading between the lines
- Because the proof relies only on the Legendre-coefficient decay of $\sigma_k$, the same linearized construction should extend to other ridge activations with comparable coefficient decay, with the exponent adjusted accordingly.
- Since $H^{(d+2k+1)/2}(\Omega)$ is shown to be the RKHS of the ReLU$^k$ kernel, standard kernel-based generalization bounds apply without further work; the paper does not exploit this connection.
- The comparison with finite elements suggests a direct numerical benchmark: for a smooth high-dimensional target, plotting least-squares error against $n$ for fixed-parameter ReLU$^k$ spaces versus degree-$k$ FEM should show the predicted polynomial gap.
- The contrast with Barron spaces implies that functions in $B_k(\Omega)\setminus H^{(d+2k+1)/2}(\Omega)$ are essentially not approximable at the same rate by fixed-parameter linear spaces; constructing such a function explicitly would mark the boundary between linear and nonlinear approximation.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies approximation properties of linearized shallow ReLU^k networks with fixed, well-distributed inner parameters. Its main results are: (i) Theorem 2.2, an approximation estimate showing that finite neuron spaces L^k_{n,M} achieve the rate O(n^{-1/2-(2k+1)/(2d)}) in L2 for functions in H^{(d+2k+1)/2}(Ω), with a coefficient bound M comparable to the Sobolev norm; (ii) Theorem 2.3, an integral representation of H^{(d+2k+1)/2}(Ω) by L2-weighted integrals of ReLU^k ridge functions, with equivalent norms; and (iii) corollaries including an embedding of this Sobolev space into the Barron space, an RKHS interpretation, and a generalization bound. The proofs transfer harmonic-analytic approximation results from the sphere to the ball through an operator T_k mapping functions on the ball to a cap of the sphere.
Significance. If the proofs are completed, the results are significant: they show that linearized networks (convex optimization over fixed weights) match the optimal approximation rates of nonlinear shallow networks on Sobolev spaces, give an explicit L2 integral representation of Sobolev spaces via ReLU^k ridge functions, and provide coefficient bounds that are useful for generalization analysis. The approximation-rate claim is falsifiable and would explain the empirical success of random feature methods as a consequence of the deterministic well-distributed nature of the parameters. The paper also reproduces the embedding H^{(d+2k+1)/2}(Ω) into B^k(Ω) and clarifies the quotient structure between Barron and Sobolev spaces. The main technical gaps, however, are two unproved load-bearing steps that must be supplied before the central claims are fully established.
major comments (2)
- [Section 5.1, proof of Theorem 2.2] The step 'By the extension theorems again, there exists a function g on S^d, which is an extension of T_k f_E, such that ||g||_{H^r(S^d)} ≲ ||T_k f_E||_{H^r(G)} ≲ ||f||_{H^r(Ω)} and g(η) = (-1)^{k+1}g(-η)' is asserted without proof or reference. This parity-preserving extension is load-bearing: Theorem 4.1 requires the parity condition (4.1), and this is the only bridge from the spherical approximation result to the ball. Standard Sobolev extension operators do not preserve exact parity, and symmetrization after extension changes the values on G unless additional construction is used. Because G and -G are disjoint, a correct construction is possible: define g on -G by g(-η)=(-1)^{k+1}T_k f_E(η), apply a bounded extension operator from the disjoint union G∪(-G) to S^d, and then use a smooth cutoff to adjust the equatorial belt while preserving the parity. However, this argument is not provided, and the current manuscript therefore leaves a gap in the proofs of both Theorem 2.2 and Theorem 2.3, which rely on the same asserted extension.
- [Section 2.2, proof of Theorem 2.4 (converse direction)] The construction of piecewise-constant densities ψ_n is only sketched, and the claim 'One could verify the limit of this subsequence is the desired function ψ' omits the key compactness argument. To complete the proof one must show that the weak limit ψ satisfies f(x) = ∫_{S^d} σ_k(θ·x̃)ψ(θ)dθ. This requires two facts: (i) the integral operator G: L^2(S^d) → L^2(Ω) is compact (it is Hilbert–Schmidt on a bounded domain), so Gψ_n → Gψ strongly whenever ψ_n → ψ weakly; and (ii) Gψ_n differs from the neural-network approximant f_n by an error of size O(n^{-1/d} ∑_j |a_j^{(n)}|) ≤ O(M n^{-1/d}) using the quasi-uniform partition of S^d and Lipschitz continuity of θ ↦ σ_k(θ·x̃). Since this direction is part of the claimed characterization in Theorem 2.4, the argument should be included explicitly.
minor comments (6)
- [Section 5.2, proof of Theorem 2.3] The notation 'T_k f' is used for f ∈ H^{(d+2k+1)/2}(Ω) without first extending f to B^d; the representation formula defines an extension implicitly, but this should be stated explicitly to avoid confusion.
- [References] References [83] and [84] are identical (both cite the same paper by Siegel, Hong, Jin, Hao, and Xu in Journal of Computational Physics, 484:112084, 2023) and should be de-duplicated.
- [Corollary 2.1] In the definition of the normalization map P, the domain is written as S^{d-1} × [1,1] but should be S^{d-1} × [-1,1]; the same display also contains a minor typo in the final bracket.
- [Section 6] The word 'follwoing' in the sentence 'Standard concentration inequalities applied to the integral representation (1.16) yield the follwoing theorem' should be 'following'.
- [Lemma 3.4(c)] In the proof, the set I_{-1,j} is defined as {i : ρ(θ*_i, -θ*_j) < h̄}, but later the notation I_{-1,j,-} appears without definition; this is a typographical slip.
- [Lemma 3.1] The phrase 'Some calculus estimation yields' in the proof of the sign of the polygamma sums is vague; adding a short justification or a reference would improve reproducibility.
Circularity Check
No significant circularity; the FNS rate and integral representation are derived from explicit Legendre/harmonic analysis, with self-citations only contextual; an unproved parity-preserving extension in Section 5.1 is a correctness gap, not a circular step.
full rationale
The main approximation theorem (Theorem 2.2, Eqs. (2.12)-(2.14)) is proved in Section 5.1 by transferring a spherical approximation theorem (Theorem 4.1) to the ball. Theorem 4.1 is self-contained: it uses the explicit Legendre coefficients of ReLU^k (Eq. (3.10), Lemma 3.1), a positive scattered-point quadrature rule (Lemma 3.2, quoted from [64]), and spectral estimates for the matrices P_beta(m) (Lemma 3.4). No target rate or Barron-space result is fed into the proof. The integral representation (Theorem 2.3, Eqs. (2.20)-(2.22)) follows from the spherical isomorphism Theorem 4.2, where the L2(S^d) to H^{(d+2k+1)/2}(S^d) equivalence is a direct multiplier computation using \hat sigma_k(m)^2 m^{d+2k+1} \simeq 1; this is not a renaming of the desired statement. The embedding H^s -> B^k, credited to the same group's [58], is re-proved as Corollary 2.2 from Theorem 2.2 rather than assumed, so [58] is not load-bearing. The lower bound (2.7) from [87] is used only to assert optimality and is an independent published result; the Barron-space surjectivity from [88] is contextual. Accordingly, self-citations are present but do not force the conclusions. One genuine concern, flagged per the reviewing rule, is a proof gap: in Section 5.1 the authors assert "By the extension theorems again, there exists a function g on S^d... g(eta)=(-1)^{k+1} g(-eta)" without proof. Exact antipodal parity is not a standard consequence of Sobolev extension theorems and may fail on the cap boundary. This threatens correctness of the ball-to-sphere transfer, but it is an omitted justification, not an instance of the paper's derivation reducing to its own inputs.
Assumptions & free parameters
assumptions (5)
- standard math Quadrature rule for scattered points on S^d with nonnegative weights and exactness up to degree O(h^{-1}) (Lemma 3.2, from Mhaskar-Narcowich-Ward).
- domain assumption Parity-preserving Sobolev extension from the cap G subset of S^d to the whole sphere with norm control, used in the proof of Theorem 2.2.
- standard math Standard extension theory for Sobolev spaces on bounded Lipschitz domains (used to extend f in H^r(Omega) to H^r(R^d)).
- standard math Legendre-coefficient formula for sigma_k with support E_sigma_k and the decay bounds from [2, Appendix D.2].
- standard math Existence of well-distributed and quasi-uniform point sets on S^d with mesh norm h comparable to n^{-1/d}.
Cite this review
Pith. "Pith review of Integral Representations of Sobolev Spaces via ReLU$^k$ Activation Function and Optimal Error Estimates for Linearized Networks." pith.science (2026). https://pith.science/paper/6GJ4YDJE
@misc{pith2026250500351,
author = {Pith},
title = {Pith review of: Integral Representations of Sobolev Spaces via ReLU$^k$ Activation Function and Optimal Error Estimates for Linearized Networks},
year = {2026},
howpublished = {\url{https://pith.science/paper/6GJ4YDJE}},
note = {Machine review of arXiv:2505.00351}
}
abstract
This paper presents two main theoretical results concerning shallow neural networks with ReLU$^k$ activation functions. We establish a novel integral representation for Sobolev spaces, showing that every function in $H^{\frac{d+2k+1}{2}}(\Omega)$ can be expressed as an $L^2$-weighted integral of ReLU$^k$ ridge functions over the unit sphere. This result mirrors the known representation of Barron spaces and highlights a fundamental connection between Sobolev regularity and neural network representations. Moreover, we prove that linearized shallow networks -- constructed by fixed inner parameters and optimizing only the linear coefficients -- achieve optimal approximation rates $O(n^{-\frac{1}{2}-\frac{2k+1}{2d}})$ in Sobolev spaces.
Forward citations
Cited by 3 Pith papers
-
Optimal Neural Network Approximation via Empirical Least Squares with Deterministic Samples
Linearized ReLU^k discrete least squares on the sphere achieves the optimal rate n^{-(r-s)/d} with m roughly n deterministic samples under a parity condition and quasi-uniform points.
-
ReLU$^k$ Neural de Rham Complexes
Fixed-neuron ReLU^k ridge functions form an exact finite-dimensional de Rham complex in any dimension, provided the lowest-order family is linearly independent.
-
Do Neural Networks Really Beat the Curse of Dimensionality? A Bit-Complexity View
When approximation quality is measured per bit instead of per parameter, neural networks do not fundamentally beat classical methods; the real limit is the metric entropy of the target function class.
Reference graph
Works this paper leans on
-
[1]
Support vector machines
Mamoun Awad and Latifur Khan. Support vector machines. In Intelligent Information Technologies: Concepts, Methodologies, Tools, and Applications, pages 1138–1146. IGI Global, 2008
2008
-
[2]
Breaking the curse of dimensionality with convex ne ural networks
Francis Bach. Breaking the curse of dimensionality with convex ne ural networks. The Journal of Machine Learning Research , 18(1):629–681, 2017
2017
-
[3]
On the equivalence between kernel quadrature r ules and random feature ex- pansions
Francis Bach. On the equivalence between kernel quadrature r ules and random feature ex- pansions. Journal of machine learning research , 18(21):1–38, 2017
2017
-
[4]
Universal approximation bounds for superpo sitions of a sigmoidal function
Andrew R Barron. Universal approximation bounds for superpo sitions of a sigmoidal function. IEEE Transactions on Information theory , 39(3):930–945, 1993
1993
-
[5]
Approximation and estimation bounds for artifi cial neural networks
Andrew R Barron. Approximation and estimation bounds for artifi cial neural networks. Machine learning , 14:115–133, 1994
1994
-
[6]
Approximation and learning by greedy algorithms
Andrew R Barron, Albert Cohen, Wolfgang Dahmen, and Ronald A D eVore. Approximation and learning by greedy algorithms. Annals of Statistics , 36(1):64–94, 2008
2008
-
[7]
Nearly-tight vc- dimension and pseudodimension bounds for piecewise linear neural ne tworks
Peter L Bartlett, Nick Harvey, Christopher Liaw, and Abbas Meh rabian. Nearly-tight vc- dimension and pseudodimension bounds for piecewise linear neural ne tworks. The Journal of Machine Learning Research, 20(1):2285–2301, 2019
2019
-
[8]
Bartlett and Shahar Mendelson
Peter L. Bartlett and Shahar Mendelson. Rademacher and gaus sian complexities: Risk bounds and structural results. Journal of Machine Learning Research , 3:463–482, 2002
2002
Show all 99 references
-
[9]
Two models of double descent for weak features
Mikhail Belkin, Daniel Hsu, and Ji Xu. Two models of double descent for weak features. SIAM Journal on Mathematics of Data Science , 2(4):1167–1180, 2020
2020
-
[10]
Springer Science & Business Media, 2012
J¨ oran Bergh and J¨ orgen L¨ ofstr¨ om.Interpolation spaces: an introduction, volume 223. Springer Science & Business Media, 2012
2012
-
[11]
Optimal asymptotic bounds for spherical designs
Andriy Bondarenko, Danylo Radchenko, and Maryna Viazovska . Optimal asymptotic bounds for spherical designs. Annals of mathematics , pages 443–452, 2013
2013
-
[12]
C oncentration inequalities
St´ ephane Boucheron, G´ abor Lugosi, and Olivier Bousquet. C oncentration inequalities. In Summer school on machine learning , pages 208–240. Springer, 2003
2003
-
[13]
Projection bodies
Jean Bourgain and Joram Lindenstrauss. Projection bodies. I n Geometric Aspects of Func- tional Analysis: Israel Seminar (GAF A) 1986–87 , pages 250–270. Springer, 2006. 35
1986
-
[14]
Weak type estimates for Cesaro sums of Jacobi polynomial series , volume 487
Sagun Chanillo and Benjamin Muckenhoupt. Weak type estimates for Cesaro sums of Jacobi polynomial series , volume 487. American Mathematical Soc., 1993
1993
-
[15]
Bridgin g traditional and machine learning-based algorithms for solving pdes: the random feature me thod
Jingrun Chen, Xurong Chi, Weinan E, and Zhouwang Yang. Bridgin g traditional and machine learning-based algorithms for solving pdes: the random feature me thod. J Mach Learn, 1:268– 98, 2022
2022
-
[16]
The random fea ture method for solving interface problems
Xurong Chi, Jingrun Chen, and Zhouwang Yang. The random fea ture method for solving interface problems. Computer Methods in Applied Mechanics and Engineering , 420:116719, 2024
2024
-
[17]
Philippe G. Ciarlet. The Finite Element Method for Elliptic Problems , volume 4 of Studies in Mathematics and its Applications . North-Holland, 1978
1978
-
[18]
Learning theory: an approximation theory viewpoint , volume 24
Felipe Cucker and Ding Xuan Zhou. Learning theory: an approximation theory viewpoint , volume 24. Cambridge University Press, 2007
2007
-
[19]
Approximation by superpositions of a sigmoida l function
George Cybenko. Approximation by superpositions of a sigmoida l function. Mathematics of control, signals and systems , 2(4):303–314, 1989
1989
-
[20]
Approximation theory and harmonic analysis on spheres and b alls
Feng Dai. Approximation theory and harmonic analysis on spheres and b alls. Springer, 2013
2013
-
[21]
Local randomized neural network s with hybridized discontin- uous petrov–galerkin methods for stokes–darcy flows
Haoning Dang and Fei Wang. Local randomized neural network s with hybridized discontin- uous petrov–galerkin methods for stokes–darcy flows. Physics of Fluids , 36(8), 2024
2024
-
[22]
Neural ne twork approximation
Ronald DeVore, Boris Hanin, and Guergana Petrova. Neural ne twork approximation. Acta Numerica, 30:327–444, 2021
2021
-
[23]
Nonlinear approximation
Ronald A DeVore. Nonlinear approximation. Acta Numerica, 7:51–150, 1998
1998
-
[24]
Constructive approximation, volume 303
Ronald A DeVore and George G Lorentz. Constructive approximation, volume 303. Springer Science & Business Media, 1993
1993
-
[25]
Some remarks on greed y algorithms
Ronald A DeVore and Vladimir N Temlyakov. Some remarks on greed y algorithms. Advances in computational Mathematics , 5(1):173–187, 1996
1996
-
[26]
Local extreme learning machines and domain decomposition for solving linear and nonlinear partial differential equations
Suchuan Dong and Zongwei Li. Local extreme learning machines and domain decomposition for solving linear and nonlinear partial differential equations. Computer Methods in Applied Mechanics and Engineering , 387:114129, 2021
2021
-
[27]
Physics informed extreme le arning machine (pielm)–a rapid method for the numerical solution of partial differential equa tions
Vikas Dwivedi and Balaji Srinivasan. Physics informed extreme le arning machine (pielm)–a rapid method for the numerical solution of partial differential equa tions. Neurocomputing, 391:96–118, 2020
2020
-
[28]
A priori estimates of the populat ion risk for two-layer neural networks
Weinan E, Chao Ma, and Lei Wu. A priori estimates of the populat ion risk for two-layer neural networks. Communications in Mathematical Sciences , 17(5):1407–1425, 2019
2019
-
[29]
The barron space and the flow-in duced function spaces for neural network models
Weinan E, Chao Ma, and Lei Wu. The barron space and the flow-in duced function spaces for neural network models. Constructive Approximation , 55:259–292, 2022
2022
-
[30]
Representation formulas and pointwise properties for barron functions
Weinan E and Stephan Wojtowytsch. Representation formulas and pointwise properties for barron functions. Calculus of Variations and Partial Differential Equations , 61(2):46, 2022. 36
2022
-
[31]
Generalisation error in learning with random features and the hidden manifold model
Federica Gerace, Bruno Loureiro, Florent Krzakala, Marc M´ e zard, and Lenka Zdeborov´ a. Generalisation error in learning with random features and the hidden manifold model. In International Conference on Machine Learning , pages 3452–3462. PMLR, 2020
2020
-
[32]
Deep neural networks with random gaussian weights: A universal classification strategy? IEEE Transactions on Signal Process- ing, 64(13):3444–3457, 2016
Raja Giryes, Guillermo Sapiro, and Alex M Bronstein. Deep neural networks with random gaussian weights: A universal classification strategy? IEEE Transactions on Signal Process- ing, 64(13):3444–3457, 2016
2016
-
[33]
Delving de ep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Delving de ep into rectifiers: Surpassing human-level performance on imagenet classification. I n Proceedings of the IEEE international conference on computer vision , pages 1026–1034, 2015
2015
-
[34]
Multilayer feedforward networks are universal approximators
Kurt Hornik, Maxwell Stinchcombe, and Halbert White. Multilayer feedforward networks are universal approximators. Neural networks , 2(5):359–366, 1989
1989
-
[35]
Universality laws for high-dimensional learn ing with random features
Hong Hu and Yue M Lu. Universality laws for high-dimensional learn ing with random features. IEEE Transactions on Information Theory , 69(3):1932–1964, 2022
1932
-
[36]
Universal approximation using incre- mental constructive feedforward networks with random hidden n odes
Guang-Bin Huang, Lei Chen, and Chee-Kheong Siew. Universal approximation using incre- mental constructive feedforward networks with random hidden n odes. IEEE transactions on neural networks , 17(4):879–892, 2006
2006
-
[37]
Extreme learning machine: theory and applications
Guang-Bin Huang, Qin-Yu Zhu, and Chee-Kheong Siew. Extreme learning machine: theory and applications. Neurocomputing, 70:489–501, 2006
2006
-
[38]
Stochastic choice of basis funct ions in adaptive function ap- proximation and the functional-link net
Boris Igelnik and Yoh-Han Pao. Stochastic choice of basis funct ions in adaptive function ap- proximation and the functional-link net. IEEE Transactions on Neural Networks , 6(6):1320– 1329, 1995
1995
-
[39]
Norming sets and spherica l cubature formulas
K Jetter, J St¨ ockler, and JD Ward. Norming sets and spherica l cubature formulas. In Advances in computational mathematics , pages 237–244. CRC Press, 2023
2023
-
[40]
A simple lemma on greedy approximation in hilbert spac e and convergence rates for projection pursuit regression and neural network tra ining
Lee K Jones. A simple lemma on greedy approximation in hilbert spac e and convergence rates for projection pursuit regression and neural network tra ining. The Annals of Statistics , 20(1):608–613, 1992
1992
-
[41]
Approximation by comb inations of relu and squared relu ridge functions with l1 and l0 controls
Jason M Klusowski and Andrew R Barron. Approximation by comb inations of relu and squared relu ridge functions with l1 and l0 controls. IEEE Transactions on Information Theory, 64(12):7649–7656, 2018
2018
-
[42]
On linear dimensionality of topolog ical vector spaces
Andrei Nikolaevich Kolmogorov. On linear dimensionality of topolog ical vector spaces. In Doklady Akademii Nauk , volume 120(2), pages 239–241. Russian Academy of Sciences, 19 58
-
[43]
Some problems in the theory of ridge functions
Sergei Vladimirovich Konyagin, Aleksandr Andreevich Kuleshov, and Vitalii Evgen’evich Maiorov. Some problems in the theory of ridge functions. Proceedings of the Steklov In- stitute of Mathematics , 301:144–169, 2018
2018
-
[44]
K˚ urkov´ a and M
V. K˚ urkov´ a and M. Sanguineti. Bounds on rates of variable basis and neural network approx- imation. IEEE Transactions on Information Theory , 47(6):2659–2665, 2001
2001
-
[45]
K˚ urkov´ a and M
V. K˚ urkov´ a and M. Sanguineti. Comparison of worst case errors in linear and neural network approximation. IEEE Transactions on Information Theory , 48(1):264–275, 2002. 37
2002
-
[46]
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. Deep learning . Nature, 521(7553):436– 444, 2015
2015
-
[47]
Multilayer feedforward networks with a nonpolynomial activation function can approximate any function
Moshe Leshno, Vladimir Ya Lin, Allan Pinkus, and Shimon Schocken. Multilayer feedforward networks with a nonpolynomial activation function can approximate any function. Neural networks, 6(6):861–867, 1993
1993
-
[48]
Approximation of funct ions of finite variation by superpositions of a sigmoidal function
Grzegorz Lewicki and Giuseppe Marino. Approximation of funct ions of finite variation by superpositions of a sigmoidal function. Applied Mathematics Letters , 17(10):1147–1152, 2004
2004
-
[49]
Towar ds a unified analysis of random fourier features
Zhu Li, Jean-Francois Ton, Dino Oglic, and Dino Sejdinovic. Towar ds a unified analysis of random fourier features. In International conference on machine learning , pages 3905–3914. PMLR, 2019
2019
-
[50]
Lower bounds of the discretiz ation error for piecewise polynomials
Qun Lin, Hehu Xie, and Jinchao Xu. Lower bounds of the discretiz ation error for piecewise polynomials. Mathematics of Computation , 83(285):1–13, 2014
2014
-
[51]
Is extreme lea rning machine feasible? a theoretical assessment (part 1)
Xiang Liu, Shouyi Lin, Jun Fang, and Zongben Xu. Is extreme lea rning machine feasible? a theoretical assessment (part 1). IEEE Transactions on Neural Networks and Learning Systems, 26(1):7–20, 2014
2014
-
[52]
Randomized nonlinear component analysis
David Lopez-Paz, Suvrit Sra, Alex Smola, Zoubin Ghahramani, an d Bernhard Sch¨ olkopf. Randomized nonlinear component analysis. In International conference on machine learning , pages 1359–1367. PMLR, 2014
2014
-
[53]
Dee p neural networks with fixed width can be universal approximators
Zuowei Lu, Zhiqiang Shen, Haizhao Yang, and Shijun Zhang. Dee p neural networks with fixed width can be universal approximators. Neural Networks , 141:202–211, 2021
2021
-
[54]
Uniform approximatio n rates and metric entropy of shallow neural networks
Limin Ma, Jonathan W Siegel, and Jinchao Xu. Uniform approximatio n rates and metric entropy of shallow neural networks. Research in the Mathematical Sciences , 9(3):46, 2022
2022
-
[55]
On the near optimality of the stocha stic approximation of smooth functions by neural networks
Vitaly E Maiorov and Ron Meir. On the near optimality of the stocha stic approximation of smooth functions by neural networks. Advances in Computational Mathematics , 13:79–103, 2000
2000
-
[56]
Random approximants and neural networks
Yuly Makovoz. Random approximants and neural networks. Journal of Approximation The- ory, 85(1):98–109, 1996
1996
-
[57]
Uniform approximation by neural networks
Yuly Makovoz. Uniform approximation by neural networks. Journal of Approximation Theory, 95(2):215–228, 1998
1998
-
[58]
Approximation rat es for shallow reluk neural networks on sobolev spaces via the radon transform
Tong Mao, Jonathan W Siegel, and Jinchao Xu. Approximation rat es for shallow reluk neural networks on sobolev spaces via the radon transform. arXiv preprint arXiv:2408.10996 , 2024
2024
-
[59]
Do neural networks have better app roximation properties than polynomials or finite elements for high-dimensional problems? preprint, 2025
Tong Mao and Jinchao Xu. Do neural networks have better app roximation properties than polynomials or finite elements for high-dimensional problems? preprint, 2025
2025
-
[60]
Rates of approximation by relu sh allow neural networks
Tong Mao and Ding-Xuan Zhou. Rates of approximation by relu sh allow neural networks. Journal of Complexity , 79:101784, 2023
2023
-
[61]
Type et cotype dans les espaces munis de structure s locales inconditionnelles
B Maurey. Type et cotype dans les espaces munis de structure s locales inconditionnelles. Seminaire Maurey-Schwartz, pages 1–25, 1973. 38
1973
-
[62]
The generalization error of ra ndom features regression: Precise asymptotics and the double descent curve
Song Mei and Andrea Montanari. The generalization error of ra ndom features regression: Precise asymptotics and the double descent curve. Communications on Pure and Applied Mathematics, 75(4):667–766, 2022
2022
-
[63]
A new function space from barron cla ss and application to neural network approximation
Yan Meng and Pingbing Ming. A new function space from barron cla ss and application to neural network approximation. Communications in Computational Physics , 32(5):1361–1400, 2022
2022
-
[64]
Spherical marcinkiewicz-z ygmund inequalities and positive quadrature
H Mhaskar, F Narcowich, and J Ward. Spherical marcinkiewicz-z ygmund inequalities and positive quadrature. Mathematics of computation , 70(235):1113–1130, 2001
2001
-
[65]
Tractability of approximatio n by general shallow net- works
Hrushikesh Mhaskar and Tong Mao. Tractability of approximatio n by general shallow net- works. arXiv preprint arXiv:2308.03230 , 2023
2023 arXiv
-
[66]
Eignets for function approximation on m anifolds
Hrushikesh N Mhaskar. Eignets for function approximation on m anifolds. Applied and Com- putational Harmonic Analysis , 29(1):63–87, 2010
2010
-
[67]
Kernel-based analysis of massive data
Hrushikesh N Mhaskar. Kernel-based analysis of massive data. Frontiers in Applied Mathe- matics and Statistics , 6:30, 2020
2020
-
[68]
Approximation properties of a mu ltilayered feedforward artificial neural network
Hrushikesh Narhar Mhaskar. Approximation properties of a mu ltilayered feedforward artificial neural network. Advances in Computational Mathematics , 1:61–80, 1993
1993
-
[69]
Weighted quadrature formulas a nd approximation by zonal function networks on the sphere
Hrushikesh Narhar Mhaskar. Weighted quadrature formulas a nd approximation by zonal function networks on the sphere. Journal of Complexity , 22(3):348–370, 2006
2006
-
[70]
Foundations of Machine Learn- ing
Mehryar Mohri, Afshin Rostamizadeh, and Ameet Talwalkar. Foundations of Machine Learn- ing. MIT Press, 2018
2018
-
[71]
The random feature mod el for input-output maps between banach spaces
Nicholas H Nelsen and Andrew M Stuart. The random feature mod el for input-output maps between banach spaces. SIAM Journal on Scientific Computing , 43(5):A3212–A3243, 2021
2021
-
[72]
Learnin g and generalization charac- teristics of the random vector functional-link net
Yoh-Han Pao, Gwang-Hoon Park, and Dejan J Sobajic. Learnin g and generalization charac- teristics of the random vector functional-link net. Neurocomputing, 6(2):163–180, 1994
1994
-
[73]
Approximation by ridge functions and neu ral networks
Pencho P Petrushev. Approximation by ridge functions and neu ral networks. SIAM Journal on Mathematical Analysis , 30(1):155–189, 1998
1998
-
[74]
Approximation theory of the mlp model in neural net works
Allan Pinkus. Approximation theory of the mlp model in neural net works. Acta numerica , 8:143–195, 1999
1999
-
[75]
Remarques sur un r´ esultat non publi´ e de B
Gilles Pisier. Remarques sur un r´ esultat non publi´ e de B. Maure y. S´ eminaire d’Analyse fonctionnelle (dit” Maurey-Schwartz”) , pages 1–12, 1981
1981
-
[76]
Random features for large-sca le kernel machines
Ali Rahimi and Benjamin Recht. Random features for large-sca le kernel machines. In Advances in Neural Information Processing Systems , pages 1177–1184, 2007
2007
-
[77]
Uniform approximation of functio ns with random bases
Ali Rahimi and Benjamin Recht. Uniform approximation of functio ns with random bases. In 2008 46th annual allerton conference on communication, con trol, and computing , pages 555–561. IEEE, 2008. 39
2008
-
[78]
Weighted sums of random kitchen sinks: Replacing mini- mization with randomization in learning
Ali Rahimi and Benjamin Recht. Weighted sums of random kitchen sinks: Replacing mini- mization with randomization in learning. Advances in neural information processing systems , 21, 2008
2008
-
[79]
On random weights and unsupervised feature learning
Andrew M Saxe, Pang Wei Koh, Zhenghao Chen, Maneesh Bhand , Bipin Suresh, and An- drew Y Ng. On random weights and unsupervised feature learning. I n Icml, volume 2, page 6, 2011
2011
-
[80]
Zu einem problem von shephard ¨ uber die projek tionen konvexer k¨ orper
Rolf Schneider. Zu einem problem von shephard ¨ uber die projek tionen konvexer k¨ orper. Mathematische Zeitschrift , 101:71–82, 1967
1967
-
[81]
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David. Understanding machine learning: From theory to algorithms. Cambridge university press, 2014
2014
-
[82]
Optimal approximation rates for deep relu ne ural networks on sobolev and besov spaces
Jonathan W Siegel. Optimal approximation rates for deep relu ne ural networks on sobolev and besov spaces. Journal of Machine Learning Research , 24(1):1–52, 2023
2023
-
[84]
Greedy training algorithms for neural networks and applications to pdes
Jonathan W Siegel, Qingguo Hong, Xianlin Jin, Wenrui Hao, and Jinc hao Xu. Greedy training algorithms for neural networks and applications to pdes. Journal of Computational Physics , 484:112084, 2023
2023
-
[85]
High-order approximation ra tes for shallow neural net- works with cosine and ReLUk activation functions
Jonathan W Siegel and Jinchao Xu. High-order approximation ra tes for shallow neural net- works with cosine and ReLUk activation functions. Applied and Computational Harmonic Analysis, 58:1–26, 2022
2022
-
[86]
Optimal convergence rates for the orthogonal greedy algorithm
Jonathan W Siegel and Jinchao Xu. Optimal convergence rates for the orthogonal greedy algorithm. IEEE Transactions on Information Theory , 68(5):3354–3361, 2022
2022
-
[87]
Sharp bounds on the approx imation rates, metric entropy, and n-widths of shallow neural networks
Jonathan W Siegel and Jinchao Xu. Sharp bounds on the approx imation rates, metric entropy, and n-widths of shallow neural networks. Foundations of Computational Mathematics , pages 1–57, 2022
2022
-
[88]
Characterization of the var iation spaces corresponding to shallow neural networks
Jonathan W Siegel and Jinchao Xu. Characterization of the var iation spaces corresponding to shallow neural networks. Constructive Approximation , 57(3):1109–1132, 2023
2023
-
[89]
Singular integrals and differentiability properties of fun ctions
Elias M Stein. Singular integrals and differentiability properties of fun ctions. Princeton university press, 1970
1970
-
[90]
Introduction to Fourier analysis on Euclidean spaces , vol- ume 1
Elias M Stein and Guido Weiss. Introduction to Fourier analysis on Euclidean spaces , vol- ume 1. Princeton university press, 1971
1971
-
[91]
Szeg¨ o.Orthogonal polynomials, volume 23 of Amer
G. Szeg¨ o.Orthogonal polynomials, volume 23 of Amer. Math. Soc. Colloq. Publ. Amer. Math. Soc., Providence, 1975
1975
-
[92]
Greedy approximation
Vladimir N Temlyakov. Greedy approximation. Acta Numerica, 17:235–409, 2008
2008
-
[93]
Wainwright
Martin J. Wainwright. High-dimensional Statistics: A Non-asymptotic Viewpoint . Cambridge Series in Statistical and Probabilistic Mathematics. Cambridge Univer sity Press, 2019. 40
2019
-
[94]
An extreme learning machine-bas ed method for computa- tional pdes in higher dimensions
Yiran Wang and Suchuan Dong. An extreme learning machine-bas ed method for computa- tional pdes in higher dimensions. Computer Methods in Applied Mechanics and Engineering , 418:116578, 2024
2024
-
[95]
Iterative methods by space decomposition and sub space correction
Jinchao Xu. Iterative methods by space decomposition and sub space correction. SIAM review, 34(4):581–613, 1992
1992
-
[96]
Finite neuron method and convergence analysis
Jinchao Xu. Finite neuron method and convergence analysis. Communications in Computa- tional Physics , 28(5):1707–1745, 2020
2020
-
[97]
Efficient and provably convergent r andomized greedy algorithms for neural network optimization
Jinchao Xu and Xiaofeng Xu. Efficient and provably convergent r andomized greedy algorithms for neural network optimization. arXiv preprint arXiv:2407.17763 , 2024
2024 arXiv
-
[98]
Optimal rates of approximat ion by shallow relu k neural networks and applications to nonparametric regression
Yunfei Yang and Ding-Xuan Zhou. Optimal rates of approximat ion by shallow relu k neural networks and applications to nonparametric regression. Constructive Approximation , pages 1–32, 2024
2024
-
[99]
Sup -norm approximation bounds for networks through probabilistic methods
Joseph E Yukich, Maxwell B Stinchcombe, and Halbert White. Sup -norm approximation bounds for networks through probabilistic methods. IEEE Transactions on Information The- ory, 41(4):1021–1027, 1995
1995
-
[100]
Trans ferable neural networks for partial differential equations
Zezhong Zhang, Feng Bao, Lili Ju, and Guannan Zhang. Trans ferable neural networks for partial differential equations. Journal of Scientific Computing , 99(1):2, 2024. 41
2024
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.