Pith. sign in

REVIEW 3 major objections 5 minor 57 references

Every vector-valued RKBS belongs to an adjoint pair with a reproducing kernel, and optimizing over one recovers shallow networks, DeepONets, and hypernetworks.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Vector-valued neural networks, DeepONets, and hypernetworks are shown to live in integral vector-valued reproducing kernel Banach spaces with representer theorems that recover the architectures.

T0 review reviewed 2026-08-04 challenge →

load-bearing objection Solid extension of vv-RKBS theory with a load-bearing gap: the sparse representer theorem assumes an extreme-point characterization of vector measures that likely needs Radon-Nikodym hypotheses. the 3 major comments →

arxiv 2509.26371 v3 pith:QXZSW4ZS submitted 2025-09-30 math.FA cs.AIcs.LGstat.ML

Vector-Valued Reproducing Kernel Banach Spaces for Neural Networks and Operators

classification math.FA cs.AIcs.LGstat.ML MSC 46E1568T0746G1046E2246B1026B40
keywords reproducing kernel Banach spacevector-valued RKBStwin operatorsintegral RKBSrepresenter theoremneural networksDeepONethypernetwork
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Vector-valued neural networks and neural operators have been hard to place in reproducing kernel theory because their outputs live in Banach spaces that may not be reflexive or separable. The paper claims a fix: define an adjoint pair of vector-valued reproducing kernel Banach spaces whose kernel K:X×Ω→Twin(U,U^⋄) is built from twin operators, and prove every vv-RKBS belongs to such a pair. It specializes to integral vv-RKBSs, where a function is an integral of a scalar feature φ against a vector-valued measure, and proves a representer theorem: the minimizer of a convex supervised loss with total-variation regularization is a sum of Nd terms a_m φ(·,w_m)u_m with u_m extreme points of the output unit ball. With neural features this representation is exactly a shallow R^d-valued network Uσ(Wx+B), and with two coupled features it is exactly a DeepONet or hypernetwork. If this is right, a single kernel-based function-space construction contains these architectures and requires none of the reflexivity, separability, symmetric-domain, or finite-output assumptions that earlier vector-valued constructions needed.

Core claim

On its own terms, the paper's central discovery is that the reproducing-kernel idea survives the passage from Hilbert to Banach outputs without reflexivity, separability, symmetric kernel domains, or finite-dimensional outputs, provided one replaces inner products by a dual pair and operator-valued kernels by twin operators. Every vv-RKBS is shown to belong to an adjoint pair with a unique reproducing kernel. For the integral subclass built from vector-valued measures, the paper establishes a general representer theorem: for data (x_n,y_n) with y_n∈R^d, convex-coercive loss, and λ||·||_B regularization, the minimizer is f†=Σ_{m=1}^{Nd} a_m φ(·,w_m)u_m with u_m extreme points of the unit ball

What carries the argument

The load-bearing object is the twin operator: a bounded bilinear map T:U^⋄×U→R that acts as a linear operator on both U and U^⋄. Twin(U,U^⋄) is the target space of the kernel and generalizes L(U) from vv-RKHS theory. The second object is the integral vv-RKBS: functions of the form (A_{Ω→X}μ)(x)=∫_Ω φ(x,w)dμ(w) with μ a U-valued measure; its reproducing kernel is K(x,w)=φ(x,w)⟨·|·⟩_U. The proof mechanism is the sparse-atomic argument: after identifying M(Ω;U) with the dual of C0(Ω;U*), the extreme points of its unit ball are taken to be single atoms δ_w u with u extreme in U, which turns any convex weak-* continuous objective into a finite atomic minimizer.

Load-bearing premise

The sparse representer theorem rests on the cited identity that the unit ball of M(Ω;U) has extreme points exactly of the form δ_wu with u an extreme point of the output unit ball; if that identity fails for a Banach output space, the solution need not collapse into the Nd-term neural/operator form.

What would settle it

Look at the proof of the extreme-point identity cited for Eq. (4.60) and check it verbatim for U-valued measures with U = L^∞([0,1]) (a non-reflexive dual space with predual L^1). If the unit ball of M(Ω;U) has an extreme point not of the form δ_wu with u extreme in U, then Theorem 4.4's atomic representation is not valid in the general setting claimed.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Every vv-RKBS — a Banach space of U-valued functions with bounded point evaluations — can be paired with a dual space and assigned a unique reproducing kernel K:X×Ω→Twin(U,U^⋄), even when U is non-reflexive and non-separable.
  • For integral vv-RKBSs, minimizers of convex, coercive, lower-semicontinuous supervised losses with total-variation regularization are atomic: f†=Σ_{m=1}^{Nd} a_m φ(·,w_m)u_m with u_m extreme points of the output unit ball.
  • For neural feature maps, the representer theorem recovers exactly a shallow R^d-valued network Uσ(Wx+B), so training in the function space is equivalent to training that architecture.
  • DeepONets and hypernetworks have a shared integral vv-RKBS function space; the weight-space and function-space formulations have the same sparse minimizer f†(z)(x)=Σ φ(z,w_m)ψ(x,θ_m)v_m.
  • Vector-valued adjoint RKBS pairs reduce to scalar adjoint RKBS pairs by augmenting the domains with output directions, extending the scalar RKBS toolbox to vector outputs.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The Nd bound in the representer theorem is an artifact of the non-group-sparse regularization; with ℓ^2-type output norms and a structured-sparsity penalty one could likely force active neurons to collapse to at most N, making the architecture bound match the RKHS case.
  • The weight-space/function-space equivalence suggests an operator-learning analogue of conditional kernel embeddings: a DeepONet is a conditional embedding of the input z into the output RKBS B, which could justify kernel-based uncertainty quantification for neural operators.
  • Because any vv-RKBS pair scalarizes without extra assumptions, numerical solvers for scalar RKBS could be reused for vector-valued problems by treating each output direction u^⋄ as an extra input coordinate.
  • The kernel chaining idea for deep scalar networks could be ported to twin-operator kernels, giving a rigorous function space for deep vector-valued networks; the paper leaves this as future work.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper develops a general framework of vector-valued reproducing kernel Banach spaces (vv-RKBS) based on adjoint pairs (B,B^⋄) and twin operators between a dual pair (U,U^⋄). It shows that every vv-RKBS admits such an adjoint pair with an associated reproducing kernel (Theorem 3.4), and that the construction reduces to scalar RKBS pairs via a scalarization theorem (Theorem 3.6). The authors then specialize to integral and neural vv-RKBSs, where functions are obtained by integrating a scalar kernel against vector-valued measures. They prove that the integral vv-RKBS is indeed an adjoint pair (Theorem 4.2), establish density and duality properties (Theorem 4.3), and state a general representer theorem (Theorem 4.4). This representer theorem is applied to R^d-valued neural networks (Corollary 5.1) and to hypernetwork/DeepONet architectures, with a joint representer theorem for weight-space and function-space formulations (Theorems 5.1 and 5.2). The advertised contributions are the kernel-based formulation without reflexivity/separability assumptions and the recovery of neural and operator architectures as solutions of convex variational problems.

Significance. If the central results hold, this is a valuable unifying contribution. The adjoint-pair/twin-operator formalism is a natural extension of vv-RKHS theory, and the explicit construction of kernels for integral vv-RKBSs goes beyond existing work by avoiding symmetry of the kernel domain and structural assumptions such as reflexivity and separability. The scalarization theorem (Theorem 3.6) and the detailed measure-theoretic verification of the pairing (Theorem B.1, Theorem 4.1, Theorem 4.2) are careful and useful. The representer theorems for finite-dimensional and operator-valued outputs are the main advertised payoff: they connect convex optimization over a Banach space directly to shallow networks, DeepONets, and hypernetworks. However, the most important new claims—Theorem 4.4 and its consequences—rest on an extreme-point characterization for vector measures that is asserted rather than proved, and that is known to fail for some Banach output spaces. The lambda=0 case is also not covered by the cited sparse-representer machinery. These are load-bearing gaps, so the paper in its present form overstates its generality, though the overall framework appears defensible with additional

major comments (3)
  1. [Section 4.4, Eq. (4.60)] The sparse representation in Theorem 4.4 depends crucially on the assertion that Ext({mu in M(Omega;U): |mu|_U(Omega)<=1}) = {delta_w u : w in Omega, u in Ext(B_U)}. The text cites [52] and Lemma 3.2 of [13] and states that the proof 'applies without modifications in the general setting considered here', but no proof is supplied. This is not a general fact for arbitrary Banach output spaces U with a predual: variants of this extreme-point theorem require a Radon-Nikodym-type condition on the relevant space, and it fails e.g. when the value space is L^infty[0,1]. Since Theorem 4.4, Corollary 5.1, and Theorem 5.2 all inherit the double-atom form from Eq. (4.60), this is load-bearing. The authors must either prove Eq. (4.60) under the hypotheses of Definition 4.3 and Theorem 4.4, or add explicit RNP-type assumptions on U (and on V in Theorem 5.2) and state the resulting limitations in the a
  2. [Theorem 4.4, Eq. (4.49)] The theorem states the result for lambda >= 0, and the proof invokes Theorem 3.3 of Bredies and Carioni [12]. That theorem is a sparse-representer result for lambda > 0; for lambda = 0 the regularizer vanishes and there is no mechanism forcing the minimizer to be a finite sum of atoms. The argument as written does not cover lambda = 0. Consequently Corollary 5.1 and Theorem 5.2, which also state lambda >= 0, inherit this gap. Either restrict the statement to lambda > 0 or provide a separate proof for the unregularized case, noting that sparsity may fail entirely in that regime.
  3. [Theorem 5.2 / Definition 5.2] In the weight-space formulation, Theorem 4.4 is applied with output space U = M(Theta;V). This is legitimate only if Eq. (4.60) holds for U = M(Theta;V), which again requires an RNP-type condition on V (or on M(Theta;V)). The theorem statement only assumes that V has a predual V^diamond; no RNP-type hypothesis is stated. For scalar-valued base networks (V=R) this is harmless, but the claimed generality for arbitrary Banach output spaces, and in particular for the function-space viewpoint where the output space is the integral vv-RKBS B, is not supported. Please state the additional hypotheses needed and adjust the claims in Remark 5.3 and Section 5.2.3 accordingly.
minor comments (5)
  1. [Global notation] Theorem 4.4 says 'with predual U^*' and then uses the canonical pairing between U^* and (U^*)^* = U. This conflicts with the standard notation U^* for the continuous dual of U. Use e.g. U_* for the predual to avoid confusion.
  2. [Eq. (4.60)] The reference in the text is given as 'Theorem 2 of Dirk [52]' but the bibliography entry is 'Dirk Werner'. Correct the in-text citation.
  3. [Section 4.3, Eq. (4.44)] In the line 'since every g in B^diamond subset C0(Omega;U)', the codomain should be U^diamond, not U, because B^diamond consists of functions with values in U^diamond. This appears to be a typo.
  4. [Eq. (5.17)] There is an extra closing parenthesis in 'M((A_{Omega->Z} mu)(z_n)))))'. Minor typographical issue.
  5. [Theorem 4.4 proof, existence] The proof begins 'Assume that there exists mu^dagger' but does not show existence. For lambda > 0, existence follows from the weak-* compactness and lower semicontinuity arguments sketched in the proof, but it would be cleaner to state and prove existence as part of the theorem, especially since the unregularized case is already excluded per the major comment above.

Circularity Check

0 steps flagged

No circularity: the representer theorems follow from external sparse-optimization and extreme-point results, not from fitted parameters or from conclusions assumed as premises.

full rationale

The central derivation chain is not circular. Theorem 4.4 poses a convex, coercive weak-* lower semicontinuous optimization problem over vector measures and invokes Bredies-Carioni [12] plus an extreme-point characterization, Eq. (4.60), cited to Werner [52] and Bredies et al. [13]. The sparse representation f^\dagger = sum a_m phi(.,w_m)u_m is obtained by substituting the resulting sparse measure into the defining integral operator A_{\Omega->X}; no parameter is fitted and no neural architecture is inserted as an assumption. The neural, DeepONet, and hypernetwork forms in Corollary 5.1 and Theorem 5.2 follow by algebraically expanding the sparse measure with the chosen phi and psi; the embedding in Remark 5.2 only shows membership of DeepONets in the constructed space, not that the minimizer is constrained to that form. Theorem 3.4, which associates every vv-RKBS with an adjoint pair, is an explicit existence construction whose adjoint space and kernel are built from the point-evaluation functionals of B; proving existence by construction is not circular. The self-citations to Heeringa et al. [24] and Bredies et al. [13] support lemmas or proof strategies, and the key extreme-point statement is additionally anchored in the external reference Werner [52]. The assertion that the proof of Lemma 3.2 of [13] 'applies without modifications' is a rigor/scope concern rather than a circularity, because the theorem's premises do not already contain its conclusion. There is no fitted-input-called-prediction step, no ansatz smuggled in via citation, and no renaming of a known empirical pattern as new structure.

Axiom & Free-Parameter Ledger

0 free parameters · 6 axioms · 1 invented entities

The central mathematical results rest on standard functional analysis (vector-measure duality, Hahn-Kolmogorov extension, Bredies-Carioni). The two domain assumptions that narrow the scope are the extreme-point characterization of vector-measure balls and the predual requirement on the output space; both are stated or cited, but the extreme-point generalization is not proved in the paper.

axioms (6)
  • standard math Vector-measure duality C0(Omega;U)* isomorphic to M(Omega;U*)
    Used to identify M(Omega;U) as a dual space in Theorem 4.4 and Theorem 5.2; cited to Meziani [37] and Singer [45].
  • domain assumption Extreme points of the unit ball of M(Omega;U) are exactly {delta_w u : w in Omega, u in Ext(B_U)}
    Load-bearing for the sparse form in Theorems 4.4 and 5.2; cited to Werner [52] and Bredies et al. [13] with the comment that the proof applies without modifications, but not proven in the paper.
  • standard math Bredies-Carioni sparse variational theorem
    Gives existence and sparsity of minimizers over the measure space; used in the proof of Theorem 4.4.
  • standard math Hahn-Kolmogorov extension for the signed vector pairing <rho|mu>_U
    Proved in Theorem B.1; needed for well-defined pairing in Definition 4.3.
  • domain assumption Input spaces for neural operators are compact or effectively finite-dimensional projections
    Remark 4.4 invokes the finite-dimensional manifold hypothesis to keep phi in C0 for operator learning; if not satisfied, the integral vv-RKBS model does not apply.
  • domain assumption Output Banach space U has a predual U*
    Theorem 4.4 and Theorem 5.2 require U=(U*)*; this excludes spaces like C(D) that are not dual spaces.
invented entities (1)
  • Twin operator space Twin(U,U^diamond) no independent evidence
    purpose: Generalizes L(U) to dual pairs, allowing a reproducing-kernel definition without reflexivity or separability
    A new mathematical construct introduced in Definition 3.5; no empirical prediction, but internally coherent and used throughout the kernel theory.

reviewed 2026-08-04 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Vector-Valued Reproducing Kernel Banach Spaces for Neural Networks and Operators." pith.science (2026). https://pith.science/paper/QXZSW4ZS

@misc{pith2026250926371,
  author       = {Pith},
  title        = {Pith review of: Vector-Valued Reproducing Kernel Banach Spaces for Neural Networks and Operators},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QXZSW4ZS}},
  note         = {Machine review of arXiv:2509.26371}
}
Share X Bluesky LinkedIn Reddit HN
abstract

Recently, there has been growing interest in characterizing the function spaces underlying neural networks. While shallow and deep scalar-valued neural networks have been linked to scalar-valued reproducing kernel Banach spaces (RKBS), $\mathbb{R}^d$-valued neural networks and neural operator models remain less understood in the RKBS setting. To address this gap, we develop a notion of adjoint pairs of vector-valued RKBSs (vv-RKBS), which inherently involves an associated reproducing kernel, and prove that every vv-RKBS belongs to such a pair. Our construction extends existing kernel definitions by avoiding restrictive assumptions such as symmetric kernel domains, finite-dimensional output spaces, reflexivity, or separability, while still recovering familiar properties of vector-valued reproducing kernel Hilbert spaces (vv-RKHS). We then show that shallow $\mathbb{R}^d$-valued neural networks are elements of a specific vv-RKBS, namely an instance of an integral vv-RKBS. To also explore the functional structure of neural operators, we analyze the DeepONet and Hypernetwork architectures and demonstrate that they too belong to an integral vv-RKBS. In all cases, we establish a representer theorem, showing that optimization over these function spaces recovers the corresponding neural architectures.

Figures

Figures reproduced from arXiv: 2509.26371 by Jos\'e A. Iglesias, Sven Dummer, Tjeerd Jan Heeringa.

Figure 1
Figure 1. Figure 1: From vv-RKHS to adjoint pair of vv-RKBS. In a vv-RKHS, functions f : X → U take values in a Hilbert space U and admit a reproducing kernel K : X × X → L(U). In the vv-RKBS setting, we instead use Banach spaces B and B ⋄ for functions mapping to a dual pair (U, U ⋄ ), and replace inner products with duality pairings. This breaks the symmetry in the domain, replacing X × X with X × Ω, and requires twin opera… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

57 extracted references · 6 linked inside Pith

  1. [1]

    Neural operator: Graph kernel network for partial differential equations

    Anima Anandkumar, Kamyar Azizzadenesheli, Kaushik Bhattacharya, Nikola Kovachki, Zongyi Li, Burigede Liu, and Andrew Stuart. Neural operator: Graph kernel network for partial differential equations. InICLR 2020 Workshop on Integration of Deep Neural Models and Differential Equations, 2019

  2. [2]

    Theory of reproducing kernels.Transactions of the American mathematical society, 68(3):337–404, 1950

    Nachman Aronszajn. Theory of reproducing kernels.Transactions of the American mathematical society, 68(3):337–404, 1950

  3. [3]

    Structured sparsity-inducing norms through submodular functions.Advances in Neural Information Processing Systems, 23, 2010

    Francis Bach. Structured sparsity-inducing norms through submodular functions.Advances in Neural Information Processing Systems, 23, 2010

  4. [4]

    Breaking the curse of dimensionality with convex neural networks.Journal of Machine Learning Research, 18(19):1–53, 2017

    Francis Bach. Breaking the curse of dimensionality with convex neural networks.Journal of Machine Learning Research, 18(19):1–53, 2017

  5. [5]

    Iglesias, Yury Korolev, Emanuele Naldi, and Stefano Vi- gogna

    Francesca Bartolucci, Marcello Carioni, Jos´ e A. Iglesias, Yury Korolev, Emanuele Naldi, and Stefano Vi- gogna. A Lipschitz spaces view of infinitely wide shallow neural networks.arXiv preprint arXiv:2410.14591, 2024

  6. [6]

    Understanding neural networks with reproducing kernel Banach spaces.Applied and Computational Harmonic Analysis, 62:194– 236, 2023

    Francesca Bartolucci, Ernesto De Vito, Lorenzo Rosasco, and Stefano Vigogna. Understanding neural networks with reproducing kernel Banach spaces.Applied and Computational Harmonic Analysis, 62:194– 236, 2023

  7. [7]

    Neural reproducing kernel Banach spaces and representer theorems for deep networks.arXiv preprint arXiv:2403.08750, 2024

    Francesca Bartolucci, Ernesto De Vito, Lorenzo Rosasco, and Stefano Vigogna. Neural reproducing kernel Banach spaces and representer theorems for deep networks.arXiv preprint arXiv:2403.08750, 2024

  8. [8]

    Kernel methods are competitive for operator learning.Journal of Computational Physics, 496:112549, 2024

    Pau Batlle, Matthieu Darcy, Bamdad Hosseini, and Houman Owhadi. Kernel methods are competitive for operator learning.Journal of Computational Physics, 496:112549, 2024

  9. [9]

    Kovachki, and Andrew M

    Kaushik Bhattacharya, Bamdad Hosseini, Nikola B. Kovachki, and Andrew M. Stuart. Model reduction and neural networks for parametric PDEs.The SMAI Journal of computational mathematics, 7:121–157, 2021

  10. [10]

    Springer, 2007

    Vladimir Bogachev.Measure Theory, volume 2. Springer, 2007

  11. [11]

    A mathematical guide to operator learning

    Nicolas Boull´ e and Alex Townsend. A mathematical guide to operator learning. InHandbook of Numerical Analysis, volume 25, pages 83–125. Elsevier, 2024

  12. [12]

    Sparsity of solutions for variational inverse problems with finite- dimensional data.Calculus of Variations and Partial Differential Equations, 59(1):14, 2020

    Kristian Bredies and Marcello Carioni. Sparsity of solutions for variational inverse problems with finite- dimensional data.Calculus of Variations and Partial Differential Equations, 59(1):14, 2020

  13. [13]

    Iglesias, and Daniel Walter

    Kristian Bredies, Jos´ e A. Iglesias, and Daniel Walter. On extremal points for some vectorial total variation seminorms.arXiv preprint arXiv:2404.12831, 2024

  14. [14]

    Functional learning through kernels.arXiv preprint arXiv:0910.1013, 2009

    St´ ephane Canu, Xavier Mary, and Alain Rakotomamonjy. Functional learning through kernels.arXiv preprint arXiv:0910.1013, 2009

  15. [15]

    Vector valued reproducing kernel Hilbert spaces of integrable functions and Mercer theorem.Analysis and Applications, 4(04):377–408, 2006

    Claudio Carmeli, Ernesto De Vito, and Alessandro Toigo. Vector valued reproducing kernel Hilbert spaces of integrable functions and Mercer theorem.Analysis and Applications, 4(04):377–408, 2006

  16. [16]

    Vector valued reproducing kernel Hilbert spaces and universality.Analysis and Applications, 8(01):19–61, 2010

    Claudio Carmeli, Ernesto De Vito, Alessandro Toigo, and Veronica Umanit´ a. Vector valued reproducing kernel Hilbert spaces and universality.Analysis and Applications, 8(01):19–61, 2010

  17. [17]

    Vector-valued reproducing kernel Banach spaces with group lasso norms.arXiv preprint arXiv:1903.00819, 2019

    Liangzhi Chen, Haizhang Zhang, and Jun Zhang. Vector-valued reproducing kernel Banach spaces with group lasso norms.arXiv preprint arXiv:1903.00819, 2019

  18. [18]

    Combettes, Saverio Salzo, and Silvia Villa

    Patrick L. Combettes, Saverio Salzo, and Silvia Villa. Regularized learning schemes in feature Banach spaces.Analysis and Applications, 16(01):1–54, 2018

  19. [19]

    Verduyn Lunel

    Odo Diekmann and Sjoerd M. Verduyn Lunel. Twin semigroups and delay equations.J. Differential Equations, 286:332–410, 2021

  20. [20]

    RDA-INR: Riemannian diffeomorphic autoen- coding via implicit neural representations.SIAM Journal on Imaging Sciences, 17(4):2302–2330, 2024

    Sven Dummer, Nicola Strisciuglio, and Christoph Brune. RDA-INR: Riemannian diffeomorphic autoen- coding via implicit neural representations.SIAM Journal on Imaging Sciences, 17(4):2302–2330, 2024

  21. [21]

    RONOM: Reduced-order neural operator modeling

    Sven Dummer, Dongwei Ye, and Christoph Brune. RONOM: Reduced-order neural operator modeling. arXiv preprint arXiv:2507.12814, 2025. 39

  22. [22]

    Springer Science & Business Media, 2011

    Mari´ an Fabian, Petr Habala, Petr H´ ajek, Vicente Montesinos, and V´ aclav Zizler.Banach space theory: The basis for linear and nonlinear analysis. Springer Science & Business Media, 2011

  23. [23]

    Georgiev, Luis S´ anchez-Gonz´ alez, and Panos M

    Pando G. Georgiev, Luis S´ anchez-Gonz´ alez, and Panos M. Pardalos. Construction of pairs of reproducing kernel Banach spaces. InConstructive Nonsmooth Analysis and Related Topics, pages 39–57. Springer, 2013

  24. [24]

    Deep networks are reproducing kernel chains.arXiv preprint arXiv:2501.03697, 2025

    Tjeerd Jan Heeringa, Len Spek, and Christoph Brune. Deep networks are reproducing kernel chains.arXiv preprint arXiv:2501.03697, 2025

  25. [25]

    Koopman and Perron–Frobenius operators on reproducing kernel Banach spaces.Chaos: An Interdisciplinary Journal of Nonlinear Science, 32(12), 2022

    Masahiro Ikeda, Isao Ishikawa, and Corbinian Schlosser. Koopman and Perron–Frobenius operators on reproducing kernel Banach spaces.Chaos: An Interdisciplinary Journal of Nonlinear Science, 32(12), 2022

  26. [26]

    Learning in latent spaces improves the predictive accuracy of deep neural operators.arXiv preprint arXiv:2304.07599, 2023

    Katiana Kontolati, Somdatta Goswami, George Em Karniadakis, and Michael D Shields. Learning in latent spaces improves the predictive accuracy of deep neural operators.arXiv preprint arXiv:2304.07599, 2023

  27. [27]

    Two-layer neural networks with values in a Banach space.SIAM Journal on Mathematical Analysis, 54(6):6358–6389, 2022

    Yury Korolev. Two-layer neural networks with values in a Banach space.SIAM Journal on Mathematical Analysis, 54(6):6358–6389, 2022

  28. [28]

    On universal approximation and error bounds for fourier neural operators.Journal of Machine Learning Research, 22(290):1–76, 2021

    Nikola Kovachki, Samuel Lanthaler, and Siddhartha Mishra. On universal approximation and error bounds for fourier neural operators.Journal of Machine Learning Research, 22(290):1–76, 2021

  29. [29]

    Fourier neural operator for parametric partial differential equations

    Zongyi Li, Nikola Kovachki, Kamyar Azizzadenesheli, Burigede Liu, Kaushik Bhattacharya, Andrew Stu- art, and Anima Anandkumar. Fourier neural operator for parametric partial differential equations. In International Conference on Learning Representations, 2021

  30. [30]

    On reproducing kernel Banach spaces: Generic definitions and unified framework of constructions.Acta Mathematica Sinica, English Series, 38(8):1459– 1483, 2022

    Rong Rong Lin, Hai Zhang Zhang, and Jun Zhang. On reproducing kernel Banach spaces: Generic definitions and unified framework of constructions.Acta Mathematica Sinica, English Series, 38(8):1459– 1483, 2022

  31. [31]

    Multi-task learning in vector-valued reproducing kernel Banach spaces with theℓ 1 norm.Journal of Complexity, 63:101514, 2021

    Rongrong Lin, Guohui Song, and Haizhang Zhang. Multi-task learning in vector-valued reproducing kernel Banach spaces with theℓ 1 norm.Journal of Complexity, 63:101514, 2021

  32. [32]

    Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators.Nature machine intelligence, 3(3):218–229, 2021

    Lu Lu, Pengzhan Jin, Guofei Pang, Zhongqiang Zhang, and George Em Karniadakis. Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators.Nature machine intelligence, 3(3):218–229, 2021

  33. [33]

    A comprehensive and fair comparison of two neural operators (with practical extensions) based on fair data.Computer Methods in Applied Mechanics and Engineering, 393:114778, 2022

    Lu Lu, Xuhui Meng, Shengze Cai, Zhiping Mao, Somdatta Goswami, Zhongqiang Zhang, and George Em Karniadakis. A comprehensive and fair comparison of two neural operators (with practical extensions) based on fair data.Computer Methods in Applied Mechanics and Engineering, 393:114778, 2022

  34. [34]

    Towards a mathematical understanding of neural network- based machine learning: what we know and what we don’t.arXiv preprint arXiv:2009.10713, 2020

    Chao Ma, Stephan Wojtowytsch, Lei Wu, et al. Towards a mathematical understanding of neural network- based machine learning: what we know and what we don’t.arXiv preprint arXiv:2009.10713, 2020

  35. [35]

    A priori estimates of the population risk for two-layer neural networks.Commu- nications in Mathematical Sciences, 17(5):1407–1425, 2019

    Chao Ma, Lei Wu, et al. A priori estimates of the population risk for two-layer neural networks.Commu- nications in Mathematical Sciences, 17(5):1407–1425, 2019

  36. [36]

    The Barron space and the flow-induced function spaces for neural network models

    Chao Ma, Lei Wu, et al. The Barron space and the flow-induced function spaces for neural network models. Constructive Approximation, 55(1):369–406, 2022

  37. [37]

    On the dual spaceC ∗ 0 (S, X).Acta Math

    Lakhdar Meziani. On the dual spaceC ∗ 0 (S, X).Acta Math. Univ. Comenian. (N.S.), 78(1):153–160, 2009

  38. [38]

    Rahul Parhi and Robert D. Nowak. Banach space representer theorems for neural networks and ridge splines.Journal of Machine Learning Research, 22(43):1–40, 2021

  39. [39]

    Rahul Parhi and Robert D. Nowak. What kinds of functions do deep neural networks learn? Insights from variational spline theory.SIAM Journal on Mathematics of Data Science, 4(2):464–489, 2022

  40. [40]

    A measure-theoretic approach to kernel conditional mean embed- dings.Advances in Neural Information Processing Systems, 33:21247–21259, 2020

    Junhyung Park and Krikamol Muandet. A measure-theoretic approach to kernel conditional mean embed- dings.Advances in Neural Information Processing Systems, 33:21247–21259, 2020

  41. [41]

    Paulsen and Mrinal Raghupathi.An introduction to the theory of reproducing kernel Hilbert spaces, volume 152 ofCambridge Studies in Advanced Mathematics

    Vern I. Paulsen and Mrinal Raghupathi.An introduction to the theory of reproducing kernel Hilbert spaces, volume 152 ofCambridge Studies in Advanced Mathematics. Cambridge University Press, 2016

  42. [42]

    Con- volutional neural operators

    Bogdan Raonic, Roberto Molinaro, Tobias Rohner, Siddhartha Mishra, and Emmanuel de Bezenac. Con- volutional neural operators. InICLR 2023 Workshop on Physics for Machine Learning, 2023

  43. [43]

    Jacob Seidman, Georgios Kissas, Paris Perdikaris, and George J. Pappas. NOMAD: Nonlinear manifold decoders for operator learning.Advances in Neural Information Processing Systems, 35:5601–5613, 2022. 40

  44. [44]

    Joseph Shenouda, Rahul Parhi, Kangwook Lee, and Robert D. Nowak. Variation spaces for multi-output neural networks: Insights on multi-task learning and network compression.Journal of Machine Learning Research, 25(231):1–40, 2024

  45. [45]

    Linear functionals on the space of continuous mappings of a compact Hausdorff space into a Banach space.Rev

    Ivan Singer. Linear functionals on the space of continuous mappings of a compact Hausdorff space into a Banach space.Rev. Math. Pures Appl., 2:301–315, 1957

  46. [46]

    Implicit neural representations with periodic activation functions.Advances in Neural Information Processing Systems, 33:7462–7473, 2020

    Vincent Sitzmann, Julien Martel, Alexander Bergman, David Lindell, and Gordon Wetzstein. Implicit neural representations with periodic activation functions.Advances in Neural Information Processing Systems, 33:7462–7473, 2020

  47. [47]

    Duality for neural networks through reproducing kernel Banach spaces.Applied and Computational Harmonic Analysis, 78:101765, 2025

    Len Spek, Tjeerd Jan Heeringa, Felix Schwenninger, and Christoph Brune. Duality for neural networks through reproducing kernel Banach spaces.Applied and Computational Harmonic Analysis, 78:101765, 2025

  48. [48]

    Reproducing kernel Hilbert spaces cannot contain all continuous functions on a compact metric space.Archiv der Mathematik, 122(5):553–557, 2024

    Ingo Steinwart. Reproducing kernel Hilbert spaces cannot contain all continuous functions on a compact metric space.Archiv der Mathematik, 122(5):553–557, 2024

  49. [49]

    American Mathematical Society, 2011

    Terence Tao.An introduction to measure theory, volume 126 ofGraduate Studies in Mathematics. American Mathematical Society, 2011

  50. [50]

    Sparse representer theorems for learning in reproducing kernel Banach spaces.Journal of Machine Learning Research, 25(93):1–45, 2024

    Rui Wang, Yuesheng Xu, and Mingsong Yan. Sparse representer theorems for learning in reproducing kernel Banach spaces.Journal of Machine Learning Research, 25(93):1–45, 2024

  51. [51]

    Hypothesis spaces for deep learning.Neural Networks, 193:107995, 2026

    Rui Wang, Yuesheng Xu, and Mingsong Yan. Hypothesis spaces for deep learning.Neural Networks, 193:107995, 2026

  52. [52]

    Extreme points in spaces of operators and vector-valued measures.Rend

    Dirk Werner. Extreme points in spaces of operators and vector-valued measures.Rend. Circ. Mat. Palermo (2), Suppl.(5):135–143, 1984

  53. [53]

    Stephan Wojtowytsch et al. On the Banach spaces associated with multi-layer ReLU networks: Function representation, approximation theory and gradient descent dynamics.CSIAM Transactions on Applied Mathematics, 1(3):387–440, 2020

  54. [54]

    Representation formulas and pointwise properties for Barron functions.Cal- culus of Variations and Partial Differential Equations, 61(2):1–37, 2022

    Stephan Wojtowytsch et al. Representation formulas and pointwise properties for Barron functions.Cal- culus of Variations and Partial Differential Equations, 61(2):1–37, 2022

  55. [55]

    Sparse machine learning in Banach spaces.Applied Numerical Mathematics, 187:138–157, 2023

    Yuesheng Xu. Sparse machine learning in Banach spaces.Applied Numerical Mathematics, 187:138–157, 2023

  56. [56]

    Reproducing kernel Banach spaces for machine learning

    Haizhang Zhang, Yuesheng Xu, and Jun Zhang. Reproducing kernel Banach spaces for machine learning. Journal of Machine Learning Research, 10(12), 2009

  57. [57]

    Vector-valued reproducing kernel Banach spaces with applications to multi-task learning.Journal of Complexity, 29(2):195–215, 2013

    Haizhang Zhang and Jun Zhang. Vector-valued reproducing kernel Banach spaces with applications to multi-task learning.Journal of Complexity, 29(2):195–215, 2013. 41

This paper was first reviewed by deepseek-v4-flash on August 4, 2026.