Pith. sign in

REVIEW 7 minor 34 references

Schoenberg characterization of continuous non-stationary isotropic positive definite kernels

T0 review · 0 major / 7 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Every continuous rotation-invariant positive definite kernel on $\mathbb{R}^d$ is a unique Gegenbauer series.

desk verdict Solid Schoenberg-type characterization for non-stationary isotropic kernels on R^d (d≥2); minor presentational flaws, core proof sound. read the letter →

arxiv 2506.22048 v2 pith:ENEPDE7F submitted 2025-06-27 math.ST stat.TH

classification math.STstat.TH MSC 33C5033C5542A8242C1043A3560G1568T07
keywords positivedefinitekernelsisotropySchoenbergcharacterizationGegenbauerpolynomialsnon-stationaryneuralnetworkstrictdefinitenessGaussianprocesses
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proves a Schoenberg-type characterization for continuous positive definite kernels on $\mathbb{R}^d$ that are invariant under the orthogonal group $O(d)$ but not necessarily stationary: such a kernel is exactly a sum over $n$ of coefficient kernels $\alpha_n(\|x\|,\|y\|)$ times normalized Gegenbauer polynomials in the cosine of the angle between $x$ and $y$. The $\alpha_n$ are unique, continuous, positive definite kernels on $[0,\infty)^2$ that vanish when one argument is zero, and the sum converges on the diagonal. This class unifies stationary isotropic kernels (functions of distance) and dot product kernels (functions of inner product), and it includes the infinite-width neural network kernels that arise in machine learning theory. The paper also gives necessary and sufficient conditions for such kernels to be strictly positive definite, phrased in terms of parity of the nonzero coefficients.

What carries the argument

The central object is the series representation with normalized Gegenbauer polynomials $\widehat{C}_n^{(\lambda)}$ evaluated at the inner product of the normalized vectors. The argument carries by the homeomorphism $T(r,v)=rv$ between $(0,\infty)\times S^{d-1}$ and $\mathbb{R}^d\setminus\{0\}$, which turns an isotropic kernel on $\mathbb{R}^d$ into a kernel on the product space of the form $f_S(r,s,\rho)=\kappa(r,s,rs\rho)$. On that product space the cited characterization supplies the series with positive definite coefficient kernels, and continuity plus the condition $\alpha_n(r,0)=\alpha_n(0,r)=0$ for $n\ge 1$ extends the representation to the origin. For $d=\infty$ the Gegenbauer polynomials are replaced by monomials $\rho^n$, giving a power series in the cosine.

What would settle it

Compute the Gegenbauer coefficients in (5) for a concrete continuous $O(d)$-invariant kernel such as $K(x,y)=\exp(-\|x-y\|^2)$. If any coefficient $\alpha_n$ fails to be positive definite on $[0,\infty)^2$, or fails to vanish at the origin, the representation (3) is false. Conversely, a continuous $O(d)$-invariant positive definite kernel whose Gegenbauer projections are not summable on the diagonal would also disprove the theorem.

Watch

Extended reading notes

Core claim

Theorem 2.1 states that for $d\in\{2,3,\ldots,\infty\}$, a continuous $K:\mathbb{R}^d\times\mathbb{R}^d\to\mathbb{C}$ is an isotropic positive definite kernel—meaning $K(Ux,Uy)=K(x,y)$ for every orthogonal $U$—if and only if it admits the representation $$K(x,y)=\sum_{n=0}^{\infty}\$alpha_n^{{(d)}}$(\|x\|,\|y\|)\,\widehat{C}$_n^{{(\lambda)}}$\!\left(\left\langle \frac{x}{\|x\|},\frac{y}{\|y\|}\right\rangle\right)$$ with $\lambda=(d-2)/2$, where the $\alpha_n^{(d)}$ are unique continuous positive definite kernels on $[0,\infty)^2$, summable on the diagonal, and zero at the origin for $n\ge 1$. For finite $d$ the coefficients are obtained by Gegenbauer projection of $\kappa(r,s,rs\rho)$ with weight $(1-\rho^2)^{\lambda-1/2}$. The same framework yields Theorem 3.1 and Theorem 3.2: strict positive definiteness holds essentially when the even and odd parts of the coefficient sequence are themselves strictly positive definite on $(0,\infty)$, with the $d=2$ case requiring intersection with every full arithmetic progression in $\mathbb{Z}$.

Load-bearing premise

The whole theorem rests on the external characterization of positive definite kernels on $X\times S^{d-1}$ (Theorem A.1); if that result were false, misstated, or inapplicable, the series representation and every strict-positive-definiteness condition derived from it would collapse.

Editorial extensions

If this is right

  • Every continuous $O(d)$-invariant positive definite kernel, including non-stationary ones, has a unique explicit series expansion, so kernel design can be done by choosing the coefficient kernels $\alpha_n$ instead of the full function $K$.
  • Stationary isotropic kernels and dot product kernels are recovered as special cases, giving a single framework for both classical classes.
  • The infinite-width neural network Gaussian process (NNGP) kernel and the neural tangent kernel are isotropic and therefore fall under the characterization, with the NNGP kernel expressed as $\sum_m \alpha_m(\|x\|,\|x'\|)\langle x/\|x\|, x'/\|x'\|\rangle^m$ in $\ell_2$.
  • Strict positive definiteness can be read off from the parity of the active coefficients: infinitely many even and odd $n$ must contribute, with $\alpha_0(0,0)>0$.
  • In infinite dimension the representation reduces to a power series in the cosine, which is the natural Schoenberg form on the infinite-dimensional sphere.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural next step, hinted at in the paper's remark on universal kernels, is a characterization of universal isotropic kernels on $\mathbb{R}^d$: universality should correspond to a density condition on the coefficient kernels $\alpha_n$ in addition to their positive definiteness.
  • The coefficient formula (5) suggests a practical numerical certificate: for any candidate isotropic kernel, compute the Gegenbauer projections and check positivity and zero-at-origin; this could be used in kernel learning or Gaussian process design.
  • If the load-bearing external product-space characterization (Theorem A.1) were to fail for some edge case, the zero-at-origin extension and therefore the full representation would break down; the $d=2$ arithmetic-progression condition in Theorem 3.2 shows where such edge effects concentrate.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

0 major / 7 minor

Summary. The paper characterizes continuous positive definite kernels on R^d that are invariant under the orthogonal group O(d), i.e. isotropic but not necessarily stationary kernels. The main result, Theorem 2.1, states that such kernels are exactly those of the form K(x,y)=sum_{n=0}^\infty \alpha_n^{(d)}(||x||,||y||) \hat C_n^{(\lambda)}(<x/||x||, y/||y||>), with unique continuous positive definite coefficient kernels \alpha_n on [0,\infty)^2 that are summable on the diagonal and vanish at the origin for n\ge 1; the case d=\infty uses monomials in place of Gegenbauer polynomials. Section 3 gives necessary and sufficient conditions for strict positive definiteness, with the case d=2 treated separately because of its arithmetic-progression structure. Section 4 shows that stationary isotropic kernels and dot product kernels are special cases and that infinite-width neural network kernels (NNGP and, in a remark, NTK) are isotropic non-stationary kernels covered by the characterization. The proofs reduce the main theorem to a restated theorem of Guella and Menegatto for kernels on (0,\infty)\times S^{d-1}, with a careful extension to the origin via Lemma A.3 and the zero-at-origin condition.

Significance. If the result holds, it fills a natural gap in the Schoenberg-type literature: it unifies the classical characterizations of stationary isotropic kernels and dot product kernels, and it applies to an important class of neural network kernels that are isotropic but non-stationary. The boundary treatment at the origin, through the zero-at-origin condition (4) and Lemma A.3, is careful and is exactly what makes the extension from (0,\infty)\times S^{d-1} to R^d work. The paper is transparent about its external dependencies: Theorem A.1 is restated from [10], and one Gram-matrix equivalence is cited from the authors' prior work [2, Appendix F]; neither is circular. The main proof is clear, and the strict positive definiteness criteria are concrete enough to be applied. The paper does not ship code or machine-checked proofs, but the analytic arguments are reproducible from the cited sources.

minor comments (7)
  1. [A.1, proof of Theorem 2.1, d=2 case] The displayed bound |\hat C_n^{(0)}(\rho)(1-\rho^2)^{-1/2}|\le 1 used for dominated convergence is false near \rho=\pm 1. A valid integrable majorant is C(1-\rho^2)^{-1/2}, which follows from the uniform bound |\hat C_n^{(0)}(\rho)|\le 1 and integrability of the weight; with this replacement the continuity-extension argument goes through unchanged.
  2. [Theorem 3.1 and Eq. (13)] The expressions c^T [\alpha_n(r_i,r_j)] c and \sum c_i c_j K(x_i,x_j) should use the conjugate transpose or \bar c, since the quadratic form for a complex positive definite kernel is Hermitian; as written, c^T A c need not be real and the condition is not equivalent to strict positive definiteness for complex-valued kernels.
  3. [Example 4.7] The arccosine kernel formula K_NNGP(x,y)=(1/\pi)||x||||y|| J_1(\cos^{-1}(\langle x,y\rangle)) is missing the normalization of the inner product; the argument should be \cos^{-1}(\langle x,y\rangle/(||x||||y||)) unless the formula is intended only for unit-norm inputs.
  4. [Section A.2, proof of Theorem 3.1] The proof relies on Theorems 3.5, 3.7, and 3.9 of [10] without stating them, and the translation from S^d to S^{d-1} is only mentioned in passing; a short statement of the exact external conditions used would make the claimed equivalences (ii) and (iii) verifiable without consulting the cited paper.
  5. [Theorem 2.1] Since (iii) alone should imply that K is continuous, it would be helpful to state explicitly that the summability of the \alpha_n on the diagonal together with the Cauchy-Schwarz inequality gives uniform convergence of the series on bounded sets; this is implicit but not spelled out.
  6. [Abstract and Theorem 2.1] The characterization excludes d=1, but the abstract says 'on R^d' without this restriction; the domain d\in\{2,3,\ldots,\infty\} should be stated in the abstract to avoid overclaiming.
  7. [Section 2] There is a typo in 'independently discoverd' (should be 'discovered').

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: the central Gegenbauer representation is imported from an external theorem, and the authors' self-citations supply only elementary intermediate equivalences.

full rationale

The main Theorem 2.1 is not circular by construction. Assertion (iii)'s expansion in normalized Gegenbauer polynomials is obtained by applying the external Theorem A.1 (Guella and Menegatto) to f_S(r,s,rho)=kappa(r,s,rs rho) on (0,infinity)xS^{d-1}, then extending the resulting coefficient kernels to the origin. The coefficient kernels alpha_n^{(d)} are uniquely determined by Gegenbauer orthogonality in finite dimension and by power-series coefficient uniqueness in d=infinity, so no fitted parameter is relabelled as a prediction and no assertion is equivalent to its inputs by definition. The authors' prior work is cited twice in the proof chain: [2, Appendix F] for the elementary equivalence (i) iff (ii), namely that an O(d)-invariant continuous kernel depends only on (||x||,||y||,<x,y>), and [2, Prop. F.4] for the standard fact that point configurations with the same Gram matrix are related by an orthogonal transformation. Both are parameter-free supporting facts that do not contain or presuppose the target series representation. The core extension from X x S^{d-1} to R^d, including the zero-at-origin condition (4), is proved within the paper. Section 3 similarly transfers strict-positive-definiteness conditions from the independent Guella-Menegatto framework rather than assuming them. The only caveat is a normal, non-circular self-citation for a preliminary equivalence; hence the low score.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

No free parameters or invented entities. The results rest on external theorems, notably the Guella-Menegatto characterization (Theorem A.1) and a self-cited equivalence from [2]; both are stated transparently.

assumptions (5)
  • domain assumption Guella-Menegatto characterization of positive definite kernels on X times S^{d-1} (Theorem A.1)
    The core external result. The paper restates it and builds Theorems 2.1 and 3.1 on it. If it is false, the paper's main theorems fail.
  • domain assumption Benning-Doring [2, Appendix F] equivalence of O(d)-invariance and dependence on (||x||,||y||,<x,y>)
    The authors cite their own prior published result for the first equivalence in Theorem 2.1 without proof. It is a self-cited, load-bearing result.
  • standard math Standard orthogonality, normalization, and uniform bounds for Gegenbauer polynomials (Reimer [18])
    Used to derive the projection formula (5), uniqueness, and the zero-at-origin condition.
  • standard math Lemma A.2: any two finite configurations with identical Gram matrices are related by an orthogonal transformation
    Used to construct kappa and to prove Lemma 2.3. This is a classical fact.
  • domain assumption Hanin's recursion for NNGP kernels (Hanin [11], Eq. (1.7) and (1.8))
    Load-bearing for Corollary 4.4. The paper does not prove the recursion and relies on its validity for the isotropy and dimension-freeness claim.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Schoenberg characterization of continuous non-stationary isotropic positive definite kernels." pith.science (2026). https://pith.science/paper/ENEPDE7F

@misc{pith2026250622048,
  author       = {Pith},
  title        = {Pith review of: Schoenberg characterization of continuous non-stationary isotropic positive definite kernels},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ENEPDE7F}},
  note         = {Machine review of arXiv:2506.22048}
}
abstract

We characterize the continuous isotropic positive definite kernels on $\mathbb{R}^d$, where isotropy refers to invariance under the orthogonal group $O(d)$ but not necessarily stationarity. Furthermore, we characterize strict positive definiteness for such kernels. The class of isotropic kernels is fairly general as it unifies stationary isotropic and dot product kernels, and includes neural network kernels that arise from infinite-width limits of neural networks. As an application, we further characterize the continuous isotropic Gaussian random functions in terms of a series representation.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

34 extracted references · 29 canonical work pages

  1. [2]

    Benning and L

    F. Benning and L. D¨ oring. Random Function Descent. In Advances in Neural Information Processing Systems , volume 37, pages 111248– 111298, Vancouver, Canada, Dec. 2024. Curran Associates, Inc. URL https://proceedings.neurips.cc/paper_files/paper/2024/hash/ c980ad0fb46f0cfb0faabcd42b30a67a-Abstract-Conference.html

  2. [10]

    J. C. Guella and V. A. Menegatto. Schoenberg’s Theorem for Positive Definite Functions on Products: A Unifying Framework. Journal of Fourier Analysis and Applications , 25(4):1424–1446, Aug. 2019. ISSN 1531-5851. doi: 10.1007/s00041-018-9631-5

  3. [1]

    Ament and C

    S. Ament and C. Gomes. Scalable First-Order Bayesian Optimization via Structured Automatic Differentiation. In Proceedings of the 39th Interna- tional Conference on Machine Learning, Baltimore, Maryland, USA, 2022

  4. [3]

    Berg and E

    C. Berg and E. Porcu. From Schoenberg Coefficients to Schoenberg Func- tions. Constructive Approximation, 45(2):217–241, Apr. 2017. ISSN 1432-

  5. [4]

    S. Bochner. Monotone Funktionen, Stieltjessche Integrale und harmonische Analyse. Mathematische Annalen, 108(1):378–410, Dec. 1933. ISSN 1432-

  6. [5]

    Positive definite matrix must be Hermitian

    chesslad. Positive definite matrix must be Hermitian. Mathematics Stack Exchange, 2019. URL https://math.stackexchange.com/q/3434192

  7. [6]

    Cho and L

    Y. Cho and L. Saul. Kernel Methods for Deep Learning. In Advances in Neural Information Processing Systems , volume 22. Curran Associates, Inc., 2009. URL https://proceedings.neurips.cc/paper/2009/hash/ 5751ec3e9a4feab575962e78e006250d-Abstract.html

  8. [7]

    de Roos, A

    F. de Roos, A. Gessner, and P. Hennig. High-Dimensional Gaussian Process Inference with Derivatives. In Proceedings of the 38th International Con- ference on Machine Learning , pages 2535–2545. PMLR, July 2021. URL https://proceedings.mlr.press/v139/de-roos21a.html

Show all 34 references
  1. [8]

    Estrade, A

    A. Estrade, A. Fari˜ nas, and E. Porcu. Covariance functions on spheres cross time: Beyond spatial isotropy and temporal stationarity. Statistics & Probability Letters, 151:1–7, Aug. 2019. ISSN 0167-7152. doi: 10.1016/j. spl.2019.03.011

  2. [9]

    Galy-Fajou, D

    T. Galy-Fajou, D. Widmann, S. Yalburgi, W. Tebbutt, st–, I. Falk, S. C. Surace, S. Ridderbusch, T. Wright, H. Ge, S. Khan, P. Monti- cone, L. Mones, david-vicente, A. R. Gnadt, J. Giersdorf, J. TagBot, R. Viljoen, S. Sch¨ olly, T. E. Fjelde, and K. ¨Ocal. JuliaGaussianPro- ces...

  3. [11]

    B. Hanin. Random neural networks in the infinite width limit as Gaussian processes. The Annals of Applied Probability, 33(6A):4798–4819, Dec. 2023. ISSN 1050-5164, 2168-8737. doi: 10.1214/23-AAP1933

  4. [12]

    Jacot, F

    A. Jacot, F. Gabriel, and C. Hongler. Neural Tangent Kernel: Convergence and Generalization in Neural Networks. In Advances in Neural Informa- tion Processing Systems, volume 31, Montr´ eal, Canada, 2018. Curran Asso- ciates, Inc. URL https://proceedings.neurips.cc/paper/2018/...

  5. [13]

    J. Lee, Y. Bahri, R. Novak, S. S. Schoenholz, J. Pennington, and J. Sohl- Dickstein. Deep Neural Networks as Gaussian Processes. In International Conference on Learning Representations, Vancouver, Canada, 2018. URL https://openreview.net/forum?id=B1EA-M-0Z

  6. [14]

    C. A. Micchelli, Y. Xu, and H. Zhang. Universal Kernels. Journal of Machine Learning Research , 7(12), 2006. URL http://www.jmlr.org/ papers/volume7/micchelli06a/micchelli06a.pdf

  7. [15]

    Panchenko

    D. Panchenko. The Sherrington-Kirkpatrick Model. Springer Monographs in Mathematics. Springer, New York, NY, 2013. ISBN 978-1-4614-6288-0 978-1-4614-6289-7. doi: 10.1007/978-1-4614-6289-7. 11

  8. [16]

    A. Pinkus. Strictly Positive Definite Functions on a Real Inner Product Space. Advances in Computational Mathematics, 20(4):263–271, May 2004. ISSN 1572-9044. doi: 10.1023/A:1027362918283

  9. [17]

    C. E. Rasmussen and C. K. Williams. Gaussian Processes for Machine Learning. Number 3 in Adaptive Computation and Machine Learning. MIT Press, Cambridge, Massachusetts, 2 edition, 2006. ISBN 0-262-18253- X. URL http://gaussianprocess.org/gpml/chapters/RW.pdf

  10. [18]

    M. Reimer. Gegenbauer Polynomials. In Multivariate Polynomial Approx- imation, pages 19–38. Birkh¨ auser, Basel, 2003. ISBN 978-3-0348-8095-4. doi: 10.1007/978-3-0348-8095-4 2

  11. [19]

    Sasv´ ari

    Z. Sasv´ ari. Multivariate Characteristic and Correlation Functions . Num- ber 50 in De Gruyter Studies in Mathematics. Walter de Gruyter, Berlin/Boston, Mar. 2013. ISBN 978-3-11-022399-6

  12. [20]

    I. J. Schoenberg. Metric spaces and positive definite functions. Transactions of the American Mathematical Society , 44(3):522–536,

  13. [21]

    I. J. Schoenberg. Positive definite functions on spheres. Duke Mathematical Journal, 9(1):96–108, Mar. 1942. ISSN 0012-7094, 1547-7398. doi: 10.1215/ S0012-7094-42-00908-6

  14. [22]

    Sch¨ olkopf and A

    B. Sch¨ olkopf and A. J. Smola.Learning with Kernels: Support Vector Ma- chines, Regularization, Optimization, and Beyond . MIT Press, Cambridge, Mass., 1 edition, 2002. ISBN 978-0-262-53657-8

  15. [23]

    J. B. Simon, S. Anand, and M. Deweese. Reverse Engineering the Neural Tangent Kernel. In Proceedings of the 39th International Conference on Machine Learning, pages 20215–20231. PMLR, June 2022. URL https: //proceedings.mlr.press/v162/simon22a.html

  16. [24]

    Williams

    C. Williams. Computing with Infinite Networks. In Advances in Neural Information Processing Systems , volume 9. MIT Press,

  17. [25]

    G. Yang. Tensor Programs I: Wide Feedforward or Recur- rent Neural Networks of Any Architecture are Gaussian Pro- cesses. In Advances in Neural Information Processing Systems , vol- ume 32, Vancouver, Canada, 2019. Curran Associates, Inc. URL https://proceedings.neurips.cc/pap...

  18. [26]

    ⇐”: This follows directly from the fact that U is a linear isometry that pre- serves norms and inner products. “⇒

    G. Yang. Tensor Programs II: Neural Tangent Kernel for Any Architecture, Nov. 2020. URL http://arxiv.org/abs/2006.14548. 12 A Appendix A.1 Proofs for Section 2 Since we are interested in characterizing isotropic kernels on Rd we consider Sd−1, whereas the source we build our p...

  19. [31]

    Case d< ∞: For all n∈ N and allr,s∈ (0,∞) we have by Theorem A.1 α(d) n (r,s ) =Z ∫1 −1 =fS(r,s,ρ) /bracehtipdownleft/bracehtipupright/bracehtipupleft/bracehtipdownright κ(r,s,rsρ ) ˆC(λ) n (ρ)w(ρ)dρ. To extendα(d) n to a continuous function on [0,∞)2, which yields (5) by defi...

  20. [32]

    (ii), (iii), (iv)⇒ (i)

    Case d =∞: First note that the α(∞) n are continuous positive definite kernels on (0,∞) [10, Example 2.4]. For the extension we select two orthonormal vectorse1,e 2 and observe K(re1,re 1)−K(re1,re 2) = ∞∑ n=0 α(∞) n (r,r )⟨e1,e 1⟩n− ∞∑ n=0 α(∞) n (r,r )⟨e1,e 2⟩n = ∞∑ n=1 α(∞)...

  21. [33]

    (i)⇒ (ii), (iii)

    on (0,∞)× Sd−1. Since these conditions are satisfied for K>0 if they are satisfied for K we thereby see that K>0 is strictly positive definite on Rd\{ 0}. 17 If m> 1 the second term in (13) is therefore non-zero. If m = 1 then the first term is m∑ i,j=1 cicjK0(xi,xj) =|c1|2α(d...

  22. [34]

    Remark A.4 (Universal kernels)

    replaced by Theorems 3.6, 3.8 and 3.9 respectively. Remark A.4 (Universal kernels) . For future work characterizing universal isotropic kernels we point out their characterization on the sphere by Micchelli et al. [14, Theorem 10]. It might be possible to extend this result to...

  23. [940]

    doi: 10.1007/s00365-016-9323-9

  24. [1807]

    doi: 10.1007/BF01452844. 10

  25. [1938]

    URL https://community.ams.org/journals/tran/1938-044-03/ S0002-9947-1938-1501980-0/S0002-9947-1938-1501980-0.pdf

  26. [1996]

    URL https://proceedings.neurips.cc/paper/1996/hash/ ae5e3ce40e0404a45ecacaaf05e5f735-Abstract.html

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.