Pith. sign in

REVIEW 3 major objections 9 minor 71 references

Categorical and geometric methods in statistical, manifold, and machine learning

T0 review · 3 major / 9 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read A map $T : X \to P(Y)$ is the correct conditional model for a joint distribution exactly when its graph sends the input marginal to the full joint; on this identity the paper builds correct loss functions and learnability.

desk verdict A competent survey of the authors' categorical framework and geometric kernels, not a new-results paper; Example 2.9 overclaims the 0-1 loss and the learnability theorem rests on deferred self-citations. read the letter →

arxiv 2505.03862 v1 pith:QJFYUEW4 submitted 2025-05-06 stat.ML cs.LGmath.CTmath.DGmath.STstat.TH

classification stat.MLcs.LGmath.CTmath.DGmath.STstat.TH MSC 68T0546E2260B05
keywords probabilisticmorphismsMarkovkernelsregularconditionalprobabilitymeasuressupervisedlearningempiricalriskminimizationkernelmeanembeddingsLog-Hilbert-Schmidtmetricmanifold
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that the category of probabilistic morphisms — Markov kernels treated as arrows between measurable spaces — gives supervised learning a unified formal setting. Its central technical claim is Theorem 2.6: a measurable map $T : X \to P(Y)$ is a regular conditional probability measure for the joint distribution $\mu$ exactly when pushing the input marginal $\mu_X$ through the graph of $T$ reproduces $\mu$, i.e. $(\Gamma_T)_*\mu_X = \mu$. From this equivalence the paper derives correct loss functions, in particular a kernel-mean-embedding loss whose minimizers are precisely the correct conditional models, and proves that under a compactness and uniform-Lipschitz condition there exists a uniformly consistent regularized empirical-risk-minimization algorithm for an overparameterized learning model (Proposition 2.18). The same language frames classical regression, since with zero-mean noise the regression function is the conditional mean of the labels. The payoff, if the framework is right, is that correct loss functions and learnability follow from one structural identity instead of being assumed problem by problem.

What carries the argument

The machinery carrying the argument is the graph of a probabilistic morphism. Given a measurable map $T : X \to P(Y)$, its graph $\Gamma_T$ joins the identity on $X$ with $T$ to form a probabilistic morphism $\Gamma_T : X \rightsquigarrow X \times Y$, and the identity doing the work is $(\Gamma_T)_*\mu_X = \mu$: Theorem 2.6 shows this equality holds exactly when $T$ is a regular conditional probability measure for $\mu$, turning the search for the correct stochastic model of labels given inputs into the search for solutions of one measure-valued equation. The learnability proof is carried by the regularizer $W(f) = \|f\|_M + L(f) + \|\Gamma_f\|_{\widetilde{K}_3,\widetilde{K}_1}$, whose sublevel sets are compact (Lemma 2.21), together with a classical consistency estimate for stochastic ill-posed problems (Proposition 2.20) that bounds the probability a regularized minimizer is far from the true conditional map by the probability that the empirical data are far from the true joint distribution; condition (L) makes that bound vanish uniformly as the sample size grows. For the geometric half, the parallel identity is the classical characterization that $\exp(-\gamma d^2)$ is positive definite for every $\gamma > 0$ exactly when the metric space embeds isometrically into a Hilbert space, which rules out squared-distance Gaussian kernels on curved Riemannian manifolds and motivates the Log-Euclidean and Log-Hilbert–Schmidt constructions with their induced vector-space structures.

What would settle it

Take $X = [0,1]$, $Y = \{0,1\}$, and a joint distribution whose conditional mean is a step function with a jump, so the conditional measure is not Lipschitz at the jump, and run the $(C,\Gamma)$-regularized ERM algorithm from Proposition 2.18 with the kernel-mean-embedding loss; if uniform consistency still holds, condition (L) is stronger than necessary, and if it fails, the compactness-and-Lipschitz assumption is doing real work. For Theorem 2.6 itself the claim follows directly from the definition of a regular conditional probability measure, so the empirically checkable part is whether the constructed losses are actually minimized only at correct conditionals on finite samples, which can be tested directly by estimating both sides of the graph equation.

Watch

Extended reading notes

Core claim

On its own terms, the paper's central discovery is that regular conditional probability measures are captured by a single graph equation. For $\mu \in P(X \times Y)$, a measurable map $T : X \to P(Y)$ is a regular conditional probability measure with respect to the projection onto $X$ if and only if $(\Gamma_T)_*\mu_X = \mu$: the graph $\Gamma_T$, the join of the identity on $X$ with $T$, pushes the input marginal forward to the full joint measure (Theorem 2.6). Two conditional models for the same $\mu$ agree $\mu_X$-almost everywhere, and any map equal $\mu_X$-almost everywhere to a regular conditional is again one. Because the condition is an equality of measures, it converts into losses: the quadratic loss on $\{0,1\}$ labels is correct because its minimizer is the conditional mean (Example 2.9), and the kernel-mean-embedding loss $R_{K_1}(h,\mu) = \|M_{K_1}((\Gamma_h)_*\mu_X) - M_{K_1}(\mu)\|_{\widetilde{K}_1}$ is correct whenever the kernel mean embedding is injective (Example 2.11). For the overparameterized model with hypothesis space $C_{\mathrm{Lip}}(X, P(Y)_{\widetilde{K}_2})$, the loss $R_{K_1}$, and a family of joint distributions satisfying condition (L) — weak*-compactness plus uniformly bounded Lipschitz constants of the conditional measures — Proposition 2.18 asserts the existence of a uniformly consistent $(C,\Gamma)$-regularized ERM algorithm, obtained by combining the graph equation in variational form with a classical consistency estimate for stochastic ill-posed problems and the compact sublevel sets of a Lipschitz-type regularizer (Lemma 2.21).

Load-bearing premise

The load-bearing premise of the learnability claim is condition (L): all possible joint distributions must lie in a weak*-compact family whose conditional label distributions all have Lipschitz constants within one fixed finite interval $[a,b]$; if real data can contain arbitrarily sharp or rough conditional distributions, the uniform consistency guarantee of Proposition 2.18 does not follow.

Editorial extensions

If this is right

  • Correct loss functions can be built directly from the graph equation rather than from an instantaneous loss: the kernel-mean-embedding loss $R_{K_1}(h,\mu) = \|M_{K_1}((\Gamma_h)_*\mu_X) - M_{K_1}(\mu)\|_{\widetilde{K}_1}$ is correct whenever the embedding is injective, and it is empirically definable, so the framework yields losses that the instantaneous-loss route does not naturally produce.
  • Overparameterized supervised learning can be learnable: under condition (L), the model $(X, Y, C_{\mathrm{Lip}}(X, P(Y)_{\widetilde{K}_2}), R_{K_1}, P_{X\times Y})$ admits a uniformly consistent $(C,\Gamma)$-regularized ERM algorithm, and the same holds for discriminative subspaces $H \subset C_{\mathrm{Lip}}(X,Y)$ (Proposition 2.18).
  • Classical regression falls out of the framework: when the noise has zero mean, the regression function is the conditional mean $r_\mu(x) = \int_Y y\, d\mu_{Y|X}(y|x)$, so the categorical formulation subsumes standard supervised learning as a special case (Example 2.7).
  • On a geodesically complete Riemannian manifold, the kernel $\exp(-\gamma d^2)$ is positive definite for all $\gamma > 0$ if and only if the manifold is isometric to Euclidean space; consequently the affine-invariant and Bures–Wasserstein distances on positive definite matrices cannot yield such kernels, while the Log-Euclidean and Log-Hilbert–Schmidt kernels are unconditionally positive definite (
  • Riemannian geometry can be learned from point clouds: a point-cloud Laplacian converges in probability to the Laplace–Beltrami operator at a suitable bandwidth scaling (Theorem 4.1), and a Riemannian manifold can be reconstructed from noisy intrinsic distances between sample points.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The graph equation suggests a model-criticism tool the paper does not develop: for any candidate $h$, the sample estimate of $\|M_{K_1}((\Gamma_h)_*\mu_X) - M_{K_1}(\mu)\|$ is a finite-sample diagnostic that is zero in expectation only at a correct conditional model, so it could be used to test generative models outside the paper's Lipschitz setting.
  • Condition (L) is likely the assumption that fails first in high-dimensional practice: uniformity over all $\mu \in P_{X\times Y}$ requires one fixed interval $[a,b]$ to contain the Lipschitz constants of every possible conditional distribution, and sharp decision boundaries would exceed any fixed bound; a local or scale-dependent Lipschitz assumption would be a natural weakening the paper does not
  • The Log-Euclidean and Log-Hilbert–Schmidt constructions encode a transferable recipe: whenever a set of non-Euclidean objects carries a commutative group operation making it a vector space, every inner-product kernel from Euclidean theory can be carried over to that set; the same recipe could apply to other cones of positive operators or matrix groups.
  • Since a posterior distribution is itself a regular conditional probability measure, the same graph equation would characterize Bayesian posteriors given data; the paper cites this connection only in passing, so formalizing Bayesian consistency through the graph identity is a natural next step.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 9 minor

Summary. The paper presents the category of probabilistic morphisms as a framework for supervised learning, reviews geometric kernel methods for positive definite matrices and operators, and surveys two manifold-learning results. Section 2 recalls regular conditional probability measures, proves a characterization via the graph probabilistic morphism (Theorem 2.6), defines generative models and correct loss functions, gives quadratic and kernel mean embedding examples, and states a learnability theorem for overparameterized models (Proposition 2.18) using condition (L) and Vapnik-Stefanyuk regularization. Section 3 summarizes positive definite kernels on SPD matrices and Log-Hilbert-Schmidt operators, and Section 4 reviews the Belkin-Niyogi Laplacian eigenmap justification and Fefferman et al. manifold reconstruction.

Significance. If the claims held, the categorical formalism would provide a unified way to discuss generative models, correct losses, and learnability, and the geometric kernel sections give a useful compendium of known constructions with explicit Gram-matrix formulas. The paper is largely expository: Theorem 2.6 is clean and correct, but the main learnability theorem and its supporting Lemma 2.21 are quoted from the authors' prior work [37], and Sections 3 and 4 mostly restate known results. The false 0-1 loss claim in Example 2.9 undercuts one of the two concrete examples of correct loss functions. There are no new experiments and no fully self-contained proofs for the central theorem. The paper is best classified as a survey/preview of the forthcoming book [40], and its final claim of having demonstrated usefulness is stronger than what is established in this manuscript.

major comments (3)
  1. [Example 2.9, last sentence] The assertion that the instantaneous 0-1 loss L_{0,1}(x,y,h)=d_{0-1}(y,h(x)) generates a correct loss function is false under Definition 2.8. For X={x0}, Y={0,1}, and μ=0.6δ_{(x0,1)}+0.4δ_{(x0,0)}, the regular conditional probability measure is μ_{Y|X}(x0)=0.6. The 0-1 risk of a hard classifier h∈Meas(X,Y) equals 0.6 if h(x0)=0 and 0.4 if h(x0)=1, so the unique minimizer is the constant classifier h≡1, not the conditional probability 0.6. Extending the loss to randomized classifiers h∈Meas(X,P(Y)) does not help, because the expected 0-1 risk is minimized at a deterministic mode. Hence no extension of this instantaneous loss satisfies the argmin condition in Definition 2.8. The quadratic-loss part of Example 2.9 and the kernel mean embedding example 2.11 are not affected, but the 0-1 statement must be corrected or removed.
  2. [Proposition 2.18 and condition (L)] The paper's central learnability theorem is not proved in this manuscript: Lemma 2.21 is deferred to [37, Proposition 6.1], Proposition 2.22 to [37, Theorem 6.5], and Proposition 2.18 is stated as [37, Corollary 6.3]. The outline in Section 2.4.2 does not supply the compactness and rate arguments needed to pass from Proposition 2.22 to uniform consistency under condition (L), nor does it specify how C and Γ are chosen. In addition, condition (L) defines L(μ_Y|X) as a function of μ, but regular conditional measures are only unique μ_X-a.e.; since P_Lip assumes full support, this is still an ambiguity unless a canonical representative is chosen. The authors should either state explicitly that this is a survey with proofs in [37], or provide the missing arguments. As written, the final sentence of Section 5 overstates what is demonstrated here.
  3. [Remark 2.19] The proposed concrete example of a family satisfying condition (L) is incomplete. The text says P_X×Y consists of all μ_f=f dxdy for which there exist c1,c0>0 with f∈C^1(X×Y), L(f)≤c1, and c1≥f≥c0. If c0,c1 are allowed to depend on f, the family is not compact in the weak*-topology and the Lipschitz constants L(μ_Y|X) are not uniformly bounded in a fixed interval [a,b] as required by condition (L). If c0,c1 are meant to be fixed uniform constants, the remark should say so and provide the verification. Without this, the example does not demonstrate that condition (L) is satisfiable in a nontrivial setting.
minor comments (9)
  1. [Abstract and Section 2.1] The abstract contains a broken phrase 'as well a s'; Section 2.1's first bullet says the spaces S(Y), M(Y), P(Y) are on X, which should be on Y.
  2. [Definition 2.8] In the definition of instantaneous loss, 'for any h∈R' should be 'for any h∈H', and the sentence about E_μ switches between f and h; please correct the variable names.
  3. [Example 2.9] There is a typo 'P(X×Y 9)' in the first paragraph, and the sentence 'any minimizer R_L^μ to Meas(X,[0,1])' should read 'any minimizer of R_L^μ over Meas(X,[0,1])'.
  4. [Section 2.4 heading] The heading 'Leanability' should be 'Learnability'.
  5. [Definition 2.16] The phrase 'a generative model ofs supervised learning' contains a typo 'ofs'.
  6. [Remark 2.12 and Proposition 2.13] Remark 2.12(1) ends with the incomplete sentence 'They proved the following beautiful result on estimating probab.'; the proposition should be introduced by a complete sentence.
  7. [Lemma 2.21] The term ||Γ_f||_{K3,K1} in equation (2.24) uses an undefined kernel or metric K3; please define it or remove it if it is a typo.
  8. [Proposition 2.22] The domain of K1 is written as 'X×Y ) × (X×Y)', which contains a misplaced parenthesis; it should be (X×Y)×(X×Y).
  9. [Section 5] The word 'languaguage' is a typo for 'language', and the spelling of 'Vapnik-Stefanyuk' is inconsistent with 'Stefanyuk' used in Proposition 2.20 and the references.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation found; Section 2's claims are either direct reformulations, purpose-built loss constructions, or externally checkable theorems cited from prior work, with a separate non-circular error in Example 2.9.

full rationale

After walking the derivation chain, I find no circular step. Theorem 2.6 is the disintegration formula (2.1) rewritten through the graph morphism; equation (2.4) is equivalent to (2.1) by definition of the pushforward along the graph, so the paper's 'characterization' is a reformulation, not a conclusion smuggled from itself. Example 2.11 defines R_K(h,μ)=||M_K((Γ_h)_*μ_X−μ)||; because M_K is injective, the loss vanishes exactly when (Γ_h)_*μ_X=μ, which by Theorem 2.6 is exactly the regular conditional condition, so the loss is correct by construction—a designed equivalence, not a predicted output. Proposition 2.18 is stated as [37, Corollary 6.3] with only an outline, and Lemma 2.21 and Proposition 2.22 are deferred to [37]; this is a load-bearing self-citation in the exposition, but the cited theorem is an externally checkable mathematical statement whose assumptions (compactness and condition (L)) do not include the conclusion, so under the hard rules it is independent support and does not itself constitute circularity. Sections 3 and 4 rely on external results (Schoenberg, Belkin–Niyogi, Fefferman et al., Sra) rather than on the authors' own fitted quantities. Separately, the closing sentence of Example 2.9 asserting that the 0-1 loss generates a correct loss function is a correctness error: for a Bernoulli label with P(Y=1)=0.6, the constant classifier 1 has 0-1 risk 0.4, while the conditional value 0.6 has larger expected 0-1 risk, so the 0-1 minimizer is the mode, not the regular conditional probability. That is a mathematical mistake, not a circular derivation, and it does not raise the circularity score.

Assumptions & free parameters 0 free parameters · 7 assumptions · 0 invented entities

No parameters are fitted to data in this paper; it is a survey. Kernel widths, regularization parameters, and Lipschitz bounds appear as existentially quantified objects in theorems, not as fitted values. No new physical or mathematical entities are introduced: the category Probm is from Lawvere/Chentsov, and the Log-Hilbert-Schmidt metric and unitized Hilbert-Schmidt operators come from Larotonda [35] and [45].

assumptions (7)
  • standard math Existence of regular conditional probability measures for probability measures on X × Y when Y is Polish.
    Invoked in Section 2.3.1 before Definition 2.8, citing Bogachev and Leao-Fragoso-Ruffino. Needed to define the target of learning.
  • standard math Gaussian kernel mean embeddings are injective on probability measures and metrize the weak*-topology.
    Used in Example 2.11 and Section 2.4.2 for the loss R_{K1}; cited to Sriperumbudur [57].
  • standard math Schoenberg's theorem: a metric space embeds isometrically into Hilbert space iff squared distance is negative definite iff the Gaussian kernel is positive definite for all gamma > 0.
    Basis for Theorems 3.1 to 3.3 on Gaussian kernels on Riemannian manifolds.
  • standard math Fan's inequality on log-concavity of the determinant and the matrix-variate Gamma integral representation.
    Used in the proof of Theorem 3.5 for the log-det kernel.
  • domain assumption The existence of a 'correct' loss function as defined in Definition 2.8 requires that the hypothesis class H can be extended to ~H containing regular conditional measures.
    Definition 2.8 imposes condition (2); the paper assumes such ~H exists for the models considered.
  • ad hoc to paper Condition (L) on P_{X×Y}: compactness in weak*-topology and uniform bounded Lipschitz constants of regular conditional measures.
    Stated in Section 2.4.2 as a condition for Proposition 2.18. It is a strong regularity assumption, not derived from first principles.
  • ad hoc to paper Lemma 2.21: the regularizer W(f) = ||f||_M + L(f) + ||Γ_f|| has compact sublevel sets.
    Needed to apply Vapnik-Stefanyuk's theorem; proof deferred to [37, Prop 6.1].

how reviews work

0 comments
Cite this review

Pith. "Pith review of Categorical and geometric methods in statistical, manifold, and machine learning." pith.science (2026). https://pith.science/paper/QJFYUEW4

@misc{pith2026250503862,
  author       = {Pith},
  title        = {Pith review of: Categorical and geometric methods in statistical, manifold, and machine learning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QJFYUEW4}},
  note         = {Machine review of arXiv:2505.03862}
}
read the original abstract

We present and discuss applications of the category of probabilistic morphisms, initially developed in \cite{Le2023}, as well as some geometric methods to several classes of problems in statistical, machine and manifold learning which shall be, along with many other topics, considered in depth in the forthcoming book \cite{LMPT2024}.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

71 extracted references · 70 canonical work pages

  1. [37]

    H. V. Lˆ e, Supervised learning with probabilistic morp hisms and kernel mean em- beddings, arXiv:2305.06348

  2. [45]

    H.Q. Minh, M. San Biagio, and V. Murino, Log-Hilbert-Sc hmidt metric between positive definite operators on Hilbert spaces. Advances in n eural information pro- cessing systems 27 (2014)

  3. [9]

    Belkin and P

    M. Belkin and P. Niyogi, Towards a theoretical foundatio n for Laplacian-based manifold methods. Journal of Computer and System Sciences, 74(8), 1289-1308, 2008

  4. [40]

    H. V. Lˆ e, H. Q. Minh, F. Protin, W. Tuschmann, Mathemati cal Foundations of Machine Learning (book in preparation, to be published by Sp ringer in 2025)

  5. [1]

    N. Ay, J. Jost, H. V. Lˆ e and L. Schwachh¨ ofer, Informatio n Geometry. Springer, 2017

  6. [2]

    Aronszajn, Theory of reproducing kernels

    N. Aronszajn, Theory of reproducing kernels. Trans. Ame r. Math. Soc. 68 (1950),337-404

  7. [3]

    Arsigny, P

    V. Arsigny, P. Fillard, X. Pennec, and N. Ayache. Geometr ic means in a novel vector space structure on symmetric positive-definite matr ices. SIAM J. on Matrix An. and App., 29(1):328-347, 2007

  8. [4]

    C. Berg, J. P. R. Christensen, and P. Ressel, Harmonic Ana lysis on Semigroups. Springer, 1984

Show all 71 references
  1. [5]

    Berner, P

    J. Berner, P. Grohs, G. Kutyniok, P. Petersen, The Modern Mathematics of Deep Learning, in: Mathematical Aspects of Deep Learning”, Edit ed by P. Grohs, G. Kutyniok, Cambridge University Press 2022, 1-111, arXiv:2 105.04026

  2. [6]

    Bhatia, T

    R. Bhatia, T. Jain, and Y. Lim. On the Bures–Wasserstein d istance between positive definite matrices. Expositiones Mathematicae, 37 (2), 165-191, 2019

  3. [7]

    Belkin, P

    M. Belkin, P. Niyogi, Laplacian eigenmaps and spectral t echniques for embedding and clustering, Advances in Neural Information Processing Systems 14(6):585-591, 2001

  4. [8]

    Belkin and P

    M. Belkin and P. Niyogi, Laplacian eigenmaps for dimensi onality reduction and data representation. Neural computation, 15(6), pp.1373- 1396, 2003. CATEGORICAL AND GEOMETRIC METHODS IN STATISTICAL LEARNING 35

  5. [10]

    V. I. Bogachev, Measure Theory I, II. Springer, 2007

  6. [11]

    Berlinet and C

    A. Berlinet and C. Thomas-Agnan, Reproducing Kernel Hi lbert Spaces in Prob- ability and Statistics. Kluwer Academic Publishers, 2004

  7. [12]

    Bregman, The relaxation method of finding the commo n point of convex sets and its application to the solution of problems in conve x programming

    L.M. Bregman, The relaxation method of finding the commo n point of convex sets and its application to the solution of problems in conve x programming. USSR computational mathematics and mathematical physics, 7(3) :200-217, 1967

  8. [13]

    Burago, S

    D. Burago, S. Ivanov and Y. Kurylev, A graph discretizat ion of the Laplace- Beltrami operator, J. Spectr. Theory 4 (2014), 675-714

  9. [14]

    Chebbi and M

    Z. Chebbi and M. Moakher. Means of Hermitian positive-d efinite matrices based on the log-determinant α -divergence function. Linear Algebra and its Applica- tions, 436 (7):1872-1889, 2012

  10. [15]

    N. N. Chentsov, The categories of mathematical statist ics. (Russian) Dokl. Akad. Nauk SSSR 164 (1965), 511-514

  11. [16]

    N. N. Chentsov, Statistical Decision Rules and Optimal Inference. Translation of mathematical monographs, AMS, Providence, Rhode Island , 1982, translation from Russian original, Nauka, Moscow, 1972

  12. [17]

    Cucker and S

    F. Cucker and S. Smale, On mathematical foundations of l earning. Bulletin of AMS, 39 (2002), 1-49

  13. [18]

    Dodziuk, Finite-Difference Approach to the Hodge The ory of Harmonic Forms, Amer

    J. Dodziuk, Finite-Difference Approach to the Hodge The ory of Harmonic Forms, Amer. J. of Math. 98, No. 1, 79-104

  14. [19]

    Dodziuk, V

    J. Dodziuk, V. Patodi, Riemannian structures and trian gulations of manifolds, J. Ind. Math. Soc. 40 (1976), p. 1-52

  15. [20]

    Gordon, D

    C. Gordon, D. L. Webb, and S. Wolpert, One cannot hear the shape of a drum, Bulletin (New Series) of the AMS, Volume 27 1992), Number 1, p . 134-138

  16. [21]

    K. Fan. On a theorem of Weyl concerning eigenvalues of li near transformations: II. Proceedings of the National Academy of Sciences of the Unite d States of America, 36 (1):31, 1950

  17. [22]

    Faraut, and A

    J. Faraut, and A. Kor´ anyi, Analysis on symmetric cones . Oxford University Press, 1994

  18. [23]

    I: The geometric Whitney problem, Foundations of Computational M athematics, 20 (5) (2020), p

    Fefferman, Charles and Ivanov, Sergei and Kurylev, Yaro slav and Lassas, Matti and Narayanan, Hariharan, Reconstruction and interpolati on of manifolds. I: The geometric Whitney problem, Foundations of Computational M athematics, 20 (5) (2020), p. 1035–1133

  19. [24]

    Recon- struction of a Riemannian manifold from noisy intrinsic dis tances

    Fefferman, Charles; Ivanov, Sergei; Lassas, Matti; Nar ayanan, Hariharan. Recon- struction of a Riemannian manifold from noisy intrinsic dis tances. SIAM J. Math. Data Sci. 2 (3) (2020) p. 770-808

  20. [25]

    Feragen, F

    A. Feragen, F. Lauze, F., and S. Hauberg. Geodesic expon ential kernels: When curvature and linearity conflict. In Proceedings of the IEEE conference on com- puter vision and pattern recognition, pp. 3032-3042, 2015

  21. [26]

    Fritz, A synthetic approach to Markov kernel, condit ional independence and theorem of sufficient statistics

    T. Fritz, A synthetic approach to Markov kernel, condit ional independence and theorem of sufficient statistics. Adv. Math. 370, 107239 (202 0), arXiv:1908.07021

  22. [27]

    Giry, A categorical approach to probability theory, In: B

    M. Giry, A categorical approach to probability theory, In: B. Banaschewski, (edi- tor) Categorical Aspects of Topology and Analysis, Lecture Notes in Mathematics 915, 68- 85, Springer, 1982

  23. [28]

    Grenander and M

    U. Grenander and M. I. Miller, Pattern theory: From Repr esentation to Inference, Oxford University Press, 2007

  24. [29]

    Hastie, R

    T. Hastie, R. Tibshirani and J. Friedman, The Elements o f Statistical Learning: Data Mining, Inference, and Prediction. Springer 2008. 36 HˆONG V ˆAN L ˆE, H `A QUANG MINH, FREDERIC PROTIN, AND WILDERICH TUSCHMANN

  25. [30]

    Jayasumana, R

    S. Jayasumana, R. Hartley, M. Salzmann, H. Li, and M. Har andi. Kernel meth- ods on Riemannian manifolds with Gaussian RBF kernels. IEEE transactions on pattern analysis and machine intelligence, 37(12), pp.246 4-2477, 2015

  26. [31]

    Joharinad and J

    P. Joharinad and J. Jost, Mathematical Principles of To pological and Geometric Data Analysis, Springer 2023

  27. [32]

    J. Jost, H. V. Lˆ e, and T. D. Tran, Probabilistic morphis ms and Bayesian non- parametrics. Eur. Phys. J. Plus 136, 441 (2021), arXiv:1905 .11448

  28. [33]

    Kadison and J.R

    R.V. Kadison and J.R. Ringrose, Fundamentals of the the ory of operator algebras. Volume I: Elementary Theory. Academic Press, 1983

  29. [34]

    Lang, Fundamentals of Differential Geometry

    S. Lang, Fundamentals of Differential Geometry. Spring er, 1999

  30. [35]

    Larotonda

    G. Larotonda. Nonpositive curvature: A geometrical ap proach to Hilbert-Schmidt operators. Differential Geometry and its Applications, 25: 679-700, 2007

  31. [36]

    W. F. Lawvere, The category of probabilistic mappings ( 1962). Unpublished, Available at https://ncatlab.org/nlab/files/lawverepro bability1962.pdf

  32. [38]

    Lˆ e, Probabilistic morphisms and Bayesian superv ised learning, Math

    H.V. Lˆ e, Probabilistic morphisms and Bayesian superv ised learning, Math. Sbornik, N5, vol. 216 (2025), 161-180

  33. [39]

    Fragoso and P

    D Leao Jr., M. Fragoso and P. Ruffino, Regular conditional probability, disinte- gration of probability and Radon spaces. Proyecciones vol. 23 (2004)Nr. 1, Uni- versidad Catolica Norte, Antofagasta, Chile, 15-29

  34. [41]

    P. Li, Q. Wang, W. Zuo, and L. Zhang. Log-Euclidean kerne ls for sparse represen- tation and dictionary learning. In International Conferen ce on Computer Vision (ICCV), pages 1601-1608, 2013

  35. [42]

    Lopez-Paz, K

    D. Lopez-Paz, K. Muandet, B. Sch¨ olkopf, and I. Tolstik hin. Towards a learning theory of cause-effect inference. In Proceedings of the 32nd International Confer- ence on Machine Learning (ICML2015), 2015

  36. [43]

    Malag` o, L., L

    L. Malag` o, L., L. Montrucchio, and G. Pistone, Wassers tein Riemannian geometry of Gaussian densities. Information Geometry, 1, 137-179, 2 018

  37. [44]

    Mathai, S.B

    A.M. Mathai, S.B. Provost, and H.J. Haubold, Multivari ate statistical analysis in the real and complex domains. Springer Nature, 2022

  38. [46]

    Minh and V

    H.Q. Minh and V. Murino. Covariances in computer vision and machine learning. Springer Synthesis Lectures on Computer Vision, 2018

  39. [47]

    Mohri, A

    M. Mohri, A. Rostamizadeh, A. Talwalkar, Foundations o f Machine Learning. MIT Press, 2nd Edition, 2018

  40. [48]

    Morse, and R

    N. Morse, and R. Sacksteder, Statistical isomorphism. Ann. Math. Statist. 37, 1 (1966), 203–214

  41. [49]

    G. Mostow. Some new decomposition theorems for semi-si mple groups. Memoirs of the American Mathematical Society, 14:31-54, 1955

  42. [50]

    Nash, The imbedding problem for Riemannian manifold s, Ann

    J. Nash, The imbedding problem for Riemannian manifold s, Ann. Math. 63(1956), 383-396

  43. [51]

    V. I. Paulsen, M. Raghupathi, An introduction to the the ory of reproducing kernel Hilbert spaces. Cambridge Studies in Advanced Mathematics , 152. Cambridge University Press, Cambridge, 2016

  44. [52]

    Petryshyn, Direct and iterative methods for the so lution of linear operator equations in Hilbert spaces

    W.V. Petryshyn, Direct and iterative methods for the so lution of linear operator equations in Hilbert spaces. Transactions of the American M athematical Society, 105:136-175, 1962. CATEGORICAL AND GEOMETRIC METHODS IN STATISTICAL LEARNING 37

  45. [53]

    Schoenberg

    I.J. Schoenberg. Metric spaces and positive definite fu nctions. Transactions of the American Mathematical Society, 44(3), 522-536, 1938

  46. [54]

    C. L. Siegel. Symplectic geometry. American Journal of Mathematics, 65(1):1-86, 1943

  47. [55]

    S. Sra. A new metric on the manifold of kernel matrices wi th application to matrix geometric means. In Advances in Neural Information Process ing Systems (NIPS), pages 144-152, 2012

  48. [56]

    S. Sra. Positive definite matrices and the S-divergence . Proceedings of the Amer- ican Mathematical Society, 144(7):2787-2797, 2016

  49. [57]

    Sriperumbudur, On the optimal estimation of probabi lity measures in weak and strong topologies

    B. Sriperumbudur, On the optimal estimation of probabi lity measures in weak and strong topologies. Bernoulli, 22(3):1839-1893, 08, 20 16

  50. [58]

    disorder

    A. R. Stefanyuk, Estimation of the likelihood ratio fun ction in the “disorder” problem of random process, Autom. Remote. Control 9 (1986), 53-59

  51. [59]

    Steinwart and A

    I. Steinwart and A. Christmann. Support vector machine s. Springer, 2008

  52. [60]

    Takatsu, Wasserstein geometry of Gaussian measures

    A. Takatsu, Wasserstein geometry of Gaussian measures . Osaka Journal of Math- ematics, 48:1005-1026, 2011

  53. [61]

    Theodoridis, Machine Learning, a Bayesian and Optim ization Perspective

    S. Theodoridis, Machine Learning, a Bayesian and Optim ization Perspective. Aca- demic Press, 2015

  54. [62]

    A. B. Tsybakov, Introduction to Nonparametric Estimat ion. Springer, 2009

  55. [63]

    A. W. van der Vaart, J.A. Wellner, Weak convergence and E mpirical Processes. 2nd Edition. Springer, 2023

  56. [64]

    Vapnik, Statistical Learning Theory

    V. Vapnik, Statistical Learning Theory. John Willey & S ons, 1998

  57. [65]

    Vapnik, The Nature of Statistical Learning Theory, S pringer, 2nd Edition, 2000

    V. Vapnik, The Nature of Statistical Learning Theory, S pringer, 2nd Edition, 2000

  58. [66]

    Vapnik and R

    V. Vapnik and R. Izmailov, Synergy of Monotonic Rules, J ournal of Machine Learning Research 17 (2016) 1-33

  59. [67]

    Vapnik and R

    V. Vapnik and R. Izmailov, Rethinking statistical lear ning theory: learning using statistical invariants, Machine Learning (2019) 108:381- 423

  60. [68]

    Vapnik and A

    V. Vapnik and A. Stefanyuk, Nonparametric methods for e stimating probability densities, Automation and Remote Control, 8 (1978), 38-52

  61. [69]

    Wald, Statistical Decision Functions

    A. Wald, Statistical Decision Functions. Wiley, New Yo rk; Chapman & Hall, London, 1950

  62. [70]

    J. Zhang. Divergence function, duality, and convex ana lysis. Neural Computation, 16 (1):159-195, 2004

  63. [71]

    S. K. Zhou and R. Chellappa. From sample similarity to en semble similarity: Probabilistic distance measures in reproducing kernel Hil bert space. IEEE Trans- actions on Pattern Analysis and Machine Intelligence, 28(6 ):917-929, 2006. Institute of Mathematics of the Czech Acad...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.