Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

Statistical Convergence Rates of Optimal Transport Map Estimation between General Distributions

T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read This paper proves convergence rates for plug-in optimal transport map estimators under far weaker assumptions than previously known, including heavy-tailed measures and non-convex or unbounded supports.

desk verdict Solid rates for OT map estimation under smoothness assumptions, but the C^2 Brenier potential requirement undermines the 'general distributions' framing. read the letter →

arxiv 2412.08064 v1 pith:Z5THE7DH submitted 2024-12-11 math.ST math.PRstat.MEstat.MLstat.TH

classification math.STmath.PRstat.MEstat.MLstat.TH MSC 62G2026D10
keywords OptimalTransportBrenierpotentialPoincaréinequalityplug-inestimatorsievesub-Weibulldistributionsheavy-tailedmultivariateranks
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper establishes non-asymptotic convergence rates for the plug-in optimal transport map estimator when the source measure is sub-Weibull or only polynomially tailed. Under $\alpha$-convexity and $\beta$-smoothness of the Brenier potential, the squared $L^2(P)$ error is $\tilde n^{-1/\gamma} + \tilde N^{-1/\gamma}$ up to logarithms for sub-Weibull $P$, and slower polynomial rates for heavy-tailed $P$. A new sieve plug-in estimator, which restricts the convex conjugate search to a growing ball, attains comparable rates without the convexity assumption, covering rank functions of normal and $t$-distributions. New Poincaré-type inequalities and a new $L^1(P)$ empirical process inequality extend the rates to heavier tails and to distributions lacking fourth moments. The results reduce the gap between existing theory and Brenier's theorem, which only requires two finite moments.

What carries the argument

The central object is the Brenier potential $\varphi_0$, the convex function whose gradient is the optimal transport map, paired with the semi-dual Kantorovich objective $P\varphi + Q\varphi^*$. The proof machinery is empirical-process control of the difference between the empirical and population objective, which requires bounding Hessians of candidate potentials and controlling the convex conjugates on the sample measure $Q_N$. The three new components are the sieve convex conjugate, which truncates the supremum in the Legendre transform to $B(0,M_n)$ and thereby acts as an implicit bound on the inverse map's range; Poincaré-type inequalities that control variances of differences of $\beta$-smooth potentials from local density boundedness and topological conditions on the support; and an $L^1(P)$ maximal inequality for function classes with polynomial envelopes.

What would settle it

Construct a valid transport problem in which the Brenier potential is convex but not twice differentiable (for example, a quantile function that is piecewise linear with a slope change) and check whether the plug-in estimator's squared $L^2(P)$ error still follows the stated $n^{-1/\gamma}+N^{-1/\gamma}$ rate. If the rate still holds, the smoothness premise is not load-bearing; if the error stalls or slows, the premise is necessary.

Watch

Extended reading notes

Core claim

The central claim is that the plug-in estimator $\nabla \hat\varphi_{n,N}$ obtained from the semi-dual objective $P_n \varphi + Q_N \varphi^*$ converges in $L^2(P)$ at rates governed by the covering entropy of the function class, the Hessian growth parameters $(\alpha,a)$ and $(\beta,a)$, and the tail of $P$. For sub-Weibull $P$ the rate is $\tilde n^{-1/\gamma} + \tilde N^{-1/\gamma}$ up to logarithmic factors, and for polynomial-tailed $P$ it degrades to a polynomial rate in the available moments. The sieve estimator $\nabla \tilde\varphi_{n,N}$, defined by $\varphi^{*,(n)}(y)=\sup_{x\in B(0,M_n)}\langle x,y\rangle-\varphi(x)$, removes the need for $\alpha$-convexity and is claimed to be the first to handle OT maps such as the rank functions of normal and $t$-distributions. With the new Poincaré-type inequalities, Donsker function classes achieve faster, nearly parametric rates. Theorem 3.20, based on the new $L^1(P)$ maximal inequality, gives rates when $P$ has only slightly more than two moments.

Load-bearing premise

The argument assumes the true transport potential is smooth enough to have a second derivative everywhere, with its Hessian bounded above and below by polynomial factors, even though the underlying theorem only guarantees a convex potential.

Editorial extensions

If this is right

  • OT-based multivariate ranks and quantiles now carry non-asymptotic convergence guarantees for distributions with unbounded or non-convex supports, not just compact convex ones.
  • Rank functions of normal and $t$-distributions, which fail $\alpha$-convexity, are covered by the sieve estimator.
  • The new Poincaré-type inequalities extend faster Donsker-class rates to sub-Weibull and polynomial-tailed measures, where the classical Poincaré inequality would force sub-exponential tails.
  • The $L^1(P)$ empirical process inequality extends rates to distributions with fewer than four moments, approaching the minimal moment requirement of Brenier's theorem.
  • A neural-network implementation of both estimators is provided, and simulations show the sieve estimator is more robust for heavy-tailed $t$ distributions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial extension: the sieve conjugate can be read as an implicit regularizer on the inverse map's range, so choosing $M_n$ as a quantile rather than the maximum may offer a bias-variance trade-off worth testing.
  • Editorial extension: the Poincaré-type inequalities proven from local density and support topology may transfer to other semi-dual estimation problems, such as generative modeling with Wasserstein objectives.
  • Editorial extension: the $L^1(P)$ maximal inequality may be usable outside optimal transport, for example in robust M-estimation where only $L^1$ integrability is available.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper studies L2(P) convergence rates of plug-in optimal transport map estimators. Under an (α,a)-convexity and (β,a)-smoothness condition on the Brenier potential, it establishes rates of order n^{-1/γ} + N^{-1/γ} for sub-Weibull P and slower polynomial rates for heavy-tailed P (Theorem 3.5). A sieve estimator (Eq. 3.8) is introduced to remove the (α,a)-convexity assumption (Theorems 3.16 and 3.17). New Poincaré-type inequalities (Section 4) and an L1(P) empirical process inequality (Lemma 3.19) extend results to Donsker classes and to distributions with fewer than four moments (Theorem 3.20). Numerical experiments illustrate the sieve estimator's robustness, especially for t-distributions.

Significance. If the theorems are correct, the paper substantially advances OT map estimation by relaxing support convexity and compactness and by accommodating heavy-tailed distributions, matching Divol et al. rates under weaker assumptions. The paper also contributes a new sieve estimator, novel Poincaré-type inequalities, and an L1 maximal inequality that may be of independent interest. The numerical experiments are thorough and include multivariate cases. However, all proofs are relegated to the supplementary material, which was not available for review, and several central assumptions conflict with the paper's stated generality, so the practical scope is narrower than advertised.

major comments (3)
  1. [Assumption 3.1 and Section 3.2] Assumption 3.1 requires the Brenier potential φ0 to be in C^2(R^d) with φ0(0)=0. This is not a consequence of Brenier's theorem and fails for elementary non-convex support pairs. For example, let P be uniform on [-2,-1] ∪ [1,2] and Q uniform on [0,1]. The monotone rearrangement is T(x) = (x+2)/2 on [-2,-1], T(x) = 1/2 on (-1,1), and T(x) = x/2 on [1,2]. This T is continuous but its derivative jumps at -1 and 1, so any convex antiderivative has no second derivative at those points. Hence Theorems 3.5, 3.8, 3.16, 3.17, and 3.20 are silent for this pair, despite P and Q being very simple distributions and the paper claiming to avoid convex and compact support assumptions. The authors should either weaken the C^2 condition or explicitly state this regularity restriction and adjust the claims in the abstract, introduction, and Table 1.
  2. [Table 1 and Theorems 3.8, 3.17] Table 1 indicates that for 'Our Results' no assumptions are made on the support or density of P and Q, but Theorems 3.8 and 3.17 require Assumption 3.7 (a Poincaré-type inequality). Proposition 4.2, which supplies sufficient conditions for Assumption 3.7, requires Assumption 4.1, including a connected, locally connected support, a Lipschitz domain for the interior (if not R^d), and a locally bounded density. Thus the support and density assumptions are still present for the faster-rate results, contradicting the table and the introductory claims of no restrictive probability-measure assumptions. This should be clarified by annotating which theorems require Assumption 4.1.
  3. [Supplementary material (all proofs)] All proofs and the key technical results Lemma 3.15, Proposition 4.2, and Lemma 3.19 are deferred to the supplementary material, which was not included in the manuscript under review. Since these results are load-bearing for Theorems 3.16, 3.8, and 3.20, respectively, and for the claimed Poincaré-type and L1 empirical process contributions, the referee cannot fully verify the correctness of the main claims from the submitted text. The authors should make the supplementary material available or include the proofs for the central lemmas in the main text.
minor comments (4)
  1. [Section 2.1] The phrase 'minimizing the the transportation cost' contains a duplicated article; it should read 'minimizing the transportation cost.'
  2. [Lemma 3.19] The condition following the integral reads 'R∞0 sqrt(H(x)/x) dx <∞ is integrable'; the phrase 'is integrable' is redundant and should be removed to avoid confusion.
  3. [Assumption 3.3] The notation E[log N(h, F, L2(Pn))] does not specify the probability space for the expectation; it should be stated that the expectation is taken with respect to the empirical measure Pn under the product measure P⊗n, or a suitable outer expectation should be used.
  4. [Notation after Eq. (1.1)] The notation 'a ≤log n b' is defined but the subscript 'log n' is not typeset cleanly in the text; consider using 'a ≤_{log n} b' or a separate sentence explaining the dependence on log n.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the rates are oracle-type bounds proved from stated entropy, convexity, and moment conditions; the self-citation to the authors' supplementary material is not load-bearing.

full rationale

The derivation chain is self-contained and non-circular. The central bounds (Theorems 3.5, 3.8, 3.16, 3.17, and 3.20) are oracle inequalities: for any candidate potential phi in the working class F, the L2(P) error is controlled by the objective gap S(phi)-S(phi0) plus an estimation term n^{-1/gamma} or n^{-2/(gamma+2)} (up to logarithms) obtained from covering/bracketing entropy (Assumption 3.3) and empirical-process maximal inequalities (Lemmas 3.15 and 3.19). No constant in these rates is fitted to data; the suppressed constants depend only on the measure, the smoothness coefficients, and the moments, not on the estimator output. Assumption 3.7 (Poincare-type inequality) is an explicit hypothesis in the Donsker-class theorems, and Section 4 then proves sufficient conditions (Assumption 4.1) under which it holds, so the result is conditional rather than a restatement of the conclusion. The sieve estimator's truncation radius M_n = max_i ||X_i|| is data-dependent, but its effect enters through explicit tail and entropy terms in Lemma 3.15, not through a fitted constant. The only self-citation is to the authors' own supplementary material (Ding et al. 2024) for a comparative proof in Remark 4.4; that proof is part of the same manuscript and is not used to import an unverified uniqueness or regularity theorem. The C^2 regularity assumption on the Brenier potential (Assumptions 3.1 and 3.14) is a genuine limitation on the advertised scope of 'general distributions', but it is an explicit theorem assumption rather than a circular reduction of the conclusions to their inputs. The paper also honestly records unproved engineering choices, such as sensitivity to ICNN architecture and alternative choices of M_n, which are limitations or open directions, not circular reasoning.

Assumptions & free parameters 0 free parameters · 7 assumptions · 0 invented entities

The central claim rests on standard optimal transport theory, on smoothness and entropy assumptions about the Brenier potential and the function class, and on the newly introduced Poincaré-type inequalities. None of these involve fitted constants or new physical entities. The most fragile additions are the C^2 regularity of the potential (Assumption 3.1) and the common-exponent smoothness/convexity condition (Assumption 3.2), which are materially stronger than the hypotheses of Brenier's theorem.

assumptions (7)
  • standard math Brenier's theorem: existence and uniqueness of a convex potential phi0 whose gradient is the OT map, given finite second moments and P absolutely continuous.
    Used throughout as the foundation for the plug-in estimators; cited from Brenier 1987/1991 as Theorem 2.1.
  • domain assumption Assumption 3.1: phi0 in C^2(R^d) with phi0(0)=0.
    All theorems in Section 3 assume this. C^2 regularity is not guaranteed by Brenier's theorem and is stronger than the theorem's hypotheses, so it is a load-bearing added condition.
  • domain assumption Assumption 3.2: phi0 and every phi in F are (beta,a)-smooth and (alpha,a)-convex with a common a >= 0.
    Used to control the objective via uniform localization; couples smoothness and convexity at the same polynomial scale. Applies to Theorems 3.5, 3.8, and 3.20.
  • domain assumption Assumption 3.3: covering or bracketing entropy of F scales as D_F h^{-gamma} log_+^{gamma'} with gamma, gamma' >= 0.
    Complexity control of the function class is the engine of the empirical process bounds. It is not proved from the data but assumed from the choice of F.
  • domain assumption Assumption 3.14: phi0 is (beta,b)-smooth in C^2(R^d).
    Replaces (alpha,a)-convexity for the sieve estimator results (Theorems 3.16, 3.17, 3.20). Still excludes potentials with Hessian growth faster than polynomial.
  • domain assumption Assumption 4.1: support Omega is closed, connected and locally connected; if Omega != R^d, the interior is a Lipschitz domain; density p is locally bounded on Omega.
    Sufficient conditions for the Poincaré-type inequalities in Proposition 4.2, which drive the faster rates in Theorems 3.8, 3.17, and 3.20.
  • standard math Caffarelli regularity / Villani (2003) Theorem 2.12(iv): gradient of the convex conjugate inverts the OT map, and the conjugate potential is the OT map in the reverse direction.
    Used in Section 3.5 to justify estimating the reverse direction when the forward potential is not smooth enough.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Statistical Convergence Rates of Optimal Transport Map Estimation between General Distributions." pith.science (2026). https://pith.science/paper/Z5THE7DH

@misc{pith2026241208064,
  author       = {Pith},
  title        = {Pith review of: Statistical Convergence Rates of Optimal Transport Map Estimation between General Distributions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Z5THE7DH}},
  note         = {Machine review of arXiv:2412.08064}
}
read the original abstract

This paper studies the convergence rates of optimal transport (OT) map estimators, a topic of growing interest in statistics, machine learning, and various scientific fields. Despite recent advancements, existing results rely on regularity assumptions that are very restrictive in practice and much stricter than those in Brenier's Theorem, including the compactness and convexity of the probability support and the bi-Lipschitz property of the OT maps. We aim to broaden the scope of OT map estimation and fill this gap between theory and practice. Given the strong convexity assumption on Brenier's potential, we first establish the non-asymptotic convergence rates for the original plug-in estimator without requiring restrictive assumptions on probability measures. Additionally, we introduce a sieve plug-in estimator and establish its convergence rates without the strong convexity assumption on Brenier's potential, enabling the widely used cases such as the rank functions of normal or t-distributions. We also establish new Poincar\'e-type inequalities, which are proved given sufficient conditions on the local boundedness of the probability density and mild topological conditions of the support, and these new inequalities enable us to achieve faster convergence rates for the Donsker function class. Moreover, we develop scalable algorithms to efficiently solve the OT map estimation using neural networks and present numerical experiments to demonstrate the effectiveness and robustness.

Figures

Figures reproduced from arXiv: 2412.08064 by the authors.

Figure 1
Figure 1. Visualization of univariate estimated OT maps in various settings for sample size [PITH_FULL_IMAGE:figures/full_fig_p037_1.png] view at source ↗
Figure 2
Figure 2. Visualization of univariate estimated OT maps in various settings for sample size [PITH_FULL_IMAGE:figures/full_fig_p038_2.png] view at source ↗
Figure 3
Figure 3. Visualization of univariate estimated OT maps in various settings for sample size [PITH_FULL_IMAGE:figures/full_fig_p039_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Visualization of univariate estimated OT maps in various settings for sample size [PITH_FULL_IMAGE:figures/full_fig_p040_4.png]
Figure 5
Figure 5. Figure 5: Visualization of 2-dimensional estimated OT maps in various settings for sample [PITH_FULL_IMAGE:figures/full_fig_p041_5.png]
Figure 6
Figure 6. Figure 6: Visualization of 2-dimensional estimated OT maps in various settings for sample [PITH_FULL_IMAGE:figures/full_fig_p042_6.png]
Figure 7
Figure 7. Figure 7: Visualization of 2-dimensional estimated OT maps in various settings for sample [PITH_FULL_IMAGE:figures/full_fig_p043_7.png]
Figure 8
Figure 8. Figure 8: Visualization of 2-dimensional estimated OT maps in various settings for sample [PITH_FULL_IMAGE:figures/full_fig_p044_8.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Statistical Inference for Optimal Transport Maps: Recent Advances and Perspectives

    math.ST 2025-06 conditional

    A survey of minimax rates and limit laws for estimating optimal transport maps from samples, covering smooth, Gaussian, semi-discrete, entropic, and divergence-regularized settings.

Reference graph

Works this paper leans on

14 extracted references · 14 canonical work pages · cited by 1 Pith paper

  1. [1]

    d-dimensional standard normal distribution

  2. [2]

    In these two cases, the coordinates of the random vectors are independently distributed according to either the standard normal distribution or the t(6) distribution

    d-dimensional t-distribution with 6 degree-of-freedom. In these two cases, the coordinates of the random vectors are independently distributed according to either the standard normal distribution or the t(6) distribution. In the univariate case, we have considered three types of OT maps to estimate: for z ∈ R,

  3. [3]

    With some abuse of notation, denote ∇φ0(z) = F(z)

    Rank function: The OT is the CDF of either standard normal or t(6) distribution, i.e., the target measure is Q = U (0, 1). With some abuse of notation, denote ∇φ0(z) = F(z). 32 Algorithm 5 Computing the sieved convex conjugate with projected gradient descent Require: Function φ; value y ∈ Rd; number of epochs T ; projection radius Mn. 1: Initialize x ← 0 ...

  4. [4]

    Linear transformation: ∇φ0(z) = 3z + 5

  5. [5]

    For the multivariate case, where the OT maps are functions between two Rd spaces, we define the OT maps to be the composition of these three cases: for z = (z1, · · ·, zd)⊤ ∈ Rd,

    Signed-quadratic transformation: ∇φ0(z) = sign(z) · z2. For the multivariate case, where the OT maps are functions between two Rd spaces, we define the OT maps to be the composition of these three cases: for z = (z1, · · ·, zd)⊤ ∈ Rd,

  6. [6]

    F(zd)  

    Rank function: ∇φ0(z) =   F(z1) F(z2) ... F(zd)  

  7. [7]

    3zd + 5  

    Linear function: ∇φ0(z) =   3z1 + 5 3z2 + 5 ... 3zd + 5  

  8. [8]

    sign(zd) · z2 d  

    Signed-quadratic function: ∇φ0(z) =   sign(z1) · z2 1 sign(z2) · z2 2 ... sign(zd) · z2 d   . We refer readers to the main paper for the architecture of the ICNN we adopted and the optimization settings. The independent empirical measures Pn and QN are randomly sampled for each experiment, with sample sizes set to be n = N = 100, 300, 500 and 10...

Show all 14 references
  1. [9]

    Effect of Sample Size: Increasing the sample size consistently decreases theL2(P ) losses and enhances the robustness of the OT map estimators, as evidenced by the reduction in standard deviation

  2. [10]

    curse of dimensionality

    Impact of Dimensionality: As the dimension d increases, the losses tend to increase as well, which is expected due to the “curse of dimensionality”

  3. [11]

    This discrepancy is particularly pronounced when estimating the signed-quadratic OT map, which has a larger b parameter in terms 35 of (β, b)-smoothness

    Comparison of Distributions: The L2(P ) losses are generally larger for the t(6) dis- tribution than for the normal distribution. This discrepancy is particularly pronounced when estimating the signed-quadratic OT map, which has a larger b parameter in terms 35 of (β, b)-smoot...

  4. [12]

    Comparison of Estimators: When comparing the two estimators, ∇ ˜φn,N and ∇ ˆφn,N , it is evident that the new estimator ∇ ˜φn,N is more efficient and robust. For example, in the univariate case for the signed-quadratic OT map under P = t(6), the mean loss for ∇ ˜φn,N with n = ...

  5. [13]

    As the sample size increases from n = N = 100 to n = N = 1000, the estimated OT maps progressively fit the ground truth more accurately. This improvement is consistent with our numerical results and theoretical conclusions, which show a decrease in L2(P ) losses as sample size...

  6. [14]

    In the outer regions, where P has a smaller probability mass, the OT map estimators tend to perform poorly

    The OT map estimators exhibit good performance in regions where the distribution P has a large probability density. In the outer regions, where P has a smaller probability mass, the OT map estimators tend to perform poorly. Specifically, the estimated OT maps ∇ ˜φn,N (x) and ∇...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.