{"id":"2d6f0aa3-d1a2-4092-8e48-423dcd2a26eb","arxiv_id":"2412.08064","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper derives convergence rates for two optimal transport map estimators that relax compact-support, convexity and sub-exponential tail assumptions, including a sieve estimator that drops strong convexity.","lead":"This paper proves how quickly optimal transport maps between probability distributions can be estimated from samples, under far weaker assumptions than previous work: data may be unbounded, non-convex, and heavy-tailed. It also introduces a new sieve estimator that handles rank functions of normal and t distributions, and provides faster rates through new Poincaré-type inequalities.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Assumption 3.1's C^2 regularity of the Brenier potential is violated by basic non-convex-support examples, so the advertised generality over non-convex supports is not delivered.","rationale":"The paper's central contribution is to extend OT map estimation rates to general distributions without compact or convex support assumptions. The reader's weakest-assumption analysis identifies Assumption 3.1's C^2 requirement as a key extra premise beyond Brenier's theorem. My stress-test confirms this is the most load-bearing issue: it is not merely a technical smoothness condition that could be relaxed by approximation, because it is violated for a two-interval uniform-to-uniform example that is squarely within the paper's motivating class of non-convex supports. Both the plug-in estimator (Theorem 3.5) and the sieve estimator (Theorem 3.16) inherit the C^2 requirement through Assumptions 3.2 and 3.14, so all main rates are conditional on this smoothness. This does not make the theorems internally inconsistent; they are stated with explicit hypotheses. However, the framing that the results hold 'without requiring restrictive assumptions on probability measures' is misleading: the restriction has merely been moved from the measures to the potential, and it fails for simple non-convex supports. The concrete test of computing the Brenier potential for the two-interval example provides a definitive check of this limitation. Since the mathematical results are conditionally correct and the assumptions are transparent, a conditional verdict remains appropriate; the paper should add a remark acknowledging that C^2 regular potentials exclude many non-convex supports and discuss whether the assumption can be weakened to piecewise smoothness or smoothness on the support. Thus I keep the reader's conditional verdict unchanged.","tokens_in":39821,"tokens_out":18423,"duration_ms":191189,"concrete_test":"Analytically compute the Brenier potential for P = Uniform([-2,-1] union [1,2]) and Q = Uniform([0,1]). Show its gradient (the OT map) has kinks at x=-1 and x=1, so the potential is not twice differentiable there. Then attempt to apply Theorem 3.5 or Theorem 3.16 to this pair; neither theorem's hypotheses hold because Assumptions 3.1 and 3.14 are violated. This settles that the theorems do not cover non-convex supports of this basic type and that the advertised generality is not achieved without an independent smoothness assumption on the potential.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Assumption 3.1 requires the Brenier potential phi0 to be C^2(R^d), not merely convex as guaranteed by Brenier's theorem. This is not a benign technicality: it fails for elementary non-convex supports that the paper motivates. Let P be uniform on [-2,-1] union [1,2] and Q uniform on [0,1]. The monotone rearrangement T(x) equals (x+2)/2 on the left interval, x/2 on the right interval, and is extended as 1/2 on the gap (-1,1). The resulting T is continuous but its derivative jumps at x=-1 and x=1 (left derivative 1/2, right derivative 0), so any convex antiderivative phi0 has a nonexistent second derivative at these support-boundary points. Hence phi0 is not in C^2(R^d), violating Assumption 3.1; the sieve estimator's Assumption 3.14 has the same C^2 requirement. Consequently, Theorems 3.5, 3.8, 3.16, 3.17, and 3.20 are silent for this pair, despite P and Q being uniform measures on very simple sets and the support being merely a disjoint union of two intervals. The claimed relaxation of compact/convex support assumptions is therefore conditional on a potential-smoothness condition that is not implied by the problem and is provably absent in elementary cases. The proven rates may be correct under the stated assumptions, but the 'general distributions' framing substantially overstates the domain of applicability.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies L2(P) convergence rates of plug-in optimal transport map estimators. Under an (α,a)-convexity and (β,a)-smoothness condition on the Brenier potential, it establishes rates of order n^{-1/γ} + N^{-1/γ} for sub-Weibull P and slower polynomial rates for heavy-tailed P (Theorem 3.5). A sieve estimator (Eq. 3.8) is introduced to remove the (α,a)-convexity assumption (Theorems 3.16 and 3.17). New Poincaré-type inequalities (Section 4) and an L1(P) empirical process inequality (Lemma 3.19) extend results to Donsker classes and to distributions with fewer than four moments (Theorem 3.20). Numerical experiments illustrate the sieve estimator's robustness, especially for t-distributions.","tokens_in":40098,"tokens_out":4913,"duration_ms":48128,"significance":"If the theorems are correct, the paper substantially advances OT map estimation by relaxing support convexity and compactness and by accommodating heavy-tailed distributions, matching Divol et al. rates under weaker assumptions. The paper also contributes a new sieve estimator, novel Poincaré-type inequalities, and an L1 maximal inequality that may be of independent interest. The numerical experiments are thorough and include multivariate cases. However, all proofs are relegated to the supplementary material, which was not available for review, and several central assumptions conflict with the paper's stated generality, so the practical scope is narrower than advertised.","major_comments":[{"comment":"Assumption 3.1 requires the Brenier potential φ0 to be in C^2(R^d) with φ0(0)=0. This is not a consequence of Brenier's theorem and fails for elementary non-convex support pairs. For example, let P be uniform on [-2,-1] ∪ [1,2] and Q uniform on [0,1]. The monotone rearrangement is T(x) = (x+2)/2 on [-2,-1], T(x) = 1/2 on (-1,1), and T(x) = x/2 on [1,2]. This T is continuous but its derivative jumps at -1 and 1, so any convex antiderivative has no second derivative at those points. Hence Theorems 3.5, 3.8, 3.16, 3.17, and 3.20 are silent for this pair, despite P and Q being very simple distributions and the paper claiming to avoid convex and compact support assumptions. The authors should either weaken the C^2 condition or explicitly state this regularity restriction and adjust the claims in the abstract, introduction, and Table 1.","section":"Assumption 3.1 and Section 3.2"},{"comment":"Table 1 indicates that for 'Our Results' no assumptions are made on the support or density of P and Q, but Theorems 3.8 and 3.17 require Assumption 3.7 (a Poincaré-type inequality). Proposition 4.2, which supplies sufficient conditions for Assumption 3.7, requires Assumption 4.1, including a connected, locally connected support, a Lipschitz domain for the interior (if not R^d), and a locally bounded density. Thus the support and density assumptions are still present for the faster-rate results, contradicting the table and the introductory claims of no restrictive probability-measure assumptions. This should be clarified by annotating which theorems require Assumption 4.1.","section":"Table 1 and Theorems 3.8, 3.17"},{"comment":"All proofs and the key technical results Lemma 3.15, Proposition 4.2, and Lemma 3.19 are deferred to the supplementary material, which was not included in the manuscript under review. Since these results are load-bearing for Theorems 3.16, 3.8, and 3.20, respectively, and for the claimed Poincaré-type and L1 empirical process contributions, the referee cannot fully verify the correctness of the main claims from the submitted text. The authors should make the supplementary material available or include the proofs for the central lemmas in the main text.","section":"Supplementary material (all proofs)"}],"minor_comments":[{"comment":"The phrase 'minimizing the the transportation cost' contains a duplicated article; it should read 'minimizing the transportation cost.'","section":"Section 2.1"},{"comment":"The condition following the integral reads 'R∞0 sqrt(H(x)/x) dx <∞ is integrable'; the phrase 'is integrable' is redundant and should be removed to avoid confusion.","section":"Lemma 3.19"},{"comment":"The notation E[log N(h, F, L2(Pn))] does not specify the probability space for the expectation; it should be stated that the expectation is taken with respect to the empirical measure Pn under the product measure P⊗n, or a suitable outer expectation should be used.","section":"Assumption 3.3"},{"comment":"The notation 'a ≤log n b' is defined but the subscript 'log n' is not typeset cleanly in the text; consider using 'a ≤_{log n} b' or a separate sentence explaining the dependence on log n.","section":"Notation after Eq. (1.1)"}],"recommendation":"major_revision","confidential_remarks":"The main technical apparatus appears sound in structure, but the paper's broad claims about 'general distributions' without support assumptions are not supported by Assumption 3.1 and the conditions in Proposition 4.2. The explicit counterexample with uniform measures on a disjoint union of intervals is convincing and should be addressed. I recommend major revision, because the theoretical results may be correct under the stated assumptions, but the framing and scope need careful adjustment, and the supplementary proofs are essential for verification."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a real technical step: the sieve plug-in estimator with the truncated convex conjugate in Eq. (3.8), the new Poincaré-type inequalities in Proposition 4.2, and the L1(P) maximal inequality in Lemma 3.19 are all genuine contributions that go beyond existing literature. The rates in Theorems 3.16, 3.17, and 3.20 cover important cases like rank functions of normal and t-distributions, where the usual (α,a)-convexity fails. The proof structure is standard—localization, entropy control, Poincaré-type inequalities—and the oracle-type bounds are transparent about their dependence on approximation error and complexity parameters. This is not a circular argument, and the paper is honest about concurrent work by Divol et al. in Remark 4.4.\n\nThe main soft spot is Assumption 3.1. Requiring the Brenier potential φ0 to be C^2(R^d) is much stronger than what Brenier's theorem guarantees, and it silently excludes simple non-convex support situations that the paper explicitly claims to handle. The two-interval example in the stress-test is correct: uniform on [-2,-1]∪[1,2] to uniform on [0,1] gives a monotone rearrangement whose derivative jumps at the gap boundaries, so no C^2 potential exists. That means the advertised relaxation of compact/convex support assumptions is overstated—it only works when the potential is smooth enough, not for general non-convex supports. This is a framing problem, not necessarily a proof problem: the theorems are conditional on the assumption, and the rates likely hold when the assumption is satisfied. But the paper should either weaken Assumption 3.1 or substantially moderate the claims about non-convex supports.\n\nTwo more minor issues. First, all proofs are in a separate supplement not included in the preprint, so Lemma 3.15, Proposition 4.2, and Lemma 3.19 cannot be checked from the main text. That is normal for this literature, but it means the current version is not fully verifiable. Second, the numerical section claims the sieve estimator consistently outperforms the original, but Table 3 shows the opposite for the normal rank function at n=500 and n=1000. That overclaim should be fixed. No code or data is provided, which is also a minor weakness.\n\nThe paper deserves serious peer review: the new techniques and rates are valuable for statisticians working on OT estimation, and the gaps are addressable rather than fatal. I would send it out with a request for the supplementary proofs and a revised discussion of Assumption 3.1.","headline":"Solid rates for OT map estimation under smoothness assumptions, but the C^2 Brenier potential requirement undermines the 'general distributions' framing.","tokens_in":40675,"tokens_out":1884,"would_cite":true,"duration_ms":20533,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G20","26D10"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proves convergence rates for plug-in optimal transport map estimators under far weaker assumptions than previously known, including heavy-tailed measures and non-convex or unbounded supports.","keywords":["Optimal Transport","Brenier potential","Poincaré inequality","plug-in estimator","sieve estimator","sub-Weibull distributions","heavy-tailed distributions","multivariate ranks"],"falsifier":"Construct a valid transport problem in which the Brenier potential is convex but not twice differentiable (for example, a quantile function that is piecewise linear with a slope change) and check whether the plug-in estimator's squared $L^2(P)$ error still follows the stated $n^{-1/\\gamma}+N^{-1/\\gamma}$ rate. If the rate still holds, the smoothness premise is not load-bearing; if the error stalls or slows, the premise is necessary.","tokens_in":39594,"feed_emoji":"📊","tokens_out":7384,"duration_ms":66294,"temperature":0.7,"pith_summary":"The paper establishes non-asymptotic convergence rates for the plug-in optimal transport map estimator when the source measure is sub-Weibull or only polynomially tailed. Under $\\alpha$-convexity and $\\beta$-smoothness of the Brenier potential, the squared $L^2(P)$ error is $\\tilde n^{-1/\\gamma} + \\tilde N^{-1/\\gamma}$ up to logarithms for sub-Weibull $P$, and slower polynomial rates for heavy-tailed $P$. A new sieve plug-in estimator, which restricts the convex conjugate search to a growing ball, attains comparable rates without the convexity assumption, covering rank functions of normal and $t$-distributions. New Poincaré-type inequalities and a new $L^1(P)$ empirical process inequality extend the rates to heavier tails and to distributions lacking fourth moments. The results reduce the gap between existing theory and Brenier's theorem, which only requires two finite moments.","feed_headline":"Optimal transport map rates reach heavy tails and unbounded supports","feed_subtitle":"New sieve estimator and Poincaré-type inequalities cover rank functions of normal and t-distributions.","key_machinery":"The central object is the Brenier potential $\\varphi_0$, the convex function whose gradient is the optimal transport map, paired with the semi-dual Kantorovich objective $P\\varphi + Q\\varphi^*$. The proof machinery is empirical-process control of the difference between the empirical and population objective, which requires bounding Hessians of candidate potentials and controlling the convex conjugates on the sample measure $Q_N$. The three new components are the sieve convex conjugate, which truncates the supremum in the Legendre transform to $B(0,M_n)$ and thereby acts as an implicit bound on the inverse map's range; Poincaré-type inequalities that control variances of differences of $\\beta$-smooth potentials from local density boundedness and topological conditions on the support; and an $L^1(P)$ maximal inequality for function classes with polynomial envelopes.","core_discovery":"The central claim is that the plug-in estimator $\\nabla \\hat\\varphi_{n,N}$ obtained from the semi-dual objective $P_n \\varphi + Q_N \\varphi^*$ converges in $L^2(P)$ at rates governed by the covering entropy of the function class, the Hessian growth parameters $(\\alpha,a)$ and $(\\beta,a)$, and the tail of $P$. For sub-Weibull $P$ the rate is $\\tilde n^{-1/\\gamma} + \\tilde N^{-1/\\gamma}$ up to logarithmic factors, and for polynomial-tailed $P$ it degrades to a polynomial rate in the available moments. The sieve estimator $\\nabla \\tilde\\varphi_{n,N}$, defined by $\\varphi^{*,(n)}(y)=\\sup_{x\\in B(0,M_n)}\\langle x,y\\rangle-\\varphi(x)$, removes the need for $\\alpha$-convexity and is claimed to be the first to handle OT maps such as the rank functions of normal and $t$-distributions. With the new Poincaré-type inequalities, Donsker function classes achieve faster, nearly parametric rates. Theorem 3.20, based on the new $L^1(P)$ maximal inequality, gives rates when $P$ has only slightly more than two moments.","pith_inferences":["Editorial extension: the sieve conjugate can be read as an implicit regularizer on the inverse map's range, so choosing $M_n$ as a quantile rather than the maximum may offer a bias-variance trade-off worth testing.","Editorial extension: the Poincaré-type inequalities proven from local density and support topology may transfer to other semi-dual estimation problems, such as generative modeling with Wasserstein objectives.","Editorial extension: the $L^1(P)$ maximal inequality may be usable outside optimal transport, for example in robust M-estimation where only $L^1$ integrability is available."],"forward_implications":["OT-based multivariate ranks and quantiles now carry non-asymptotic convergence guarantees for distributions with unbounded or non-convex supports, not just compact convex ones.","Rank functions of normal and $t$-distributions, which fail $\\alpha$-convexity, are covered by the sieve estimator.","The new Poincaré-type inequalities extend faster Donsker-class rates to sub-Weibull and polynomial-tailed measures, where the classical Poincaré inequality would force sub-exponential tails.","The $L^1(P)$ empirical process inequality extends rates to distributions with fewer than four moments, approaching the minimal moment requirement of Brenier's theorem.","A neural-network implementation of both estimators is provided, and simulations show the sieve estimator is more robust for heavy-tailed $t$ distributions."],"supporting_citations":[{"why":"Provides the existence and uniqueness of the convex Brenier potential whose gradient is the OT map, the object being estimated.","marker":"Brenier 1987, 1991"},{"why":"Sets up the plug-in OT map estimation problem and the minimax rates on compact hypercubes that this paper generalizes.","marker":"Hütter & Rigollet 2021"},{"why":"Supplies the general function-space empirical-process framework and the $(\\alpha,a)$-convexity condition whose rates this paper extends to heavier tails.","marker":"Divol et al. 2022"},{"why":"Motivates the sieve plug-in estimator that removes the convexity assumption.","marker":"Shen & Wong 1994"},{"why":"Shows the classical Poincaré inequality implies sub-exponential tails, motivating the new Poincaré-type inequalities.","marker":"Bobkov & Ledoux 1997"},{"why":"Supplies the empirical-process localization and entropy tools used in the proofs.","marker":"Vaart & Wellner 2023"},{"why":"Provides the earlier OT rank/quantile rate results for compact convex supports that the sieve results extend.","marker":"Ghosal & Sen 2022"}],"fun_headline_variants":["Sieve estimator hits OT maps for heavy-tailed distributions","OT map rates relax support and convexity assumptions","Faster OT map estimation without bi-Lipschitz or compact support","Nearly parametric rates via new Poincaré inequalities"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument assumes the true transport potential is smooth enough to have a second derivative everywhere, with its Hessian bounded above and below by polynomial factors, even though the underlying theorem only guarantees a convex potential.","fun_headline_variants_meta":{"raw":{"variants":["Sieve estimator hits OT maps for heavy-tailed distributions","OT map rates relax support and convexity assumptions","Faster OT map estimation without bi-Lipschitz or compact support","Nearly parametric rates via new Poincaré inequalities"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000736,"raw_usage":{"total_tokens":3333,"prompt_tokens":1032,"completion_tokens":2301,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":648,"completion_tokens_details":{"reasoning_tokens":2236}},"tokens_in":648,"tokens_out":2301,"duration_ms":17573,"temperature":1.0,"reasoning_tokens":2236,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T18:15:21.722426+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Construct a valid transport problem in which the Brenier potential is convex but not twice differentiable (for example, a quantile function that is piecewise linear with a slope change) and check whether the plug-in estimator's squared $L^2(P)$ error still follows the stated $n^{-1/\\gamma}+N^{-1/\\gamma}$ rate. If the rate still holds, the smoothness premise is not load-bearing; if the error stalls or slows, the premise is necessary.","supporting_citations":[],"review_version":1}