{"id":"4204db9a-b8e6-48d9-829f-241df8484b5a","arxiv_id":"1908.10324","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Estimating the Wasserstein-1 metric between two beta-Holder smooth densities has minimax error n^{-(beta+1)/(2beta+d)} up to a log log n / log n factor.","lead":"This paper proves a lower bound on how accurately one can estimate the Wasserstein-1 distance between two smooth probability distributions from samples. It shows the distance is not much easier to estimate than the full distributions, so plug-in estimators are nearly optimal.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The lower bound's factorization in Eq. (2.12) requires disjoint supports for the scale-J wavelets; standard smooth wavelet bases do not satisfy this, so the hard-instance likelihood/TV calculation is unsupported as written.","rationale":"The reader's CONDITIONAL verdict is appropriate. The upper bound is a standard plug-in estimate, but the lower bound has a real missing specification: the wavelet family used in the hard instances must have disjoint supports for the likelihood factorization and telescoping to go through, yet standard wavelet bases do not have disjoint supports and no alternative family is provided. The integer-β Besov embedding flagged by the reader is a secondary issue that the same corrected construction would need to address. I do not see circularity or incorrect asymptotics elsewhere; the flaw is a proof gap, not evidence that the theorem is false, so the verdict should remain conditional pending a precise construction and verification.","tokens_in":10861,"tokens_out":25140,"duration_ms":256744,"concrete_test":"Choose a standard smooth wavelet basis (e.g., Daubechies-4) at scale J satisfying (2.11), and pick a point x in the overlap of the supports of two adjacent h_{Jk}. Compare the additive density 1 + n^{-1/2} ∑_k θ_k h_{Jk}(x) with the product in Eq. (2.12). Any difference shows the factorization fails for that basis. To settle the intended construction, give an explicit family of pairwise disjoint C^r compactly supported zero-mean functions h_{Jk} with ‖h_{Jk}‖_2 = 1 and |∫ f h_{Jk}| ≍ 2^{-J(1+d/2)} for 1-Lipschitz f; verify that (2.11) bounds the C^β norm and that the product factorization is exact. If no such family exists, the lower-bound calculation must be reworked.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Eq. (2.12) asserts dνθ/dx = ∏_k (1 + θ_k n^{-1/2} h_{Jk}(x)), which is algebraically equivalent to the additive density in Eq. (2.10) only if at every x at most one h_{Jk}(x) is nonzero, i.e., the supports of the h_{Jk} are pairwise disjoint. This identity is then used in Eqs. (2.13)-(2.29) to factor the likelihood over k and to apply the telescoping Proposition 2.3. For standard compactly supported orthonormal wavelet bases (Daubechies, symmlets, coiflets, tensor-product constructions), wavelets at a fixed scale have overlapping supports; exact disjointness at a dyadic scale forces the Haar wavelet, which is not smooth enough for the C^β construction when β>0. The paper never specifies a single-scale basis with disjoint C^r elements and the required coefficient decay, so as written the hard-instance construction is incomplete and the TV bound does not follow. This is the central load-bearing step: without the product factorization, the composite-hypothesis likelihood and the n^{-c} indistinguishability bound collapse. The gap is likely repairable using a custom family of smooth zero-mean functions supported on disjoint dyadic cubes, but that must be stated and checked.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the minimax rate of estimating the Wasserstein-1 distance W(μ,ν) between two unknown probability measures on [0,1]^d whose densities belong to a Hölder class Gβ. Theorem 1 claims that the minimax risk is n^{-(β+1)/(2β+d)} up to a multiplicative factor of order loglog(n)/log(n), so that plug-in estimation of the metric is essentially as hard as estimating the measures themselves under W1. The upper bound is obtained by a wavelet-thresholded plug-in estimator and is derived in Section 2.2. The lower bound is constructed in Section 2.1 through composite hypotheses: perturbations νθ = μ + n^{-1/2} Σ_k θ_k h_{Jk}, priors on θ with matching moments up to order 2K, a direct total-variation bound on the induced likelihoods, and Le Cam's argument.","tokens_in":11167,"tokens_out":12949,"duration_ms":127070,"significance":"If the theorem is correct, it resolves a natural question in nonparametric functional estimation: for the Wasserstein-1 metric, estimating the value of the functional is not significantly easier than estimating the underlying distributions, in contrast to known results for quadratic functionals. The proof strategy is conceptually attractive, and the upper-bound part is a useful self-contained derivation of the plug-in rate. The lower-bound framework using moment-matched priors and a telescoping TV bound is elegant and transparent. However, the lower-bound construction as written contains a load-bearing gap concerning the supports of the wavelets used in Eq. (2.12), and the treatment of integer smoothness is not fully justified. The result is plausible and likely repairable, but the proof is currently incomplete.","major_comments":[{"comment":"The factorization dνθ/dx = ∏_k (1 + θ_k n^{-1/2} h_{Jk}(x)) is asserted 'due to the separation of support for wavelets,' but standard compactly supported orthonormal wavelet bases at a fixed scale do not have pairwise disjoint supports; the Haar basis is disjoint but is not smooth enough for the C^β construction when β>0. The manuscript does not specify a wavelet basis with pairwise disjoint C^r elements and the required coefficient decay. This product identity is used in Eqs. (2.13)-(2.15) to factor the likelihood over k and in Step 4 to replace νθ_{-k} by μ when averaging functions of h_{Jk}(Y_i). Without the product factorization, the likelihood does not split into a product of single-coordinate factors, the telescoping TV bound does not follow, and the entire Le Cam argument collapses. The gap is likely repairable by explicitly constructing smooth zero-mean functions supported on disjoint dyadic cubes, but such a construction is not an orthonormal wavelet basis in the standard sense, so the identification of W with the Besov-type norm d_B in Step 1 would also need to be rederived for the new basis. This is a load-bearing omission rather than a cosmetic issue.","section":"Section 2.1, Eqs. (2.12), (2.26)-(2.29)"},{"comment":"Proposition 2.1 states the Besov-Hölder equivalence B^{β,∞}_∞ = C^β only for β not an integer, while Theorem 1 claims the full range β ≥ 0. In Step 2 the text asserts 'dνθ/dx ∈ B^{β,∞}_1 ⊆ C^β' without proof. For integer β this inclusion is not covered by the stated proposition, and it depends on the nonstandard definition of C^β in Eq. (1.4), where β - ⌊β⌋ = 0. Since the lower bound for integer smoothness is part of the theorem, the embedding and the convention for Hölder regularity at integer exponents must be stated and proved, or the theorem must be restricted to non-integer β. This issue is separate from the disjoint-support gap but also affects the validity of the hard-instance construction for the full parameter range.","section":"Section 2.1, Eqs. (2.3), (2.10)-(2.11)"}],"minor_comments":[{"comment":"The notation for the spaces d_B^{γ,p}_q is inconsistent: the printed text mixes the symbol ∞ printed as 8 with subscripts and superscripts, making the definition of d_B^{γ,∞}_1 versus d_B^{γ,∞}_∞ hard to parse. Please define the Besov norm once with unambiguous indices and use that notation consistently throughout.","section":"Section 2.1, notation"},{"comment":"The reference to 'Tribel' should be 'Triebel'; this typo appears both in the text and in the bibliography entry.","section":"References"},{"comment":"The notation 'τ — 1' is informal; please state explicitly that τ = 1 or that any fixed positive τ works after rescaling.","section":"Section 2.1, Step 2, Eq. (2.7)"},{"comment":"The proof that dνθ/dx ≥ 1 − √(2^{dJ}/n) implicitly uses a sup-norm bound on the functions h_{Jk} that is never stated. Please state the normalization and sup-norm bound for the chosen basis or family of functions.","section":"Section 2.1, Step 2, nonnegativity of νθ"},{"comment":"The expression 'inf_{pTn}' appears to be a typo for 'inf_{\\hat T_n}'; please correct it.","section":"Section 2.1, Step 6, Eq. (2.39)"}],"recommendation":"major_revision","confidential_remarks":"The main gap is a missing construction rather than a disagreement with the consensus; if the author supplies an explicit disjoint-support smooth basis or otherwise justifies the product factorization, and handles the integer-β embedding, the result is likely publishable. I would not reject on the current text, but the proof is incomplete as written."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper asks a good question: when µ and ν are β-Hölder smooth, can we estimate W1(µ,ν) faster than we can estimate µ and ν themselves? Liang answers no, up to a loglog n / log n factor. The upper bound is the usual plug-in rate; the real contribution is the lower bound, built from Lepski-style moment-matching priors and a telescoping TV calculation. The idea is clean and the exposition is honest—no fitted constants, no circularity. The author cites the relevant plug-in literature (Weed–Bach, Liang, Singh et al., Weed–Berthet), and the upper bound is derived in Section 2.2 rather than imported.\n\nThe lower bound has two soft spots. The minor one, which the reader's report caught, is that the Besov–Hölder embedding is only quoted for non-integer β. The theorem covers all β ≥ 0, so the integer case needs an argument. That's probably a small fix. The larger problem is in Eq. (2.12): the density is factored as ∏(1 + θ_k n^{-1/2} h_{Jk}), which is only true if the supports of the h_{Jk} at scale J are pairwise disjoint. Standard smooth wavelet bases don't have disjoint supports—Haar does, but Haar isn't C^β for β>0. The paper never specifies a basis that has both disjointness and smoothness. Without that, the product factorization, the likelihood separation, and the TV bound don't go through. This is the load-bearing step. I think the construction is repairable: take smooth zero-mean bump functions on disjoint dyadic cubes instead of an orthonormal wavelet basis, and check that the Besov/Hölder norm bounds still hold. But that has to be written down.\n\nSo my read is more skeptical than the reader's: not 7/10 soundness, more like 5/10 as written. Not because the theorem looks false—I suspect it's true—but because the proof currently has a hole at its center. I'd still send it to peer review. The question is worth settling, the technique is interesting, and the gap is the kind a good referee can push the author to close.\n\nIf you're working on minimax rates for Wasserstein estimation, this is worth reading and discussing, just with a careful eye on Section 2.1.","headline":"Clever lower bound for an important question, but the proof as written needs an explicit disjoint-support basis; likely fixable and worth refereeing.","tokens_in":11663,"tokens_out":2926,"would_cite":false,"duration_ms":30917,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G05","62G20","62G07","60B10"],"pacs":[],"model":"deepseek-v4-flash","headline":"Estimating the Wasserstein-1 distance between smooth distributions is essentially as hard as estimating the distributions themselves.","keywords":["Wasserstein-1 metric","minimax estimation","Hölder smoothness","Besov spaces","wavelet thresholding","plug-in estimator","optimal transport","empirical measures"],"falsifier":"Take $\\beta=1$, $d=2$, fix $J$ and coefficients $\\theta_k$ at the paper's maximal allowed size, and compute the Lipschitz constant of $\\nu_\\theta=1+n^{-1/2}\\sum_k \\theta_k h_{Jk}$ across the boundary between two adjacent wavelet supports. If the constant exceeds the $C^1$ bound $M$, the constructed measures leave $G^\\beta$ and the lower-bound family collapses; conversely, a uniform bound over all such $\\theta$ would resolve the integer-smoothness gap.","tokens_in":10678,"feed_emoji":"📏","tokens_out":9580,"duration_ms":87622,"temperature":0.7,"pith_summary":"The paper claims that the minimax risk of estimating the Wasserstein-1 metric $W(\\mu,\\nu)$ between two unknown probability measures on $[0,1]^d$ ($d\\ge 2$) with $\\beta$-Hölder smooth densities is, up to a factor $\\log\\log(\\min(n,m))/\\log(\\min(n,m))$, the same as the risk of estimating the measures themselves under the metric. In concrete terms, no estimator can beat a rate of about $(\\min(n,m))^{-(\\beta+1)/(2\\beta+d)}$, and a plug-in estimator built from smoothed empirical measures achieves that rate. The consequence is that estimating this particular scalar functional is not significantly easier than estimating the underlying density. The result matters because it closes a natural question in nonparametric two-sample testing and optimal transport: for the Wasserstein-1 metric, the plug-in approach is essentially optimal.","feed_headline":"Wasserstein distance is barely easier to estimate than the densities","feed_subtitle":"Plug-in estimators reach the minimax rate up to a loglog/log factor, so faster methods can only win logarithmically.","key_machinery":"The argument runs on four linked pieces: (i) the wavelet multiresolution expansion of densities on $[0,1]^d$, which turns the Wasserstein-1 distance into a weighted $\\ell^1$ sum of wavelet coefficients; (ii) the single-scale perturbation family $\\nu_\\theta=\\mu+n^{-1/2}\\sum_{k}\\theta_k h_{Jk}$ with $2^{dJ}\\asymp n^{1/(2\\beta/d+1)}$, for which the paper verifies $W(\\mu,\\nu_\\theta)=d_B^{1,q}(\\mu,\\nu_\\theta)=(2^{-dJ})^{-(\\beta+1)/d}(2^{-dJ})\\sum_k|\\theta_k|$; (iii) two symmetric priors $q_0,q_1$ on $[-\\tau,\\tau]$ whose moments match through order $2K$ but whose first absolute moments differ by about $K^{-1}$, making the null and alternative Wasserstein values differ by the target rate while the induced sample distributions stay close in total variation; (iv) a telescoping product inequality plus an $\\ell^2$ computation that bounds the total variation by $2^{dJ}\\cdot 2\\tau^{2K}/\\sqrt{(2K)!}$, which forces $K\\asymp \\log n/\\log\\log n$ and produces the $\\log\\log/\\log$ correction.","core_discovery":"Formally, Theorem 1 establishes that for $d\\ge 2$ and $\\beta\\ge 0$, $$\\frac{\\log\\log(\\min(n,m))}{\\log(\\min(n,m))}\\,(\\min(n,m))^{-\\frac{\\$\\beta$+1}{2\\$\\beta$+d}} \\;\\lesssim\\; \\inf_{\\widehat T}\\sup_{\\mu,\\nu\\in G^\\$\\beta$}\\mathbb{E}\\,|\\widehat T - W(\\mu,\\nu)| \\;\\lesssim\\; (\\min(n,m))^{-\\frac{\\$\\beta$+1}{2\\$\\beta$+d}},$$ where $G^\\beta$ is the class of probability measures whose densities lie in the Hölder space $C^\\beta(M)$. The lower bound is the paper's main technical contribution: it constructs two composite hypotheses on the density $\\nu$ (with $\\mu$ fixed as the uniform measure) using single-scale wavelet perturbations $\\nu_\\theta = \\mu + n^{-1/2}\\sum_k \\theta_k h_{Jk}$, chooses priors on $\\theta$ whose moments match through order $2K$, and relies on the equality $W(\\mu,\\nu_\\theta)=d_B^{1,q}(\\mu,\\nu_\\theta)$ holding on these hard instances. The upper bound is achieved by plugging in wavelet-truncated empirical measures. Together these bounds show that the plug-in estimator is minimax optimal up to the $\\log\\log/\\log$ factor.","pith_inferences":["Editorial extension: the $\\log\\log/\\log$ gap appears to be an artifact of the moment-matching construction, which can only match moments up to $K\\asymp \\log n/\\log\\log n$; a different hypothesis pair might close the gap and yield a minimax rate of exactly $(\\min(n,m))^{-(\\beta+1)/(2\\beta+d)}$.","Editorial extension: the same single-scale wavelet construction should transfer to other integral probability metrics whose dual balls are Besov balls, provided the equality between the metric and the Besov norm holds on the hard instances; the paper's argument gives a template rather than a Wasserstein-specific trick.","Editorial extension: if the Besov–Hölder inclusion fails at integer $\\beta$, the lower bound as stated would only cover non-integer smoothness; a direct check for $\\beta=1$ is a small calculation and would settle the gap."],"forward_implications":["Plug-in estimation is essentially optimal: using the smoothed empirical measures in place of the unknown densities attains the minimax rate up to the $\\log\\log/\\log$ factor, so any improved estimator can only remove a logarithmic factor.","Knowing one of the two measures exactly does not break the rate: with $\\mu$ fixed and only $\\nu$ unknown, the minimax rate for estimating the transport cost is governed by the smoothness of the unknown measure.","When the two measures have different smoothness levels $\\beta_1,\\beta_2$, the estimation rate is governed by the smaller smoothness $\\beta=\\min(\\beta_1,\\beta_2)$.","In dimension $d\\ge 2$, raw empirical plug-in attains only $n^{-1/d}$ even for smooth measures, while the paper's smoothed estimator improves to $n^{-(\\beta+1)/(2\\beta+d)}$, so smoothing matters."],"supporting_citations":[{"why":"Shows raw empirical plug-in has rate $n^{-1/d}$ even for smooth measures, motivating the search for better estimators and the upper-bound design.","marker":"Dudley (1969)"},{"why":"Supplies the symmetric priors $q_0,q_1$ with matching moments up to $2K$ and separated absolute first moments, the core of the composite-hypothesis construction.","marker":"Lepski et al. (1999)"},{"why":"Gives the Besov-space and wavelet characterization used to identify the Hölder class and to compute the single-scale norms.","marker":"Tribel (1980)"},{"why":"Provides the wavelet-basis expansion and Besov norm estimates used in both the lower and upper bound proofs.","marker":"Donoho et al. (1996)"},{"why":"One of the density-estimation results the plug-in upper bound is built on.","marker":"Liang (2018)"},{"why":"Establishes nonparametric density estimation under adversarial losses, the framework the upper-bound argument follows.","marker":"Singh et al. (2018)"},{"why":"Provides sharp rates for estimating smooth densities in Wasserstein distance, used for the upper bound.","marker":"Weed and Berthet (2019)"}],"fun_headline_variants":["Wasserstein estimation is no shortcut: plug-in is near-optimal","Plug-in estimators match optimal Wasserstein rate to a loglog factor","Estimating Wasserstein distance: no easier than estimating measures","No free lunch: Wasserstein metric estimation is as hard as density estimation","Wasserstein estimation: plug-in is minimax up to a loglog/log gap"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The lower bound rests on the claim that every perturbed density built from a single scale of wavelet bumps lies in the Hölder class $C^\\beta$ for every $\\beta\\ge 0$, including integer $\\beta$, although the paper's stated Besov–Hölder embedding covers only non-integer $\\beta$.","fun_headline_variants_meta":{"raw":{"variants":["Wasserstein estimation is no shortcut: plug-in is near-optimal","Plug-in estimators match optimal Wasserstein rate to a loglog factor","Estimating Wasserstein distance: no easier than estimating measures","No free lunch: Wasserstein metric estimation is as hard as density estimation","Wasserstein estimation: plug-in is minimax up to a loglog/log gap"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000937,"raw_usage":{"total_tokens":3984,"prompt_tokens":900,"completion_tokens":3084,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":516,"completion_tokens_details":{"reasoning_tokens":2989}},"tokens_in":516,"tokens_out":3084,"duration_ms":21140,"temperature":1.0,"reasoning_tokens":2989,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T10:48:10.348427+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take $\\beta=1$, $d=2$, fix $J$ and coefficients $\\theta_k$ at the paper's maximal allowed size, and compute the Lipschitz constant of $\\nu_\\theta=1+n^{-1/2}\\sum_k \\theta_k h_{Jk}$ across the boundary between two adjacent wavelet supports. If the constant exceeds the $C^1$ bound $M$, the constructed measures leave $G^\\beta$ and the lower-bound family collapses; conversely, a uniform bound over all such $\\theta$ would resolve the integer-smoothness gap.","supporting_citations":[{"cited_title":"The speed of mean glivenko-cantelli convergence","cited_arxiv_id":null,"evidence_quote":"Shows raw empirical plug-in has rate $n^{-1/d}$ even for smooth measures, motivating the search for better estimators and the upper-bound design."},{"cited_title":"On estimation of the l r norm of a regression function","cited_arxiv_id":null,"evidence_quote":"Supplies the symmetric priors $q_0,q_1$ with matching moments up to $2K$ and separated absolute first moments, the core of the composite-hypothesis construction."},{"cited_title":"Theory of interpolation, functional spaces and differential operators","cited_arxiv_id":null,"evidence_quote":"Gives the Besov-space and wavelet characterization used to identify the Hölder class and to compute the single-scale norms."},{"cited_title":"Density estimation by wavelet thresholding","cited_arxiv_id":null,"evidence_quote":"Provides the wavelet-basis expansion and Besov norm estimates used in both the lower and upper bound proofs."},{"cited_title":"Nonparametric density estimation under adversarial losses","cited_arxiv_id":null,"evidence_quote":"Establishes nonparametric density estimation under adversarial losses, the framework the upper-bound argument follows."}],"review_version":1}