Pith. sign in

REVIEW 2 major objections 4 minor 1 cited by

Minimax Optimal Rates for Regression on Manifolds and Distributions

T0 review · 2 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read The minimax rate for distribution regression on manifolds is the sum of three smoothness-geometry terms, and a wavelet-adversarial estimator attains it.

desk verdict Careful, complete minimax theory paper with new rates for covariate-dependent manifold supports; the main soft spot is the ambient-smoothness assumption in Regime 3b, but it is explicit and I see no internal contradiction. read the letter →

arxiv 2506.07504 v1 pith:ION333QM submitted 2025-06-09 math.ST stat.TH

classification math.STstat.TH MSC 62G0562G2062R30
keywords conditionaldistributionestimationregressionmanifoldminimaxrateintegralprobabilitymetricwaveletestimatorgenerativemodelssmoothsubmanifolds
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks how many samples are needed to learn the full conditional distribution of a response $Y$ given a covariate $X$ when both variables live on unknown low-dimensional manifolds, and when the response manifold itself may vary smoothly with $X$. The answer it proves is a minimax rate that is the sum of three terms: $n^{-\alpha_X/(2\alpha_X+d_X)}$ for recovering the covariate dependence, $n^{-(\alpha_Y+\gamma)/(2\alpha_Y+d_Y+(\alpha_Y/\alpha_X)d_X)}$ for nonparametric conditional density estimation, and $n^{-\gamma\beta_Y/d_Y}$ (or $n^{-\gamma/(d_Y/\beta_Y+d_X/\beta_X)}$ when the support depends on $X$) for estimating the response support manifold. Matching lower and upper bounds are established for three regimes, including a new manifold-regression problem as an ingredient. The paper constructs a wavelet-based hybrid estimator that alternates density regression in the ambient space with density regression in a learned latent tangent space, and shows it attains the rate up to logarithmic factors. If the rates are right, they provide a benchmark for conditional generative models and make precise how geometry of the support and smoothness of the conditional law trade off against each other.

What carries the argument

The load-bearing object is the conditional expectation functional $J(f,x)=E_{\mu^*_{Y|X=x}}[f(Y)]$ evaluated over the $\gamma$-Hölder test class, whose supremum defines the integral probability metric $d_\gamma$. The paper estimates $J(f,x)$ by a wavelet multiresolution decomposition of $f$: low-resolution (coarse-scale) wavelet coefficients are estimated by joint least-squares mean regression over the covariate $x$, while high-resolution (fine-scale) coefficients are estimated after mapping the response into a learned $d_Y$-dimensional latent space using local tangent-space projections and conditional decoders learned by least squares. A truncation level $J$ is chosen so that the two estimation errors balance, and an adversarial step converts the estimated functional back into a conditional distribution estimate that is simultaneously near-optimal for all $\gamma>0$.

What would settle it

Take a Regime 3b family where the conditional density is $H^{\alpha_Y,\alpha_X}$-smooth only as a function on the joint support manifold, with no smooth extension to the ambient space; if the Section 5 estimator still attains the claimed rate on this family, the ambient-smoothness assumption is unnecessary, and if it does not, the assumption is load-bearing.

Watch

Extended reading notes

Core claim

The central claim is that distribution regression with manifold structure decomposes additively into three statistical tasks, each with its own irreducible rate. For a conditional distribution supported on a fixed unknown $\beta_Y$-smooth manifold, the minimax rate under the $\gamma$-Hölder integral probability metric is $n^{-\alpha_X/(2\alpha_X+d_X)} + n^{-(\alpha_Y+\gamma)/(2\alpha_Y+d_Y+(\alpha_Y/\alpha_X)d_X)} + n^{-\gamma\beta_Y/d_Y}$; when the support is a $(\beta_Y,\beta_X)$-smooth family of manifolds indexed by $x$, the last term becomes $n^{-\gamma/(d_Y/\beta_Y+d_X/\beta_X)}$. The paper proves the lower bounds by reduction to multiple testing over local bump perturbations, proves matching upper bounds with a wavelet estimator, and identifies the phase transitions: for small $\gamma$ the error is governed by support recovery, for intermediate $\gamma$ by conditional density estimation, and for large $\gamma$ by the conditional mean trend at the classical nonparametric rate.

Load-bearing premise

The load-bearing premise is that, in the covariate-dependent-support case, the conditional density $u^*(y|x)$ is the restriction to the joint support of a function that is jointly smooth over the whole ambient space $\mathbb{R}^{D_Y}\times\mathbb{R}^{D_X}$, not merely smooth along the support manifold; if this ambient extension fails, the proof that the pushed-forward latent density remains smooth breaks down.

Editorial extensions

If this is right

  • When the covariate is discrete or $Y$ is independent of $X$ ($d_X=0$), the rate reduces to the unconditional manifold-distribution rate $n^{-1/2}+n^{-(\alpha_Y+\gamma)/(2\alpha_Y+d_Y)}+n^{-\gamma\beta_Y/d_Y}$, so the conditional result contains the unconditional one as a boundary case.
  • For $\gamma\geq d_Y\alpha_X/(2\alpha_X+d_X)$, the density and support terms are dominated by $n^{-\alpha_X/(2\alpha_X+d_X)}$, the classical rate for estimating an $\alpha_X$-smooth regression function; smoother test functions only see the conditional mean.
  • At $\gamma=1$, the support term in Regime 2 is $n^{-\beta_Y/d_Y}$, matching the known Hausdorff minimax rate for manifold estimation, and in Regime 3 it matches the paper's new manifold-regression rate $n^{-1/(d_Y/\beta_Y+d_X/\beta_X)}$.
  • For $\gamma=1$, the rate agrees (up to logs) with the convergence rate proved for conditional diffusion models on manifolds, which indicates those models are minimax optimal under the expected $1$-Wasserstein metric.
  • The estimator's two-step structure—ambient density regression for coarse scales and latent-space density regression for fine scales—shows that rate-optimal conditional distribution estimation does not require knowing the support manifold in advance.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The additive form of the rate suggests that in practice the three tasks (covariate regression, density estimation, support recovery) can be solved by separate modules without losing minimax optimality; this is an implicit design principle the paper does not state.
  • The ambient-smoothness assumption on $u^*$ in Regime 3b may be stronger than needed; a plausible extension is that only intrinsic smoothness on the joint support $M$ matters, and the paper's compatibility conditions $\beta_X\geq\alpha_X+\alpha_X/\alpha_Y$ and $\alpha_Y\geq\alpha_X$ could be relaxed if the analysis is redone in intrinsic coordinates.
  • Viewing the wavelet level $j$ as analogous to diffusion time ties this estimator to score-based generative models; one could test whether replacing the least-squares wavelet step with a conditional score estimator preserves the rate in the noisy/deconvolution setting the paper leaves open.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper studies minimax rates for distribution regression, i.e., estimation of the conditional distribution μ_{Y|X} under the Hölder integral probability metric d_γ, when the covariate X lies on a low-dimensional set and the response Y lies on a low-dimensional manifold that may itself depend on X. Three regimes are analyzed: Euclidean response spaces (Theorem 1), covariate-independent manifold response spaces (Theorem 2), and covariate-dependent manifold families (Theorems 3 and 4). A separate minimax rate for manifold regression is established in Theorem 3. The upper bounds are achieved by a wavelet-based hybrid estimator that combines ambient-space joint mean regression with latent-space density regression on estimated local charts; convergence rates are stated in Theorems 5 and 6 and Corollaries 1–2. The proofs are given in a lengthy appendix, with lower bounds obtained by Fano-type reductions and upper bounds by oracle inequalities for joint mean regression. The stated rates recover known results as special cases: unconditional distribution estimation on manifolds when d_X=0, classical conditional density estimation when d_Y=D_Y, and manifold estimation when the density is trivial. The central mathematical claims appear internally consistent, although full verification of every appendix lemma is beyond the scope of a single review.

Significance. This is a substantial contribution to the theory of conditional distribution estimation under manifold structure. The paper provides a nearly complete minimax characterization for a wide family of regimes, identifies the role of intrinsic dimensions and smoothness parameters, and matches the lower bounds with explicit estimators up to logarithmic factors. The phase-transition discussion is useful and the recovery of known rates for unconditional manifold estimation and Euclidean density regression gives the results credibility. The proofs are detailed and appear to be carefully assembled, with the estimator construction being genuinely nontrivial, especially the ambient-versus-latent decomposition and the treatment of covariate-dependent supports. The main qualification is that the headline covariate-dependent result, Theorem 4, is proved under an ambient-smoothness extension assumption on the conditional density that is stronger than intrinsic smoothness on the joint support; this is not an internal inconsistency, but it is a real limitation on the scope of the claimed minimax characterization.

major comments (2)
  1. [§4.2, Theorem 4, Remark 2] The upper bound for Regime 3b rests on Item 2 of Regime 3b: the conditional density u*(y|x), defined with respect to the volume measure of M_{Y|x}, is assumed to be the restriction of a function in H^{α_Y,α_X}_L(R^{D_Y}, R^{D_X}) on the full ambient product space. This is load-bearing: Lemma 5 and the proof of Theorem 6 push the density forward to tangent coordinates through the chart Φ_{w0}, and the H^{α_Y,α_X}-smoothness of the pushed-forward density is obtained from the ambient extension of u* together with the chart smoothness. Without the extension, I do not see how the proof establishes the stated rate. The theorem is internally consistent, but the paper's framing in the abstract and introduction suggests a result under intrinsic manifold smoothness alone. Please either state prominently that Regime 3b requires this ambient extension and discuss whether the rate is expected to change under intrinsic-only assumptions, or provide a version of the theorem under chart-adapted intrinsic smoothness. As written, the generality of the headline covariate-dependent result is narrower than the presentation suggests.
  2. [§4.2, Theorem 4 and Remark 2] The compatibility conditions β_Y ≥ 2∨(α_Y+1)∨β_X, β_X ≥ α_X + α_X/α_Y, and α_Y ≥ α_X exclude several natural parameter combinations, including α_Y < α_X and β_X < α_X even when the manifold family is very smooth in x. These conditions are used in the same chart-pushforward argument as the previous comment, so they are not merely technical conveniences. The paper should discuss whether the displayed three-term rate is expected to hold outside this parameter range, and identify which proof step first fails when, for example, α_Y < α_X. At minimum, the limitations should be acknowledged in the main text rather than only in Remark 2.
minor comments (4)
  1. [Appendix B.3, Eq. (22)] The index set for the local polynomial order is printed as "|l| ≤ ⌊ eβX ⌋2 + ⌊αX⌋", which appears to be a typesetting error; the intended expression should be clarified.
  2. [Appendix C.2.1] The definition of em2 mixes notations (\(2\alpha_X + d_X + \frac{\alpha_X}{\alpha_Y} D_Y\) versus \(2\alpha_X + d_X + \frac{\alpha_X}{\alpha_Y} d_Y\)); please check and unify the exponents.
  3. [§6] The admission that the rate-optimal estimator is primarily theoretical and not computationally efficient is helpful, but it should be stated earlier, perhaps at the start of Section 5, so readers are not misled by the phrase "propose a new hybrid estimator" in the abstract.
  4. [References] There are duplicate or near-duplicate entries for the authors' own work (Tang and Yang 2023a/2023b) and for Song et al.; these should be consolidated.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the minimax rates are derived from explicit smoothness and manifold assumptions with self-contained proofs; self-citations are independent published results used as benchmarks or limiting-case reductions.

full rationale

The paper's central claims are minimax upper and lower bounds derived from explicit model classes (Regimes 1, 2, 3a, 3b) and proved in the appendices. The estimator in Section 5 is constructed from wavelet thresholding, local polynomial regression, and explicit manifold-chart estimation; the rates in Theorems 5 and 6 follow from approximation and concentration bounds (Lemmas 9, 12, 14, Theorem 8), not from fitting a parameter to the target quantity and then renaming the fit as a prediction. Self-citations to Tang and Yang (2023a,b), Tang et al. (2024, 2025), and Divol (2022) are used for benchmark rates and for definitions and geometric lemmas that are stated and proved in the supplementary material (e.g., Lemmas 3-6, E.5-E.8); these are independent published results with proofs, not placeholders for the new rates. The most delicate step, the Regime 3b upper bound, relies on the explicit assumption that the conditional density u*(y|x) is the restriction of a jointly H^{alpha_Y,alpha_X}-smooth function on the full ambient space (Regime 3b, Item 2; Remark 2), and the proof (Lemma 5) shows that this assumption is sufficient to keep the pushed-forward density smooth. This assumption is load-bearing but it is an explicit premise, not a conclusion obtained circularly from the theorem being proved. The paper also explicitly flags its own limitations (e.g., the noisy observation case is deferred, and the covariate density lower bound in Regime 3a is conjectured to be relaxable), which further supports that the derivation is not masking a circular step. I therefore find no circularity of any of the seven enumerated kinds.

Assumptions & free parameters 0 free parameters · 7 assumptions · 0 invented entities

The central rates rest on the manifold hypothesis, smoothness definitions, local mass lower bounds, and standard wavelet and Fano tools. The paper proves the auxiliary lemmas in appendices, so these are background assumptions rather than post hoc fudge factors.

assumptions (7)
  • domain assumption Data lie on smooth submanifolds with positive reach and intrinsic dimensions d_X and d_Y (Definitions 3, 4).
    This is the manifold hypothesis the whole paper is built on; it is assumed, not derived.
  • domain assumption Conditional distributions admit densities with respect to volume measures that are alpha_Y-smooth in y and alpha_X-smooth in x (Definitions 1, 12).
    These smoothness classes are the basis of all rates.
  • domain assumption Local mass lower bounds: mu_X(B(x,r)) >= g(r) r^{d_X}/L and mu_{Y|x}(B(y,r)) >= g(r) r^{d_Y}/L for small r (Regime 2/3 conditions).
    These conditions ensure enough samples near every point, needed for worst-case minimax analysis.
  • domain assumption In Regime 3b, the conditional density extends to a jointly H^{alpha_Y,alpha_X}-smooth function on ambient space (Regime 3b definition).
    This ambient extension is used to push densities to latent coordinates (Lemma 5, Remark 2).
  • standard math Standard wavelet and Besov theory: existence of an orthonormal wavelet basis with Holder regularity, locality, and coefficient decay (Lemma 7, from Triebel and Gine-Nickl).
    The estimator and all MSE bounds are expressed in wavelet coefficients.
  • standard math Standard information-theoretic tools: Fano's lemma and Varshamov-Gilbert packing used in lower bound proofs.
    Used in Sections C.2, D.4, D.5.
  • standard math Geometric facts about smooth submanifolds with positive reach (chart inverses, volume growth, projection invertibility) as in Divol and Tang-Yang.
    Background for Definitions 3-4 and Lemma 4.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Minimax Optimal Rates for Regression on Manifolds and Distributions." pith.science (2026). https://pith.science/paper/ION333QM

@misc{pith2026250607504,
  author       = {Pith},
  title        = {Pith review of: Minimax Optimal Rates for Regression on Manifolds and Distributions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ION333QM}},
  note         = {Machine review of arXiv:2506.07504}
}
read the original abstract

Distribution regression seeks to estimate the conditional distribution of a multivariate response given a continuous covariate. This approach offers a more complete characterization of dependence than traditional regression methods. Classical nonparametric techniques often assume that the conditional distribution has a well-defined density, an assumption that fails in many real-world settings. These include cases where data contain discrete elements or lie on complex low-dimensional structures within high-dimensional spaces. In this work, we establish minimax convergence rates for distribution regression under nonparametric assumptions, focusing on scenarios where both covariates and responses lie on low-dimensional manifolds. We derive lower bounds that capture the inherent difficulty of the problem and propose a new hybrid estimator that combines adversarial learning with simultaneous least squares to attain matching upper bounds. Our results reveal how the smoothness of the conditional distribution and the geometry of the underlying manifolds together determine the estimation accuracy.

Figures

Figures reproduced from arXiv: 2506.07504 by the authors.

Figure 1
Figure 1. Diagram for the minimax rate under Regime 2 for fixed [PITH_FULL_IMAGE:figures/full_fig_p015_1.png] view at source ↗
Figure 2
Figure 2. Diagram for the minimax rate under Regime 3 for varying [PITH_FULL_IMAGE:figures/full_fig_p019_2.png] view at source ↗

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Semi-Supervised Conditional Diffusion via Label Augmentation

    stat.ML 2026-07 conditional novelty 5.0 of 10

    Attaching a trivial ∅ label to unlabeled data and running joint denoising score matching provably accelerates TV convergence of conditional generation whenever classes share a baseline component.

Reference graph

Works this paper leans on

51 extracted references · 48 canonical work pages · cited by 1 Pith paper

  1. [1]

    Ifh≤ τ 4 , then there exist some constants(c, C)so that for anyx∈ M, c hd ≤vol M(BM(x, h))≤C hd, wherevol M denotes the volume measure ofM

  2. [2]

    For anyh≤r 0 =τ 1 ∧((τ∧L)/4)andx∈ M,B M(x, h)⊂ϕ x BTxM(0, h) ⊂B M(x,8h/7)

  3. [3]

    Using the partition of unity, one can glue constructions in the local charts to form a global construc- tion on the manifold

    Each pointx∈ Mhas a neighborhood which intersectssupp(ρ λ)for only finitely manyλ∈Λ. Using the partition of unity, one can glue constructions in the local charts to form a global construc- tion on the manifold. Such a global construction usually does not rely on the choice of the partition of unity. Conversely, the partition of unity enables the decomposi...

  4. [4]

    The proof of Lemma 7 is provided in Appendix E.1

    (Index of Wavelet basis) for anyR ′ >0, letΨ d j ={ψ∈ Ψ d j : supp(ψ)∩B Rd(0, R′)̸=∅}, then Ψd j can be written as an index set Ψd j ={ψ jι(·) :ι∈I j ⊂[0,1] d+1}, whereI j isC I 2−j/(R′ +C L)-separated. The proof of Lemma 7 is provided in Appendix E.1. The following lemma presents the wavelet trunca- tion approximation for marginal smooth functions, the p...

  5. [8]

    35 A.2 Smooth submanifold family and smooth conditional distributions Firstly we recall the definition of(β Y , βX )-smooth manifold family defined in Definition 4 of the main text

    IfProj M(z) =xfor somezsatisfyingdist(z,M)< τ, thenz−x∈T xM⊥. 35 A.2 Smooth submanifold family and smooth conditional distributions Firstly we recall the definition of(β Y , βX )-smooth manifold family defined in Definition 4 of the main text. Definition((β Y , βX )-smooth submanifold family).A submanifold family MY|x :x∈ M X is said to belong toM βY βX τ...

  6. [9]

    There exist constants(τ, τ1, L)so that{MY|x }x∈MX ∈M βY βX τ,τ1,L (dY , DY ,M X )

  7. [10]

    (Existence ofx-dependentH βY ,βX -smooth local charts) There exist constants(eτ ,eτ1, eL)so that for anyw 0 = (x 0, y0)∈ M, there exists a neighborhood eUy0 ofy 0 onM Y such that for any x∈B MX (x0,eτ), it holds thatBMY|x (y0,eτ)⊂ eUy0 ∩ MY|x ⊂R DY and there exists a uniformly eL-Lipschitz diffeomorphism eQω0(·, x)that mapseUy0 ∩ MY|x toB RdY (0,eτ1)with ...

  8. [11]

    (Solution manifold withH βY ,βX -smooth defining functions) There exist constants( τ ,τ 1, L)so thatM Y ⊂B RDY (0, L)and for anyω 0 = (x 0, y0)∈ M, there exists a functionF ω0 ∈ HβY ,βX L,DY −dY (BRDY (y0, τ),B MX (x0, τ))so that for anyx∈B MX (x0, τ), it holds thatB MY|x (y0, τ) = {y∈B RDY (y0, τ) :F ω0(y, x) =0}, and for any(x, y)∈B MX (x0, τ)×B RDY (y0...

Show all 51 references
  1. [12]

    For anyx∈B MX (x0, τ),BMY|x (y0, τ)⊂U ω0 ∩ MY|x . 2. The functionV T ω0(· −y0), when restricted to domainU ω0 ∩ MY|x , is a diffeomorphism onto its image, with the inverse function denoted byg ω0,x, defined onB RdY (0, τ1). 3. The functionG ω0 :B RdY (0, τ1)×B MX (x0, τ)→R DY ...

  2. [13]

    Then there existsL′ so that{µ ∗ Y|x }x∈MX ∈C βY ,βX ,αY ,αX τ,τ1,L′ (dY , DY ,M X )

    If for anyω 0 = (x0, y0)∈ M, the push-forward measure[ProjTy0 MY|x 0 (· −y0)]#(µ∗ Y|x |Uω0 ∩MY|x )‡ exists with a density function with respect to the volume measure ofTMY|x 0 y0, denoted asν ω0(·|x), and it satisfies thatν ω0(z,|, x)∈ HαY ,αX L (BTMY|x 0 y0(0, τ1),B MX (x0, τ...

  3. [14]

    If{µ ∗ Y|x }x∈MX ∈C βY ,βX ,αY ,αX τ,τ1,L (dY , DY ,M X ). Then there exists a constantL ′ so that for any ω0 = (x0, y0)∈ Mand any eQω0 that satisfies the conditions specified in Point 2 of Lemma 3, the density of the push forward measure[ eQω0(·, x)]#(µ∗ Y|x | eU ω0 Y|x )with...

  4. [15]

    (Regularity)sup x∈Rd |ψ(l)(x)| ≤CR2j|l|+ d j 2 holds for anyl∈N d 0 with|l| ≤αandψ∈ Ψ d j

  5. [16]

    38 (d) for anyj≥1andx∈R d, {ψ∈ Ψ d j :I ψ ∩B Rd(x,2 −(j−1))̸=∅} ≤C ‡ L

    (Locality) for anyψ∈ Ψ d j , there exists a rectangleI ψ such that (a) for anyl∈N d 0 with|l| ≤α,supp(ψ (l))⊂I ψ and the diameter ofI ψ is smaller thanC L2−j (b)sup x∈Rd P ψ∈Ψd j 1(x∈I ψ)≤C ′ L (c) for anyR≥1, {ψ∈ Ψ d j :I ψ ∩B Rd(0, R)̸=∅} ≤C † LR2 jd. 38 (d) for anyj≥1andx∈R...

  6. [17]

    (Wavelet coefficients of smooth function) for anyα 1 ≤α,r >0andf∈ H α1 r (Rd), it holds for anyψ∈ Ψ d j that the wavelet coefficientf ψ = R Rd f(x)ψ(x) dxis bounded byC W r2− d j 2 −jα1 in absolute value

  7. [19]

    It holds for anyS∈ Sthatsup (x,y)∈M P λ∈Λ S(λ, x)2 +|ψ λ(y)S(λ, x)| ≤C

  8. [20]

    Denoteℓ(x, y, S) =P λ∈Λ ∥S(λ, x)∥2 −2ψ λ(y)T S(λ, x), then for anyS, S′ ∈ S, it holds that Eµ∗ h ℓ(X, Y, S)−ℓ(X, Y, S′) 2i ≤CE µ∗ X h X λ∈Λ S(λ, X)−S ′(λ, X) 2i

  9. [21]

    Define the distanced n asd n(S, S′) = q 1 n Pn i=1(ℓ(Xi, Yi, S)−ℓ(Xi, Yi, S′))2 andN(S, dn, ε) be theε-covering number ofSwith respect tod n, Then, for some termsW n, Tn >1that may depend onn, it holds for any0< ε≤sup S,S′∈S dn(S, S′)that N(S, dn, ε)≤( Tn ε )Wn. 39 Then for an...

  10. [22]

    For anyk∈ bKandγ 1 ∈(0,1], Eµ∗ X Eµ∗ Y|X [∥Y− bG[k]( bQ[k](Y))∥ γ1 ·1(X∈B RDX (xk,2τ 2))1(Y∈B RDY (yk,2τ 2))] ≲    C (logn) 1+γ1 √n dY βY ≤2γ 1, C(logn∧ 1 dY −2γ1βY ) )1+γ1 ·n − γ1 dY βY dY βY >2γ 1

  11. [23]

    For anyk∈ bK, there exists(x ∗ k, y∗ k)∈B M((xk, yk), √ 2τ2)such that bV T [k]P ∗ [k] bV[k] ≳C 1IdY , whereP ∗ [k] is the projection matrix ofT MY y∗ k. 56 Given the assumption thatM Y|x =M Y for anyx∈ M x, and note that if a functionf(y, x) isH βY ,βX -smooth for someβ X >0, ...

  12. [24]

    for anyγ 1 ∈(0,1]andk∈ bK, Eµ∗ X Eµ∗ Y|X [∥Y− bG[k]( bQ[k](Y))∥ γ1 ·1(X∈B RDX (xk,2τ 2))1(Y∈B RDY (yk,2τ 2))] ≲    C (logn) 1+γ1 √n dY βY ≤2γ 1, C(logn∧ 1 dY −2γ1βY ) )1+γ1 ·n − γ1 dY βY dY βY >2γ 1

  13. [25]

    for anyj∈ {0} ∪[J]withJ=⌈ 1 2αY +dY +dX αY αX ·log 2( n logn )⌉, Eµ∗ X h X ψ∈ΨdY j (bvkψ(X)−E µ∗ Y|X [ψ( bQ[k](Y))ρ [k](X, Y)])2 i ≲2 2jαX dY 2αX +dX ( n logn )− 2αX 2αX +dX

  14. [26]

    for anyj∈N,ψ∈Ψ dY j andx∈ MX, Eµ∗ Y|x [ψ( bQ[k](Y))ρ [k](x, Y)] ≲2 − dY j 2 −jαY

  15. [27]

    for anyk∈[K]\ bK,E µ∗[ρ[k](X, Y)]≲ q logn n . So for any1< γ≤ dY αY 2αX +dX , (EB)≲E µ∗[ X k∈[K]\ bK ρ[k](X, Y)] +Eµ∗ h sup f∈H γ 1 (RDY ) X k∈ bK ρ[k](X, Y) f ⊥ J (Y)−f ⊥ J ( bG[k]( bQ[k](Y))) i ≲ r logn n + (logn)·2 −J(γ−1) X k∈ bK Eµ∗ ρ[k](X, Y)∥Y− bG[k]( bQ[k](Y))∥ +n − γ ...

  16. [28]

    For anyk∈ bKandγ 1 ∈(0,1], Eµ∗ X Eµ∗ Y|X [∥Y− bG[k]( bQ[k](Y), X)∥γ1 ·1(X∈B RDX (xk,2τ 2))1(Y∈B RDY (yk,2τ 2))] ≲    (logn) 1+γ1 √n dY βY + dX βX ≤2γ 1, (logn∧ 1 βY (dY /βY +dX /βX −2γ1) )1+γ1 + (logn) γ1)·n − γ1 dX βX + dY βY dY βY + dX βX >2γ 1

  17. [29]

    Moreover, since for anyk∈[K]\ bK, it holds that 1 n X i∈I1 ρ[k](Xi, Yi)≤ 1 n X i∈I1 1(∥(Xi, Yi)−(x k, yk)∥ ≤ √ 2τ2) = 0

    For anyk∈ bK, there exists(x ∗ k, y∗ k)∈B M((xk, yk), √ 2τ2)such that bV T [k]P ∗ [k] bV[k] ≳C 1IdY , whereP ∗ [k] is the projection matrix ofT MY|x ∗ k y∗ k. Moreover, since for anyk∈[K]\ bK, it holds that 1 n X i∈I1 ρ[k](Xi, Yi)≤ 1 n X i∈I1 1(∥(Xi, Yi)−(x k, yk)∥ ≤ √ 2τ2) = ...

  18. [30]

    For eachω∈ {0,1}mdY 1 ×mdX 2 , define¯ω= 1−ωin the element-wise manner

    for anyj, k∈[H 0]withj̸=k, the Hamming distance∥ω (j) −ω (k)∥H betweenω (j) andω (k) satisfies mdY 1 mdX 2 4 ≤ ∥ω(j) −ω (k)∥H ≤ 3mdY 1 mdX 2 4 . For eachω∈ {0,1}mdY 1 ×mdX 2 , define¯ω= 1−ωin the element-wise manner. We may expand the above H0 tensors intoH= 2H 0 ones, ordered...

  19. [31]

    Consequently, we have for anyx∈ M X, 1 c4 f(·, x)∈ Hγ 1 (RD Y )(recall that we only considerγ <1)

    Putting pieces together, we have that for any y, y′ ∈R DY andx∈R DX , there exist constantsc 3, c4 such that |f(y, x)−f(y ′, x)| ≤mγ−γβ Y +βY 1 ef(y 1:dY , x)· h ydY +1 −q(y 1:dY ) −h y′ dY +1 −q(y ′ 1:dY ) + h y′ dY +1 −q(y ′ 1:dY ) ·( ef(y 1:dY , x)− ef(y ′ 1:dY , x)) ≤c 3 ∥...

  20. [32]

    Consider anyγ 1 ∈(0,1]and denoteN(G, d γ1 ∞, ε)as theε-covering number of Gwith respect to thed γ1 ∞ distance, whered γ1 ∞(G1, G2) = sup z∈RdY ,x∈RDX ∥G1(z, x)−G2(z, x)∥γ1

    If there existsG∈ GandV∈O(D Y , dY )such that for any(x, y)∈ M={(x, y) :x∈ MX , y∈ MY|x }withx∈B RDX (x0,2τ 2)andy∈B RDY (y0,2τ 2), it holds that∥y−G(V T (y− y0), x)∥ ≤ε∗. Consider anyγ 1 ∈(0,1]and denoteN(G, d γ1 ∞, ε)as theε-covering number of Gwith respect to thed γ1 ∞ dist...

  21. [33]

    If there exists(x ∗, y∗)∈B M((x0, y0), √ 2τ2), andτ 2 < τ1∧τ 2 . Then letP ∗ be the projection matrix ofT MY|x ∗ y∗, there exist positive constantsc, c1 so that ifE µ∗[∥Y− bG( bV T (Y−y 0), X)∥ · 1(X∈B RDX (x0,2τ 2))1(Y∈B RDY (y0,2τ 2))]≤c, then bV T P ∗ bV T ≥c 1IdY . The pro...

  22. [34]

    Ifx ∗ ∈N x εx j , and∥x ∗ −x ∗ k∥ ≤τ2 + 2εx j , then considering the following local approximation to G∗ [k] andv ∗ [k]: G† [k],x∗(z, x) = jX s=0 X ψ∈ eΨdYs X l∈NDX 0 |l|<βX Z RdY 1 l! G∗(0,l) [k] (t, x∗)(x−x ∗)lψ(t) dt·ψ(z) 85 and v† [k],x∗(z, x) = jX s=0 X ψ∈ eΨdYs X l∈NDX 0...

  23. [35]

    LetN z c2 −j be ac2 −j-covering set ofB RdY (0, τ1), contained withinB RdY (0, τ1), wherecis a small enough positive constant

    Ifx ∗ /∈Nx εx j , orx ∗ ∈N x εx j , but∥x ∗ −x ∗ k∥> τ 2 + 2εx j , we defineG † [k],x∗(z, x)≡0 DY and v† [k],x∗(z, x)≡0. LetN z c2 −j be ac2 −j-covering set ofB RdY (0, τ1), contained withinB RdY (0, τ1), wherecis a small enough positive constant. For anyx ∗ ∈ MX, denote ΨDY j...

  24. [36]

    It holds for anyx∈B MX (x∗,2ε x j )that K∗ X k=1 Z BRdY (0,τ1) 2 j(dY −DY ) 2 ψ∗(G∗ [k](z, x))v∗ [k](z, x) dz − K∗ X k=1 Z BRdY (0,τ1) 2 j(dY −DY ) 2 ψ∗(G† [k],x∗(z, x))v† [k],x∗(z, x) dz+ X s∈NDX 0 0≤s≤⌊ eβX ⌋2+⌊αX ⌋ a∗ ψ∗,x∗,s(x−x ∗)s ≤C 2 (logn)·2 − jdY 2 (εx j )αX . 86

  25. [37]

    The proof of Lemma 19 is provided in Appendix D.14

    Ifψ ∗ ∈Ψ DY j \Ψ DY j (x∗), then it holds for anyx∈B MX (x∗, εx j )andx ′ ∈B Nx εx j (x,2ε x j )that, K∗ X k=1 Z BRdY (0,τ1) 2 j(dY −DY ) 2 ψ∗(G∗ [k](z, x))v∗ [k](z, x) dz= 0 and K∗ X k=1 Z BRdY (0,τ1) 2 j(dY −DY ) 2 ψ∗(G† [k],x′(z, x))v† [k],x′(z, x) dz+ X s∈NDX 0 0≤s≤⌊ eβX ⌋...

  26. [38]

    Hencesup x∈Rd P ψ∈Ψd j 1(x∈I ψ)≤C ′ L

    orϕ [d] k (x)̸= 0(j= 0). Hencesup x∈Rd P ψ∈Ψd j 1(x∈I ψ)≤C ′ L. Moreover, ifI ϕ[d] k ∩BRd(0, R)̸=∅, thenk∈[−C−R, C+R] d; ifI ψ[d] ljk ∩B Rd(0, R)̸=∅, thenk∈[2 j−1(−C−R),2 j−1(C+R)] d, so {ψ∈ Ψ d j :I ψ ∩B Rd(0, R)̸=∅} ≤(2 d −1)(2 j(C+R) + 1) d ≤(2 d −1)(C+ 2) dRd2jd; if Iψ[d] ...

  27. [39]

    Ifj= 0, Z Rd f(x)ψ(x) dx= Z Iψ f(x)ψ(x) dx≤ sZ Iψ ψ2(x) dx Z Iψ f 2(x) dx≤(2C) d 2 r

  28. [40]

    116 For the last statement

    Ifj >0, then we have for anyl∈N d 0 with|l|< α 1, Z Rd xlψ(x) dx= 0 and thus for anyx 0 ∈I ψ, we have Z Rd f(x)ψ(x) dx = Z Rd (f(x)−f(x 0))ψ(x) dx = Z Iψ (f(x)−f(x 0))ψ(x) dx = Z Rd X s∈Nd 0 1≤|s|<α1 f (s)(x0) s! (x−x 0)sψ(x) dx+ Z Iψ f(x)− X s∈Nd 0 1≤|s|<α1 f (s)(x0) s! (x−x ...

  29. [41]

    for any(l 1, l2)∈ Jd1,d2 α1,α2 with |l1| α1 + |l2| α2 + 1 α1 ≥1, | X ψ∈Ψd1 j1 X ϕ∈Ψd2 j2 efψ,ϕψ(l1)(x)ϕ(l2)(y)− X ψ∈Ψd1 j1 X ϕ∈Ψd2 j2 efψ,ϕψ(l1)(x′)ϕ(l2)(y)| ≤L 4∥x−x ′∥α1−|l1|− α1 α2 |l2|

  30. [42]

    for any(l 1, l2)∈ Jd1,d2 α1,α2 with |l1| α1 + |l2| α2 + 1 α2 ≥1, | X ψ∈Ψd1 j1 X ϕ∈Ψd2 j2 efψ,ϕψ(l1)(x)ϕ(l2)(y)− X ψ∈Ψd1 j1 X ϕ∈Ψd2 j2 efψ,ϕψ(l1)(x)ϕ(l2)(y′)| ≤L 4∥y−y ′∥α2−|l2|− α2 α1 |l1|. Then given Claim 1, we can derive that for any(l 1, l2)∈ Jd1,d2 α1,α2 with |l1| α1 + |l...

  31. [43]

    If |l1| βY + |l2| βX + 1 βY ≥1, then for anyz, z ′ ∈B RdY (0, τ 2)andx∈B RDX (x0, τ 2), ∥Fω0 (l1,l2)((z, s(z, x)), x)−Fω0 (l1,l2)((z′, s(z′, x)), x)∥ ≲∥z−z ′∥βY −|l1|− βY βX |l2| +∥s(z, x)−s(z ′, x)∥βY −|l1|− βY βX |l2| ≲∥z−z ′∥βY −|l1|− βY βX |l2| ≲∥z−z ′∥βY −|j1|− βY βX |j2|

  32. [44]

    If |l1| βY + |l2| βX + 1 βY <1, then for anyz, z ′ ∈B RdY (0, τ 2)andx∈B RDX (x0, τ 2), ∥Fω0 (l1,l2)((z, s(z, x)), x)−Fω0 (l1,l2)((z′, s(z′, x)), x)∥ ≲∥z−z ′∥+∥s(z, x)−s(z ′, x)∥ ≲∥z−z ′∥≲∥z−z ′∥βY −|j1|− βY βX |j2|. Therefore, there exists a constantL 3 so that for anyk∈[D Y ...

  33. [45]

    If |l1| βY + |l2| βX + 1 βY ≥1, then for anyz∈B RdY (0, τ 2)andx, x′ ∈B RDX (x0, τ 2), ∥Fω0 (l1,l2)((z, s(z, x)), x)−Fω0 (l1,l2)((z, s(z, x′)), x′)∥ ≲∥s(z, x)−s(z, x′)∥βY −|l1|− βY βX |l2| +∥x−x ′∥βX −|l2|− βX βY |l1| ≲∥x−x ′∥βX −|l2|− βX βY |l1| ≲∥x−x ′∥βX −|j2|− βX βY |j1|. 125

  34. [46]

    If |l1| βY + |l2| βX + 1 βY <1and |l1| βY + |l2| βX + 1 βX ≥1, then for anyz, z ′ ∈B RdY (0, τ 2)andx∈ BRDX (x0, τ 2), ∥Fω0 (l1,l2)((z, s(z, x)), x)−Fω0 (l1,l2)((z, s(z, x′)), x′)∥ ≲∥s(z, x)−s(z, x′)∥+∥x−x ′∥βX −|l2|− βX βY |l1| ≲∥x−x ′∥βX −|l2|− βX βY |l1| ≲∥x−x ′∥βX −|j2|− β...

  35. [47]

    If |l1| βY + |l2| βX + 1 βX <1, then for anyz, z ′ ∈B RdY (0, τ 2)andx∈B RDX (x0, τ 2), ∥Fω0 (l1,l2)((z, s(z, x)), x)−Fω0 (l1,l2)((z, s(z, x′)), x′)∥ ≲∥s(z, x)−s(z, x′)∥+∥x−x ′∥ ≲∥x−x ′∥βX −|j2|− βX βY |j1|. Therefore, there exists a constantL 3 so that for anyk∈[D Y −d Y ], t...

  36. [48]

    • if∥z−z 0∥<1, then |G(ej +j1,j2) [k] (z, x)−G(ej +j1,j2) [k] (z0, x)| ≤ p dY L∥z−z 0∥ ≤ p dY L∥z−z 0∥αY −|j1|− αY αX |j2|

    if |j1|+1 βY + |j2| βX + 1 βY <1, then for anyj∈[d Y ],z, z0 ∈R dz withz̸=z 0 andx∈R DX , • if∥z−z 0∥ ≥1, then |G(ej +j1,j2) [k] (z, x)−G(ej +j1,j2) [k] (z0, x)| ≤2L≤2L∥z−z 0∥αY −|j1|− αY αX |j2|. • if∥z−z 0∥<1, then |G(ej +j1,j2) [k] (z, x)−G(ej +j1,j2) [k] (z0, x)| ≤ p dY L∥...

  37. [49]

    • if∥z−z 0∥<1, then |G(ej +j1,j2) [k] (z, x)−G(ej +j1,j2) [k] (z0, x)| ≤L∥z−z0∥βY −|j1|−1− βY βX |j2| ≤L∥z−z 0∥αY −|j1|− αY αX |j2|

    if |j1|+1 βY + |j2| βX + 1 βY ≥1, then since βY −(|j 1|+ 1)− βY βX |j2| =β Y (1− |j1|+ 1 βY − |j2| βX ) ≥(α Y + 1)(1− |j1|+ 1 βY − |j2| βX ) ≥(α Y + 1)(1− |j1|+ 1 αY + 1 − |j2| αX + αX αY ) =α Y − |j1| −αY αX |j2|, we have for anyj∈[d Y ],z, z0 ∈R dz withz̸=z 0 andx∈R DX , • i...

  38. [50]

    Furthermore, given that for anyk∈[K ∗],G ω∗ k (z, x)∈ HβY ,βX L (BRdY (0, τ1), BMX (x0, τ))and vω∗ k (z,|, x)∈ HαY ,αX L1 (BRdY (0, τ1), BMX (x0, τ)), there existG∗ [k] ∈ HβY ,βX L (RdY ,R DX )andeν ∗ [k] ∈ HαY ,αX L1 (RdY ,R DX )such that for anyz∈B RdY (0, τ1)andx∈B MX (x0, ...

  39. [51]

    Z rn 0 r Wn log 4Tnρ ε2 dε # = C4√n Eµ∗,⊗n

    Therefore it follows that theε-covering number of G ∗ with respect tod g n satisfies that,N( G ∗ , dg n, ε)≤N(G ∗, dg n, ε 2 )· 2ρ ε . Therefore, we can obtain for any0< ε≤r n logN( G ∗ , dg n, ε)≤logN(G ∗, dg n, ε 2 ) + log 2ρ ε = logN(S, dn, ε 2 ) + log 2ρ ε ≤W n log 2Tn ε +...

  40. [2013]

    A crossvalidation method for estimating conditional densities

    124 Jianqing Fan and Tsz Ho Yim. A crossvalidation method for estimating conditional densities. Biometrika, 91(4):819–834, 2004. 4 Herbert Federer. Curvature measures.Transactions of the American Mathematical Society, 93(3):418– 491, 1959. 9 Reinaldo Padilha Franc ¸a, Ana Caro...

  41. [2018]

    Number 37

    15 Yves Meyer.Wavelets and operators: volume 1. Number 37. Cambridge university press, 1992. 20, 37, 115 Mehdi Mirza and Simon Osindero. Conditional generative adversarial nets.arXiv preprint arXiv:1411.1784, 2014. 4, 5 Alfred M¨uller. Integral probability metrics and their ge...

  42. [2019]

    Minimax Optimal Rates for Regression on Manifolds and Distributions

    4 Xingyu Zhou, Yuling Jiao, Jin Liu, and Jian Huang. A deep generative approach to conditional sampling. Journal of the American Statistical Association, pages 1–12, 2022. 5 33 Supplementary Materials to “Minimax Optimal Rates for Regression on Manifolds and Distributions” Not...

  43. [2021]

    13 Ilya Tolstikhin, Olivier Bousquet, Sylvain Gelly, and Bernhard Schoelkopf

    URLhttps://proceedings.neurips.cc/paper_files/paper/2021/file/ cfe8504bda37b575c70ee1a8276f3486-Paper.pdf. 13 Ilya Tolstikhin, Olivier Bousquet, Sylvain Gelly, and Bernhard Schoelkopf. Wasserstein auto-encoders. arXiv preprint arXiv:1711.01558, 2017. 24 Hans Triebel.Bases in f...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.