REVIEW 3 major objections 4 minor 68 references
Tightening Causal Bounds via Covariate-Aware Optimal Transport
T0 review · 3 major / 4 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read A penalty on duplicate covariates converts conditional optimal transport into ordinary optimal transport, producing valid partial-identification bounds that tighten toward the sharp covariate-adjusted limit.
desk verdict Clean interpolation result with a real but incomplete finite-sample guarantee; worth refereeing and probably citing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the mirror relaxation Vip(η): it introduces mirror covariate copies Z(0) and Z(1), requires that the marginals of (Y(0),Z(0)) and (Y(1),Z(1)) match the observed covariate-outcome laws, and minimizes h(Y(0),Y(1)) plus η‖Z(0)−Z(1)‖²₂ over all such couplings. This is an unconditional optimal transport problem on the augmented space (Y×Z)², so off-the-shelf OT solvers apply; the penalty term carries the covariate information, forcing the two covariate copies to align as η grows and thereby recovering the conditional transport constraint Z(0)=Z(1) almost surely in the limit. For quadratic cost functions, the finite-sample analysis is carried by the Brenier potential between the transformed marginal distributions, whose strong convexity and smoothness constant λ enters the rate.
What would settle it
Compute or estimate the curvature λ of the Brenier potential between $P_{A_{12}^\top Y(0),\,\eta Z}$ and $P_{Y(1),Z}$ for a non-Gaussian family as $\eta$ grows and check whether $\lambda=O(\eta)$; a superlinear growth would break the uniformity of the stated bound. Alternatively, simulate the plug-in estimator at fixed $\eta$ across increasing sample sizes $n,m$ and compare the empirical $L^1$ error to $C\lambda\eta\,\gamma_{N,d}$.
Extended reading notes
Core claim
The paper's central result, Proposition 3.6, is that the mirror relaxation Vip(η) interpolates between the two partial-identification boundaries: Vip(0)=Vu and lim_{η→∞} Vip(η)=Vc, with Vu≤Vip(η)≤Vc and monotonic continuity in η. In words, a single family of ordinary optimal transport problems, each minimizing h(Y(0),Y(1)) plus η‖Z(0)−Z(1)‖²₂ over couplings of the two observed covariate-outcome laws, produces a valid lower bound that is never worse than the covariate-free bound and can be driven arbitrarily close to the sharp conditional bound by increasing the penalty. The paper also constructs a plug-in estimator from the empirical covariate-outcome measures and proves consistency; for quadratic cost functions, the expected absolute error is bounded by Cλη·γ_{N,d}, where γ_{N,d} is the empirical Wasserstein convergence rate and λ is the curvature of the associated Brenier potential.
Load-bearing premise
The finite-sample rate theorem assumes the Brenier potential between the transformed covariate-outcome measures is globally strongly convex and smooth, and the paper has no general bound on how its curvature constant λ scales with the penalty η, so if λ grows faster than linearly the advertised rate degrades.
Editorial extensions
If this is right
- For any finite η the mirror-relaxation bound is a valid partial-identification bound and is never wider than the covariate-free OT bound, so covariates can be exploited without estimating conditional marginals.
- Increasing η moves the bound monotonically toward the sharp covariate-adjusted COT bound, so one algorithmic template covers the full range from Vu to Vc.
- The plug-in estimator is consistent and can be computed with standard OT solvers, avoiding the per-covariate computation and nuisance-function estimation required by direct COT approaches.
- For quadratic causal estimands such as Neymanian variance, null-effect tests, and the correlation of potential outcomes, the finite-sample error decays at the empirical Wasserstein rate γ_{N,d} up to the factor λη.
- In the STAR data experiment the method narrows the Neymanian confidence interval and produces a correlation upper bound below one, which the paper reads as evidence of treatment-effect heterogeneity.
Reading between the lines
- The mirror penalty can be viewed as a Lagrangian penalty on the constraint Z(0)=Z(1); a natural extension the paper does not develop is to select η by monitoring the expected mismatch E‖Z(0)−Z(1)‖² under the optimal coupling, treating it as a duality gap.
- The same penalty-on-copies device should apply to other hard OT variants, including the causal or adapted OT relaxation sketched in the appendix, suggesting a general recipe for converting conditional constraints into unconditional OT penalties.
- Because the rate γ_{N,d} worsens with dimension, the practical benefit of covariate adjustment may be offset by estimation error unless covariate selection, which the paper lists as future work, is applied.
- The population interpolation result is proved under compactness and smoothness assumptions; the Gaussian example shows the conclusion can survive without compactness, but a general extension would require explicit tail or moment conditions.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a covariate-aware relaxation of conditional optimal transport (COT) for partial identification of causal estimands. For a penalty parameter η, the method introduces mirror covariates Z(0), Z(1) and solves a standard (unconditional) optimal transport problem with augmented cost h(Y(0),Y(1)) + η∥Z(0)-Z(1)∥². The population value Vip(η) is shown to be a valid lower bound that is no smaller than the covariate-free OT bound Vu and no larger than the sharp COT bound Vc, and to interpolate between Vu and Vc as η ranges from 0 to infinity. A plug-in estimator based on empirical marginals is proposed, implemented through a standard LP/Sinkhorn solver, and analyzed for consistency and convergence rate. Synthetic and real-data experiments (including the STAR dataset) illustrate that the method can substantially tighten the partial identification interval relative to covariate-free OT and to a competing COT-based method.
Significance. If fully established, the population-level interpolation result would be a clean and practically useful bridge: it turns the harder COT problem into a sequence of standard OT problems while preserving covariate adjustment, and it gives a valid, monotone family of lower bounds. The proof of Proposition 3.6 uses compactness and uniqueness arguments that appear correct, and the estimator is simple to implement with off-the-shelf OT packages. The paper also provides reproducible code and a useful set of quadratic-function examples (Neymanian variance, correlation of potential outcomes, null-effect testing). The main weakness is that the finite-sample rate in Theorem 4.3 is not uniform in the penalty η, because the curvature constant λ is allowed to depend on η without a general bound; this blocks the advertised 'asymptotically exact as η→∞' statistical claim. This is a completeness issue in the finite-sample theory rather than an error in the population-level interpolation.
major comments (3)
- [Section 4.2, Theorem 4.3] The finite-sample guarantee is not uniform in the penalty η, and this is load-bearing for the paper's central statistical claim. The bound is E[|Vip,n,m(η)-Vip(η)|] ≤ C λ η γ_{N,d}, where λ appears in the curvature assumption 1/λ I ⪯ ∇²φ̃ ⪯ λ I on the Brenier potential between P_{A12^T Y(0), ηZ} and P_{Y(1),Z}. Since the source measure depends on η, λ is generally a function of η, but the paper states in Section 4.2 that it 'currently lack[s] a general method to bound the curvature λ with respect to η'; Appendix C.7 offers only an informal scaling heuristic and Lemma 4.6 covers only jointly Gaussian marginals. The population gain Vip(η)↑Vc is realized only as η→∞, yet the error bound contains η and any η-dependent λ; if λ grows faster than O(η), the bound diverges exactly in the regime that approaches the sharp COT bound. Moreover, no bias-variance tradeoff or data-driven rule for selecting η is provided. The theorem as stated is a fixed-η statement, but the abstract and Remark 3.7 use the limit η→∞ as a central motivation, so the finite-sample theory is incomplete for that regime.
- [Section 4.2, Lemma 4.5] The appeal to Caffarelli regularity does not close the gap left by Theorem 4.3. Lemma 4.5 asserts the existence of some λ > 0 under C² uniformly convex supports and Hölder densities, but it gives no quantitative control of λ in terms of η, the covariance structure, or the distributions. Consequently, combining Lemma 4.5 with Theorem 4.3 still yields an unspecified constant and does not establish a rate that remains meaningful when η is sent to infinity with the sample size. The paper should either prove a general bound, e.g., λ(η)=O(η) under explicit conditions, or restrict the asymptotic-exactness claim to the population level and present the finite-sample theory for fixed η with a separate treatment of the bias Vc - Vip(η).
- [Section 4.2, Theorem 4.3 and Section 5.2] The experimental section evaluates the estimator by L1 error against the oracle Vc and reports that the error 'decreases quickly for small η and then stabilizes' as η grows. This empirical behavior is consistent with a favorable λ(η), but it does not substitute for the missing general bound. In addition, the experiments always choose η from a fixed grid (0 to 100) rather than by a principled rule, so the finite-sample message in Section 5.2 is stronger than what Theorem 4.3 currently guarantees. I would ask the authors to either provide a practical selection procedure that balances the bias term Vc - Vip(η) against the variance term Cληγ_{N,d}, or to state explicitly that the rate theory covers only fixed η and that the near-exact regime is justified only at the population level.
minor comments (4)
- [Abstract and Remark 3.7] The words 'narrower PI intervals for any value of the penalty parameter' should be 'no wider', because Proposition 3.6 guarantees only Vip(η) ≥ Vu and equality can occur, for example when the covariates are independent of the potential outcomes.
- [Appendix C.2, C.3, C.4, Remark 4.4] There are several typos that should be corrected in a revision: 'Founier' should be 'Fournier' in the C.2 heading, 'uniqueneness' should be 'uniqueness' in the proof of Proposition 3.5, 'certein' should be 'certain' in Remark 4.4, and 'satistified' should be 'satisfied' in the proof of Proposition 3.5.
- [Algorithm 1, Eq. (4)] The constraint notation '1^T π = (1/m)1^T, π1 = (1/n)1' is easy to misread; I suggest writing π ∈ R^{n×m}_+ with row sums 1/m and column sums 1/n, and stating explicitly that the displayed constraint is over matrix couplings.
- [Section 5.3.1, Table 2] The 'relative sample size' row in Table 2 is not defined in the text. The authors should state that the ratio is Vu/Vip,n,m(η) (or the equivalent normalized variance ratio) and explain why it corresponds to the relative sample size required for a given level of statistical power.
Circularity Check
No significant circularity: the interpolation result Vip(0)=Vu, Vip(eta)->Vc is proven from stated marginals and cost, not assumed, and the finite-sample rate is a genuine extension of external OT stability bounds; the unquantified curvature lambda(eta) is a completeness gap, not a circular reduction. Only a minor, non-load-bearing self-citation (Gao et al. 2025) is noted.
-
other
[Sections 2.1 and 5.2 (Gao et al. 2025 baseline citation)]
"More motivating examples can be found in (Ji et al., 2023; Gao et al., 2025), which propose to obtain the partial identification sets by solving the related (multi-marginal) OT problems. ... comparing our proposal with (i) an existing method for estimating Vc in our causal setting, that is, (Ji et al., 2023) ... and (ii) the unconditional OT method (Gao et al., 2025) estimating Vu."
Gao et al. (2025) shares the present second author (Zijun Gao). It is cited as motivation and as an experimental baseline ('the unconditional OT method (Gao et al., 2025) estimating Vu'), not as an input to the derivation. Vip(eta), Vu, and Vc are each defined directly from the stated marginals P_{Y(0)}, P_{Y(1)}, P_{Y(0),Z}, P_{Y(1),Z} and the cost h (Definitions in Eqs. 1, 2, 3.4); Proposition 3.6 is proven in Appendix C.4 by the set inclusions Pi_ip subset of Pi'_u and Pi'_c subset of Pi_ip combined with a Prokhorov compactness argument showing E||Z(0)-Z(1)||^2 -> 0, so any limit point lies in Pi'_c. No claim in the chain reduces to this self-citation, so it is recorded here only to justify the score of 1 rather than 0 under the rubric; no circular reduction is exhibited.
full rationale
The central claim chain is self-contained and does not reduce to its inputs. Vip(eta) is defined (Definition 3.4) from the same stated marginals and cost as Vu (Eq. 1) and Vc (Eq. 2), with an explicit penalty eta||Z(0)-Z(1)||^2; Vc is not defined in terms of Vip. Proposition 3.6(ii) is proven in Appendix C.4, not assumed: the paper shows Pi_ip subset of Pi'_u (giving Vu <= Vip(0)) plus a gluing construction (giving Vip(0) <= Vu), and for the limit shows eta*E||Z(0)-Z(1)||^2 <= Vc forces E||Z(0)-Z(1)||^2 -> 0, so every weak limit point lies in Pi'_c and the sandwich Vc <= E[h] <= lim Vip(eta) <= Vc closes. The Gaussian example (Example 3.8) independently confirms the 1/eta gap. Theorem 4.3's rate is internally consistent with the paper's own stability bounds (Props. C.4-C.6 extending Manole et al. 2024): Term C <= lambda * Delta and Delta <= 2lambda * E[W2^2] with E[W2^2] of order eta^2 gamma^2 yields the stated lambda*eta*gamma. The estimator is a plug-in empirical OT value with no parameter fitted to reproduce Vc, so nothing is a fitted input renamed as a prediction. The genuine weakness is a completeness gap, not circularity: Section 4.2 concedes 'we currently lack a general method to bound the curvature lambda with respect to eta', and Appendix C.7 offers only an informal scaling argument, so the advertised finite-sample guarantee is not uniform in the penalty that drives Vip toward Vc; this should be weighed as a correctness risk, not as circularity. The only self-citation is Gao et al. (2025), co-authored by the present second author, used as a baseline and motivation rather than as load-bearing support; score 1 reflects that minor self-citation only.
Assumptions & free parameters
free parameters (1)
- eta (penalty parameter) =
user-chosen; eta=50 in STAR experiment
assumptions (6)
- standard math Brenier's theorem on uniqueness of optimal transport maps for quadratic costs
- standard math Fournier-Guillin empirical Wasserstein convergence rates
- domain assumption Completely randomized design / unconfoundedness (Assumption 3.1)
- domain assumption Compact support and densities (Assumption 3.2)
- domain assumption Cost function smooth with injective gradient (Assumption 3.3)
- domain assumption Curvature bound on the Brenier potential (Theorem 4.3 condition)
invented entities (1)
-
Mirror covariates Z(0), Z(1)
Cite this review
Pith. "Pith review of Tightening Causal Bounds via Covariate-Aware Optimal Transport." pith.science (2026). https://pith.science/paper/RD7BMZBA
@misc{pith2026250201164,
author = {Pith},
title = {Pith review of: Tightening Causal Bounds via Covariate-Aware Optimal Transport},
year = {2026},
howpublished = {\url{https://pith.science/paper/RD7BMZBA}},
note = {Machine review of arXiv:2502.01164}
}
read the original abstract
Causal estimands can vary significantly depending on the relationship between outcomes in treatment and control groups, potentially leading to wide partial identification (PI) intervals that impede decision making. Incorporating covariates can substantially tighten these bounds, but requires determining the range of PI over probability models consistent with the joint distributions of observed covariates and outcomes in treatment and control groups. This problem is known to be equivalent to a conditional optimal transport (COT) optimization task, which is more challenging than standard optimal transport (OT) due to the additional conditioning constraints. In this work, we study a tight relaxation of COT that effectively reduces it to standard OT, leveraging its well-established computational and theoretical foundations. Our relaxation incorporates covariate information and ensures narrower PI intervals for any value of the penalty parameter, while becoming asymptotically exact as a penalty increases to infinity. This approach preserves the benefits of covariate adjustment in PI and results in a data-driven estimator for the PI set that is easy to implement using existing OT packages. We analyze the convergence rate of our estimator and demonstrate the effectiveness of our approach through extensive simulations, highlighting its practical use and superior performance compared to existing methods.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...
-
[2]
Agueh, M. and Carlier, G. Barycenters in the W asserstein space. SIAM Journal on Mathematical Analysis, 43 0 (2): 0 904--924, 2011
work page 2011
-
[3]
Incentives and services for college achievement: Evidence from a randomized trial
Angrist, J., Lang, D., and Oreopoulos, P. Incentives and services for college achievement: Evidence from a randomized trial. American Economic Journal: Applied Economics, 1 0 (1): 0 136--163, 2009
work page 2009
-
[4]
Wasserstein generative adversarial networks
Arjovsky, M., Chintala, S., and Bottou, L. Wasserstein generative adversarial networks. In International Conference on Machine Learning, pp.\ 214--223. PMLR, 2017
work page 2017
-
[5]
Aronow, P. M., Green, D. P., and Lee, D. K. Sharp bounds on the variance in randomized experiments. Annals of Statistics, 42 0 (3): 0 850--871, 2014
work page 2014
-
[6]
Conservative Inference for Counterfactuals
Balakrishnan, S., Kennedy, E., and Wasserman, L. Conservative inference for counterfactuals. arXiv preprint arXiv:2310.12757, 2023
work page Pith review arXiv 2023
-
[7]
On the representation and learning of monotone triangular transport maps
Baptista, R., Marzouk, Y., and Zahm, O. On the representation and learning of monotone triangular transport maps. Foundations of Computational Mathematics, 24 0 (6): 0 2063--2108, 2024 a
work page 2024
-
[8]
Baptista, R., Pooladian, A.-A., Brennan, M., Marzouk, Y., and Niles-Weed, J. Conditional simulation via entropic optimal transport: Toward non-parametric estimation of conditional B renier maps. arXiv preprint arXiv:2411.07154, 2024 b
arXiv 2024
Show all 68 references
-
[9]
E., Gerber, M., and Robert, C
Bernton, E., Jacob, P. E., Gerber, M., and Robert, C. P. On parameter estimation with the W asserstein distance. Information and Inference: A Journal of the IMA, 8 0 (4): 0 657--676, 2019
2019
-
[10]
J., Klaassen, C
Bickel, P. J., Klaassen, C. A., Bickel, P. J., Ritov, Y., Klaassen, J., Wellner, J. A., and Ritov, Y. Efficient and A daptive E stimation for S emiparametric M odels , volume 4. Springer, 1993
1993
-
[11]
Convergence of P robability M easures
Billingsley, P. Convergence of P robability M easures . John Wiley & Sons, 2013
2013
-
[12]
Fliptest: fairness testing via optimal transport
Black, E., Yeom, S., and Fredrikson, M. Fliptest: fairness testing via optimal transport. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, pp.\ 111--121, 2020
2020
-
[13]
Optimal transport-based distributionally robust optimization: Structural properties and iterative schemes
Blanchet, J., Murthy, K., and Zhang, F. Optimal transport-based distributionally robust optimization: Structural properties and iterative schemes. Mathematics of Operations Research, 47 0 (2): 0 1500--1529, 2022
2022
-
[14]
Polar factorization and monotone rearrangement of vector-valued functions
Brenier, Y. Polar factorization and monotone rearrangement of vector-valued functions. Communications on P ure and A pplied M athematics , 44 0 (4): 0 375--417, 1991
1991
-
[15]
Caffarelli, L. A. Boundary regularity of maps with convex potentials. Communications on P ure and A pplied M athematics , 45 0 (9): 0 1141--1151, 1992 a
1992
-
[16]
Caffarelli, L. A. The regularity of mappings with a convex potential. Journal of the American Mathematical Society, 5 0 (1): 0 99--104, 1992 b
1992
-
[17]
Caffarelli, L. A. Boundary regularity of maps with convex potentials-- II . Annals of M athematics , 144 0 (3): 0 453--496, 1996
1996
-
[18]
From K nothe's transport to B renier's map and a continuation method for optimal transport
Carlier, G., Galichon, A., and Santambrogio, F. From K nothe's transport to B renier's map and a continuation method for optimal transport. SIAM Journal on Mathematical Analysis, 41 0 (6): 0 2554--2576, 2010
2010
-
[19]
Optimal transport for counterfactual estimation: A method for causal inference
Charpentier, A., Flachaire, E., and Gallic, E. Optimal transport for counterfactual estimation: A method for causal inference. In Optimal Transport Statistics for Economics and Related Topics, pp.\ 45--89. Springer, 2023
2023
-
[20]
Conditional W asserstein distances with applications in B ayesian OT flow matching
Chemseddine, J., Hagemann, P., Steidl, G., and Wald, C. Conditional W asserstein distances with applications in B ayesian OT flow matching. arXiv preprint arXiv:2403.18705, 2024
2024 arXiv
-
[21]
Monge-- K antorovich depth, quantiles, ranks and signs
Chernozhukov, V., Galichon, A., Hallin, M., and Henry, M. Monge-- K antorovich depth, quantiles, ranks and signs. Annals of Statistics, 45 0 (1): 0 223--256, 2017
2017
-
[22]
Stargan: Unified generative adversarial networks for multi-domain image-to-image translation
Choi, Y., Choi, M., Kim, M., Ha, J.-W., Kim, S., and Choo, J. Stargan: Unified generative adversarial networks for multi-domain image-to-image translation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp.\ 8789--8797, 2018
2018
-
[23]
Sinkhorn distances: Lightspeed computation of optimal transport
Cuturi, M. Sinkhorn distances: Lightspeed computation of optimal transport. Advances in Neural Information Processing Systems, 26, 2013
2013
-
[24]
Transport-based counterfactual models
De Lara, L., Gonz \'a lez-Sanz, A., Asher, N., Risser, L., and Loubes, J.-M. Transport-based counterfactual models. Journal of Machine Learning Research, 25 0 (136): 0 1--59, 2024
2024
-
[25]
and Landau, B
Dowson, D. and Landau, B. The F r \'e chet distance between multivariate normal distributions. Journal of Multivariate Analysis, 12 0 (3): 0 450--455, 1982
1982
-
[26]
and Pammer, G
Eckstein, S. and Pammer, G. Computational methods for adapted optimal transport. Annals of Applied Probability, 34 0 (1A): 0 675--713, 2024
2024
-
[27]
and Park, S
Fan, Y. and Park, S. S. Sharp bounds on the distribution of treatment effects and their statistical inference. Econometric Theory, 26 0 (3): 0 931--951, 2010
2010
-
[28]
Partial identification and inference in moment models with incomplete data
Fan, Y., Shi, X., and Tao, J. Partial identification and inference in moment models with incomplete data. Journal of Econometrics, 235 0 (2): 0 418--443, 2023
2023
-
[29]
Z., Boisbunon, A., Chambon, S., Chapel, L., Corenflos, A., Fatras, K., Fournier, N., et al
Flamary, R., Courty, N., Gramfort, A., Alaya, M. Z., Boisbunon, A., Chambon, S., Chapel, L., Corenflos, A., Fatras, K., Fournier, N., et al. POT : Python optimal transport. Journal of Machine Learning Research, 22 0 (78): 0 1--8, 2021
2021
-
[30]
and Guillin, A
Fournier, N. and Guillin, A. On the rate of convergence in W asserstein distance of the empirical measure. Probability Theory and Related Fields, 162 0 (3): 0 707--738, 2015
2015
-
[31]
Optimal T ransport M ethods in E conomics
Galichon, A. Optimal T ransport M ethods in E conomics . Princeton University Press, 2018
2018
-
[32]
Finite-sample guarantees for W asserstein distributionally robust optimization: B reaking the curse of dimensionality
Gao, R. Finite-sample guarantees for W asserstein distributionally robust optimization: B reaking the curse of dimensionality. Operations Research, 71 0 (6): 0 2291--2306, 2023
2023
-
[33]
Bridging multiple worlds: Multi-marginal optimal transport for causal partial-identification problem
Gao, Z., Ge, S., and Qian, J. Bridging multiple worlds: Multi-marginal optimal transport for causal partial-identification problem. In Proceedings of the 28th International Conference on Artificial Intelligence and Statistics, 2025
2025
-
[34]
On H \"o lder continuity-in-time of the optimal transport map towards measures along a curve
Gigli, N. On H \"o lder continuity-in-time of the optimal transport map towards measures along a curve. Proceedings of the Edinburgh Mathematical Society, 54 0 (2): 0 401--409, 2011
2011
-
[35]
Gulrajani, I., Ahmed, F., Arjovsky, M., Dumoulin, V., and Courville, A. C. Improved training of W asserstein GAN s. Advances in Neural Information Processing Systems, 30, 2017
2017
-
[36]
J., Smith, J., and Clements, N
Heckman, J. J., Smith, J., and Clements, N. Making the most out of programme evaluations and social experiments: Accounting for heterogeneity in programme impacts. The Review of Economic Studies, 64 0 (4): 0 487--535, 1997
1997
-
[37]
and Lemar\'echal, C
Hiriart-Urruty, J.-B. and Lemar\'echal, C. Fundamentals of Convex Analysis. Springer Science & Business Media, 2004
2004
-
[38]
W., and Taghvaei, A
Hosseini, B., Hsu, A. W., and Taghvaei, A. Conditional optimal transport on function spaces. SIAM/ASA Journal on Uncertainty Quantification, 13 0 (1): 0 304--338, 2025
2025
-
[39]
Unsupervised multi-domain image translation with domain-specific encoders/decoders
Hui, L., Li, X., Chen, J., He, H., and Yang, J. Unsupervised multi-domain image translation with domain-specific encoders/decoders. In 2018 24th International Conference on Pattern Recognition (ICPR), pp.\ 2044--2049. IEEE, 2018
2018
-
[40]
Imbens, G. W. and Rubin, D. B. Causal I nference in S tatistics, S ocial, and B iomedical S ciences . Cambridge university press, 2015
2015
-
[41]
Model-agnostic covariate-assisted inference on partially identified causal effects
Ji, W., Lei, L., and Spector, A. Model-agnostic covariate-assisted inference on partially identified causal effects. arXiv preprint arXiv:2310.08115, 2023
2023 arXiv
-
[42]
Dynamic conditional optimal transport through simulation-free flows
Kerrigan, G., Migliorini, G., and Smyth, P. Dynamic conditional optimal transport through simulation-free flows. Advances in Neural Information Processing Systems, 37: 0 93602--93642, 2024
2024
-
[43]
and Tamer, E
Kline, B. and Tamer, E. Recent developments in partial identification. Annual Review of Economics, 15: 0 125--150, 2023
2023
-
[44]
and Smith, C
Knott, M. and Smith, C. S. On the optimal mapping of distributions. Journal of Optimization Theory and Applications, 43: 0 39--49, 1984
1984
-
[45]
R., Thorpe, M., Slepcev, D., and Rohde, G
Kolouri, S., Park, S. R., Thorpe, M., Slepcev, D., and Rohde, G. K. Optimal mass transport: Signal processing and machine-learning applications. IEEE Signal Processing, 34 0 (4): 0 43--59, 2017
2017
-
[46]
Laan, M. J. and Robins, J. M. Unified M ethods for C ensored L ongitudinal D ata and C ausality . Springer, 2003
2003
-
[47]
A two-step computation of the exact GAN W asserstein distance
Liu, H., Xianfeng, G., and Samaras, D. A two-step computation of the exact GAN W asserstein distance. In International Conference on Machine Learning, pp.\ 3159--3168. PMLR, 2018
2018
-
[48]
Plugin estimation of smooth optimal transport maps
Manole, T., Balakrishnan, S., Niles-Weed, J., and Wasserman, L. Plugin estimation of smooth optimal transport maps. Annals of Statistics, 52 0 (3): 0 966--998, 2024
2024
-
[49]
Manski, C. F. The mixing problem in programme evaluation. The Review of Economic Studies, 64 0 (4): 0 537--553, 1997
1997
-
[50]
K., Biswas, S., and Jagarlapudi, S
Manupriya, P., Das, R. K., Biswas, S., and Jagarlapudi, S. N. Consistent optimal transport with empirical conditional measures. In International Conference on Artificial Intelligence and Statistics, pp.\ 3646--3654. PMLR, 2024
2024
-
[51]
and Kuhn, D
Mohajerin Esfahani, P. and Kuhn, D. Data-driven distributionally robust optimization using the W asserstein metric: Performance guarantees and tractable reformulations. Mathematical Programming, 171 0 (1): 0 115--166, 2018
2018
-
[52]
Nelsen, R. B. An I ntroduction to C opulas . Springer, 2006
2006
-
[53]
and Cuturi, M
Peyr \'e , G. and Cuturi, M. Computational optimal transport: With applications to data science. Foundations and Trends in Machine Learning , 11 0 (5-6): 0 355--607, 2019
2019
-
[54]
E., Berar, M., and Courty, N
Rakotomamonjy, A., Flamary, R., Gasso, G., Alaya, M. E., Berar, M., and Courty, N. Optimal transport for conditional domain matching and label shift. Machine Learning, pp.\ 1--20, 2022
2022
-
[55]
On W asserstein two-sample testing and related families of nonparametric tests
Ramdas, A., Garc \' a Trillos, N., and Cuturi, M. On W asserstein two-sample testing and related families of nonparametric tests. Entropy, 19 0 (2): 0 47, 2017
2017
-
[56]
Rubin, D. B. Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of Educational Psychology, 66 0 (5): 0 688, 1974
1974
-
[57]
Si, N., Murthy, K., Blanchet, J., and Nguyen, V. A. Testing group fairness via optimal transport projections. In International Conference on Machine Learning, pp.\ 9649--9659. PMLR, 2021
2021
-
[58]
and Munk, A
Sommerfeld, M. and Munk, A. Inference for empirical W asserstein distances on finite spaces. Journal of the Royal Statistical Society. Series B (Statistical Methodology), 80 0 (1): 0 219--238, 2018
2018
-
[59]
On the application of probability theory to agricultural experiments
Splawa-Neyman, J. On the application of probability theory to agricultural experiments. E ssay on principles. section 9. Statistical Science, pp.\ 465--472, 1923
1923
-
[60]
G., Trigila, G., and Zhao, W
Tabak, E. G., Trigila, G., and Zhao, W. Data driven conditional optimal transport. Machine Learning, 110: 0 3135--3155, 2021
2021
-
[61]
Empirical optimal transport on countable metric spaces: Distributional limits and statistical applications
Tameling, C., Sommerfeld, M., and Munk, A. Empirical optimal transport on countable metric spaces: Distributional limits and statistical applications . Annals of Applied Probability, 29 0 (5): 0 2744 -- 2781, 2019
2019
-
[62]
Taskesen, B., Blanchet, J., Kuhn, D., and Nguyen, V. A. A statistical test for probabilistic fairness. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pp.\ 648--665, 2021
2021
-
[63]
An optimal transport approach to estimating causal effects via nonlinear difference-in-differences
Torous, W., Gunsilius, F., and Rigollet, P. An optimal transport approach to estimating causal effects via nonlinear difference-in-differences. Journal of Causal Inference, 12 0 (1): 0 20230004, 2024
2024
-
[64]
Van der Vaart, A. W. Asymptotic S tatistics , volume 3. Cambridge U niversity P ress, 2000
2000
-
[65]
Villani, C. et al. Optimal T ransport: O ld and N ew , volume 338. Springer, 2009
2009
-
[66]
O., Baptista, R., Marzouk, Y., Ruthotto, L., and Verma, D
Wang, Z. O., Baptista, R., Marzouk, Y., Ruthotto, L., and Verma, D. Efficient neural network approaches for conditional optimal transport with applications in B ayesian inference. arXiv preprint arXiv:2310.16975, 2023
2023 arXiv
-
[67]
K., Munn, M., and Acciaio, B
Xu, T., Wenliang, L. K., Munn, M., and Acciaio, B. COT-GAN : Generating sequential data via causal optimal transport. Advances in Neural Information Processing Systems, 33: 0 8798--8809, 2020
2020
-
[68]
and Watanabe, S
Yamada, T. and Watanabe, S. On the uniqueness of solutions of stochastic differential equations. Journal of Mathematics of Kyoto University, 11 0 (1): 0 155--167, 1971
1971
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.