REVIEW 2 major objections 3 minor 2 cited by
Wasserstein projections in the convex order: regularity and characterization in the quadratic Gaussian case
T0 review · 2 major / 3 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper proves that, for Gaussian measures, the two quadratic Wasserstein projections onto convex-order sets have explicitly characterized Gaussian covariance matrices, obtained by rotating both covariances into a common…
desk verdict The regularity half is solid and self-contained; the Gaussian characterization is elegant but rests on a companion-paper theorem, so the paper is conditionally ready for serious refereeing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing tool is the common-correlation-matrix representation of the Bures-Wasserstein distance. Two matrices $\Sigma_1,\Sigma_2$ share a correlation matrix $C\in C_d$ when $\Sigma_i=\operatorname{dg}(\Sigma_i)^{1/2}C\operatorname{dg}(\Sigma_i)^{1/2}$; the paper states that some orthogonal $O$ always makes $O^*\Sigma_1O$ and $O^*\Sigma_2O$ share a correlation matrix, and in that frame the Bures-Wasserstein distance equals $\sum_i(\sqrt{(O^*\Sigma_1O)_{ii}}-\sqrt{(O^*\Sigma_2O)_{ii}})^2$. Together with the diagonal rescaling matrix $D$, this reduces the projection formulas to one-dimensional comparisons of diagonal entries and supplies the optimal transport maps $x\mapsto ODO^*x$ and their inverses.
What would settle it
Take two non-commuting singular covariance matrices, compute their Bures-Wasserstein distance directly from the formula $\operatorname{tr}(\Sigma_1+\Sigma_2-2(\Sigma_1^{1/2}\Sigma_2\Sigma_1^{1/2})^{1/2})$, and check whether any orthogonal matrix $O$ makes this equal to the sum of squared differences of the square roots of the diagonal entries of $O^*\Sigma_1O$ and $O^*\Sigma_2O$; a pair where no such $O$ exists would falsify Theorem 4.3 and the formulas built on it.
Extended reading notes
Core claim
The main discovery is an explicit characterization of the covariance matrices of the two quadratic Wasserstein projections for Gaussian measures. The paper proves that for any covariance matrices $\Sigma_\mu, \Sigma_\nu\in S_d^+$ there exists an orthogonal matrix $O$ such that, with $D=\operatorname{diag}\left(1\wedge \sqrt{(O^*\Sigma_\nu O)_{ii}/(O^*\Sigma_\mu O)_{ii}},\ldots\right)$, the inequality $DO^*\Sigma_\mu OD\le O^*\Sigma_\nu O$ holds, and for any such $O$, the projected covariances are $I_2(\Sigma_\mu,\Sigma_\nu)=ODO^*\Sigma_\mu ODO^*$ and $J_2(\Sigma_\nu,\Sigma_\mu)=O\tilde\Sigma_J O^*$ with $\tilde\Sigma_J$ given explicitly in terms of $O^*\Sigma_\mu O$, $O^*\Sigma_\nu O$, and $D$. The paper also shows that when $\Sigma_\nu$ is singular and $d\ge 2$, non-Gaussian upper projections can exist, but they all share the same covariance and the Gaussian representative is unique; a necessary and sufficient condition is given for the Gaussian projection to be the only projection at all.
Load-bearing premise
The entire Gaussian characterization rests on the claim that one can always rotate both covariance matrices into a common correlation-matrix frame and that the Bures-Wasserstein distance then decomposes as a sum of one-dimensional squared differences of square roots of diagonal entries.
Editorial extensions
If this is right
- For Gaussian inputs, both projection covariance matrices can now be computed by an orthogonal diagonalization plus an entrywise rescaling, with no iterative optimization needed beyond finding the orthogonal matrix.
- When $\Sigma_\nu$ is singular, the paper settles the previously open uniqueness question: the Gaussian upper projection is always unique, and all non-Gaussian projections, when they exist, have the same covariance matrix and the same lower-dimensional Gaussian marginals.
- The inverse optimal transport maps between the lower and upper projections give concrete convex contraction and expansion maps, linking the Gaussian result to known backward and forward projection structure in convex order.
- The projected-gradient algorithm in the final section computes $J_2$ first and then derives $I_2$ from the explicit formula, giving a practical route for numerical calculation in dimension higher than one.
- Trace identities from the formulas show that the second moments of the two projections sum to the second moments of the original measures, a multidimensional Gaussian analogue of the one-dimensional identity for the projections on the line.
Reading between the lines
- This reader's inference: the explicit orthogonal frame turns the Gaussian projection problem into a diagonal semidefinite program, so the same formulas could accelerate the projected-gradient scheme in Section 4.4 and make high-dimensional Gaussian projections numerically routine.
- This reader's inference: the construction of non-Gaussian projections for singular $\nu$ uses martingale couplings between Gaussian conditionals; a testable extension is whether the same mechanism describes non-Gaussian projections for elliptically contoured or conditionally Gaussian inputs.
- This reader's inference: the fact that the optimal transport maps between the original measures and the two projections are mutual inverses suggests that convex-order projection could be used to build explicit one-sided martingale couplings with exact covariance control in martingale optimal transport problems.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies Wasserstein projections onto sets of probability measures ordered in the convex order. It proves continuity of the projections when they are unique, a 1-Lipschitz (non-expansive) dependence of I_2(μ,ν) on μ, a local Hölder-1/2 dependence on ν, and then analyzes the Gaussian case. For Gaussian laws, it proves that I_2 is Gaussian, studies uniqueness of J_2 in the singular-covariance regime, gives a characterization of when non-Gaussian projections with the same covariance exist, and states an explicit orthogonal-diagonalization formula for the covariance matrices of both projections. The main technical engine for the Gaussian characterization is Theorem 4.3, a common-correlation reduction for the Bures-Wasserstein distance, whose proof is only sketched and partly deferred to a companion paper.
Significance. If fully established, the results are significant. The regularity results in Sections 1–2 are clean and useful, and the Gaussian analysis appears to settle a previously open uniqueness question for J_2 in the singular Gaussian case. The explicit formulas in Theorem 4.1, modulo Theorem 4.3, give a striking reduction to an axis-aligned computation and also support a practical projected-gradient algorithm. The proofs of Propositions 1.1, 2.2, 2.9, and 4.7 are careful and largely self-contained, and the paper includes explicit examples and a concrete numerical procedure. The main weakness is that the central Gaussian characterization depends on a theorem whose proof is not contained in the manuscript.
major comments (2)
- [Section 4.1 (Theorem 4.3)] Theorem 4.3 is load-bearing: Proposition 4.9 invokes it to construct the orthogonal matrix O for positive-definite matrices and then extends by compactness to semidefinite ones; Theorem 4.1 uses (4.10) via Theorem 4.3; Corollary 4.11 and Proposition 3.3 both rely on Theorem 4.1. However, the proof of Theorem 4.3 is only sketched in the positive-definite case via (4.2), and the singular case is explicitly deferred with the sentence "we show (see [5])". The singular case is not a cosmetic edge case: it is exactly the regime needed for the J_2 formula (4.1) and for the new uniqueness results. As it stands, the Gaussian characterization is conditional on an unproved theorem, and the manuscript should either include a complete proof of Theorem 4.3 or clearly restate it with a proof in an appendix.
- [Section 3.3 (Proposition 3.3)] Proposition 3.3, which characterizes when the Gaussian projection is the unique W_2-projection in the singular case, is proved using Corollary 4.11 and Theorem 4.1, and therefore inherits the dependency on Theorem 4.3. Since Theorem 4.3 is not proved in this preprint, the uniqueness characterization is also conditional. This should be fixed together with the gap in Theorem 4.3; alternatively, these results should be moved to or explicitly attributed to the companion paper with a full proof provided.
minor comments (3)
- [Section 2, proof of Proposition 2.9] In the display after equation (2.5), the upper bound is written as "(W_2(μ,I_2(μ,ν)) + W_2(μ,I_2(μ,ν))) W_2(ν,\tilde ν)", but the second term should be W_2(μ,I_2(μ,\tilde ν)) to match the statement (2.4). This is a typo, but it makes the displayed estimate weaker than claimed.
- [Section 4.3 (Corollary 4.14)] In the statement of Corollary 4.14, the assumption reads "(O^*Σ_ν O)_{ii} > 0 and (O^*Σ_ν O)_{ii} > 0 for all i", where the second inequality should refer to (O^*Σ_μ O)_{ii}. Please correct this typo.
- [Notation and Section 2] The symbol W_2 is overloaded: it denotes the usual Wasserstein distance, and after Proposition 2.1 it is also used for the centered version W_2(μ,ν) = (W_2^2(μ,ν) − |m(μ)−m(ν)|^2)^{1/2}. The authors do warn the reader, but the reuse of the same symbol in equations such as (2.4) and its proof is a source of confusion; introducing a different symbol for the centered distance would improve readability.
Circularity Check
No circularity: Theorem 4.1 formulas follow from independent Bures-Wasserstein and projection facts; companion-paper deferral is a proof dependency, not a self-referential reduction.
full rationale
I traced the derivation chain of the main Gaussian characterization, Theorem 4.1. Proposition 4.7 computes I2(Σµ,Σν) and the projection distance from the inequality DO*ΣµOD ≤ O*ΣνO using the derivative formula in Lemma 4.5, convexity of bw2, and a limiting argument for singular matrices; none of these steps presupposes the theorem's conclusion. Proposition 4.9 establishes the existence of the orthogonal matrix O by applying Theorem 4.3 to the pair (Σν, J2(Σν,Σµ)) and then using the first-order optimality of J2 together with Lemma 4.5; it does not assume the formulas for I2 or J2 that are being proved. Theorem 4.1 then derives the J2 formula by checking the distance equality (4.10) through Theorem 4.3 and the already-established identity W2(µ,I2(µ,ν)) = W2(ν,J2(ν,µ)) from [3]. The reliance on Theorem 4.3, whose full proof is deferred to the companion paper [5], is a proof dependency rather than a circular one: Theorem 4.3 is a parameter-free statement about the Bures-Wasserstein distance and common correlation matrices, and its assumptions do not include the projection values or the target formulas. The other load-bearing inputs—the Lipschitz estimate (0.8) from [24], uniqueness from [3], and the Gaussian Wasserstein distance formula from [16,18,26]—are external results or are proved self-contained in Sections 1–2 (Propositions 2.2 and 2.9 are proved directly). No fitted parameter is renamed as a prediction, no ansatz is smuggled in through a citation, and no known result is merely relabeled. The deferred singular-case proof is a completeness or correctness risk, not a circularity, and the manuscript itself flags it by citing the companion paper for that case.
Assumptions & free parameters
assumptions (5)
- domain assumption Uniqueness and existence of the Wasserstein projections I_ρ and J_ρ under the stated conditions (ρ>1 and, for J_ρ, d=1 or ν has a density).
- domain assumption Lipschitz estimate (0.8)-(0.9) for projection distances from [24].
- domain assumption Theorem 4.3: existence of a common correlation matrix representation and the related Bures-Wasserstein formula, with proof deferred to companion paper [5].
- domain assumption Martingale coupling between two centered Gaussians from [23, Remark 1.17] used to construct non-Gaussian projections in Proposition 3.3.
- standard math Standard weak compactness and tightness results from [13] and [28].
Cite this review
Pith. "Pith review of Wasserstein projections in the convex order: regularity and characterization in the quadratic Gaussian case." pith.science (2026). https://pith.science/paper/RL6AJ6PZ
@misc{pith2026250623981,
author = {Pith},
title = {Pith review of: Wasserstein projections in the convex order: regularity and characterization in the quadratic Gaussian case},
year = {2026},
howpublished = {\url{https://pith.science/paper/RL6AJ6PZ}},
note = {Machine review of arXiv:2506.23981}
}
abstract
In this paper, we first show continuity of both Wasserstein projections in the convex order when they are unique. We also check that, in arbitrary dimension $d$, the quadratic Wasserstein projection of a probability measure $\mu$ on the set of probability measures dominated by $\nu$ in the convex order is non-expansive in $\mu$ and H\"older continuous with exponent $1/2$ in $\nu$. When $\mu$ and $\nu$ are Gaussian, we check that this projection is Gaussian and also consider the quadratic Wasserstein projection on the set of probability measures $\nu$ dominating $\mu$ in the convex order. In the case when $d\ge 2$ and $\nu$ is not absolutely continuous with respect to the Lebesgue measure where uniqueness of the latter projection was not known, we check that there is always a unique Gaussian projection and characterize when non Gaussian projections with the same covariance matrix also exist. Still for Gaussian distributions, we characterize the covariance matrices of the two projections. It turns out that there exists an orthogonal transformation of space under which the computations are similar to the easy case when the covariance matrices of $\mu$ and $\nu$ are diagonal.
Forward citations
Cited by 2 Pith papers
-
A Brenier-Strassen Theorem on CAT(kappa) Spaces
On CAT(0) spaces, every probability measure has a unique Wasserstein projection onto the set of measures dominated by ν in convex order, and the optimal transport is a 1-Lipschitz map.
-
On the quadratic barycentric transport problem
The quadratic barycentric transport cost equals the infimum of an expected kinetic-energy integral over semimartingales, and the optimal processes are geodesics with Markovian dynamics.
Reference graph
Works this paper leans on
-
[5]
A. Alfonsi and B. Jourdain. Quadratic Wasserstein distance between Gaussian laws revisited with correlation. Arxiv Preprint, 2025. 30 AUR ´ELIEN ALFONSI AND BENJAMIN JOURDAIN
work page 2025
-
[22]
B. Jourdain, W. Margheriti, and G. Pammer. Lipschitz continuity of the Wasserstein projections in the convex order on the line. Electron. Commun. Probab., 28:Paper No. 18, 13, 2023
work page 2023
-
[3]
A. Alfonsi, J. Corbetta, and B. Jourdain. Sampling of probability measures in the convex order by Wasserstein projection. Ann. Inst. Henri Poincar´e Probab. Stat., 56(3):1706–1729, 2020
work page 2020
-
[1]
A. Alfonsi, J. Corbetta, and B. Jourdain. Sampling of probability measures in the convex order and approximation of Martingale Optimal Transport problems. ArXiv e-prints, Sept. 2017
work page 2017
-
[2]
A. Alfonsi, J. Corbetta, and B. Jourdain. Sampling of one-dimensional probability measures in the convex order and computation of robust option price bounds. Int. J. Theor. Appl. Finance, 22(3):1950002, 41, 2019
work page 2019
-
[4]
A. Alfonsi and B. Jourdain. Squared quadratic Wasserstein distance: optimal couplings and Lions differentiability. ESAIM: Probability and Statistics, 24:703–717, 2020
work page 2020
-
[6]
A. Alfonsi, A. Kebaier, and C. Rey. Maximum likelihood estimation for Wishart processes. Stochastic Process. Appl., 126(11):3243–3282, 2016
work page 2016
-
[7]
J.-J. Alibert, G. Bouchitt´e, and T. Champion. A new class of costs for optimal transport planning. European Journal of Applied Mathematics, 30(6):1229–1263, 2019
work page 2019
Show all 28 references
-
[8]
Ambrosio, N
L. Ambrosio, N. Gigli, and G. Savar ´e. Gradient flows in metric spaces and in the space of probability measures. Lectures in Mathematics ETH Z¨ urich. Birkh¨auser Verlag, Basel, second edition, 2008
2008
-
[9]
Backhoff, M
J. Backhoff, M. Beiglb ¨ock, and G. Pammer. Existence, duality, and cyclical monotonicity for weak transport costs. Calculus of Variations and Partial Differential Equations , 58(6):1–28, 2019
2019
-
[10]
Backhoff, M
J. Backhoff, M. Beiglb ¨ock, and G. Pammer. Weak monotone rearrangement on the line. Elec- tronic Communications in Probability, 25, 2020
2020
-
[11]
Backhoff-Veraguas and G
J. Backhoff-Veraguas and G. Pammer. Applications of weak transport theory. Bernoulli, 28(1):370–394, 2022
2022
-
[12]
Backhoff-Veraguas and G
J. Backhoff-Veraguas and G. Pammer. Stability of martingale optimal transport and weak optimal transport. Ann. Appl. Probab., 32(1):721–752, 2022
2022
-
[13]
Beiglb ¨ock, B
M. Beiglb ¨ock, B. Jourdain, W. Margheriti, and G. Pammer. Approximation of martingale couplings on the line in the adapted weak topology. Probab. Theory Related Fields , 183(1- 2):359–413, 2022
2022
-
[14]
Bhatia, T
R. Bhatia, T. Jain, and Y. Lim. On the Bures-Wasserstein distance between positive definite matrices. Expo. Math., 37(2):165–191, 2019
2019
-
[15]
Carlier, T
G. Carlier, T. Lachand-Robert, and B. Maury. A numerical approach to variational problems subject to convexity constraint. Numer. Math., 88(2):299–318, 2001
2001
-
[16]
D. C. Dowson and B. V. Landau. The Fr´echet distance between multivariate normal distributions. J. Multivariate Anal., 12(3):450–455, 1982
1982
-
[17]
Fathi, N
M. Fathi, N. Gozlan, and M. Prod’homme. A proof of the Caffarelli contraction theorem via entropic regularization. Calculus of Variations and Partial Differential Equations , 59:1–18, 2020
2020
-
[18]
C. R. Givens and R. M. Shortt. A class of Wasserstein metrics for probability distributions. Michigan Math. J., 31(2):231–240, 1984
1984
-
[19]
Gozlan and N
N. Gozlan and N. Juillet. On a mixture of Brenier and Strassen theorems. Proceedings of the London Mathematical Society, 120(3):434–463, 2020
2020
-
[20]
Gozlan, C
N. Gozlan, C. Roberto, P.-M. Samson, Y. Shu, and P. Tetali. Characterization of a class of weak transport-entropy inequalities on the line. Ann. Inst. Henri Poincar´e Probab. Stat., 54(3):1667– 1693, 2018
2018
-
[21]
Gozlan, C
N. Gozlan, C. Roberto, P.-M. Samson, and P. Tetali. Kantorovich duality for general transport costs and applications. J. Funct. Anal., 273(11):3327–3405, 2017
2017
-
[23]
Jourdain and G
B. Jourdain and G. Pag `es. Convex comparison of Gaussian mixtures. J. Multivariate Anal. , 209:Paper No. 105448, 2025
2025
-
[24]
Kim, Y.-H
J. Kim, Y.-H. Kim, Y. Ruan, and A. Warren. Statistical inference of convex order by Wasserstein projection. arXiv:2406.02840, 2024
2024 arXiv
-
[25]
Kim and Y
Y.-H. Kim and Y. Ruan. Backward and forward Wasserstein projections in stochastic order. J. Funct. Anal., 286(2):Paper No. 110201, 63, 2024
2024
-
[26]
Olkin and F
I. Olkin and F. Pukelsheim. The distance between two random vectors with given dispersion matrices. Linear Algebra Appl., 48:257–263, 1982
1982
-
[27]
L. C. G. Rogers and D. Williams. Diffusions, Markov processes and martingales: Volume 2, Itˆo calculus, volume 2. Cambridge university press, 2000
2000
-
[28]
C. Villani. Optimal Transport. Old and New, volume 338 of Grundlehren der mathematischen Wissenschaften. Springer, 2009
2009
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.