{"id":"c6fc21f1-32ed-4239-a303-9ab0fe971f2c","arxiv_id":"2506.23981","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"For Gaussian measures, the two Wasserstein projections in the convex order are explicit functions of the covariance matrices, and non-uniqueness occurs only in a characterized singular case.","lead":"This paper proves new regularity properties for Wasserstein projections onto sets of probability measures ordered in the convex order. It also gives explicit formulas for these projections when the measures are Gaussian, making the projections easy to compute and showing when uniqueness can fail.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 4.1 rests on Theorem 4.3 (the common-correlation reduction), whose full proof, especially for singular matrices, is deferred to companion paper [5]; if that theorem fails, the formulas for I2 and J2 have no basis.","rationale":"Read in good faith: Sections 1 and 2 are self-contained, and I find no serious defect in the continuity, non-expansiveness, or Hölder arguments; Proposition 2.2's Hilbert-space argument is valid after the translation reduction, and Proposition 2.9's estimates check out. Sections 3's Gaussian projection reductions are also internally consistent. The only place where the central claim can break is the common-correlation reduction, Theorem 4.3. The paper's own text defers its proof to the companion paper [5] and gives only the positive-definite idea (4.2). The theorem is plausible — for positive-definite matrices one can diagonalize T = A^{−1/2}(A^{1/2}BA^{1/2})^{1/2}A^{−1/2}, which satisfies TAT = B and therefore yields a common correlation matrix — but the singular case is not supplied in this preprint. Proposition 4.9, Theorem 4.1's formulas, Corollary 4.11, and §3.2's uniqueness discussion all import this theorem. Hence the conditional verdict is appropriate. I do not see a reason to move to reject or unverified: the dependency is explicit, the companion paper presumably contains the missing proof, and the rest of the paper is coherent. The abstract's omission of 'locally' before 'Hölder continuous' is a presentation issue, not a load-bearing concern.","tokens_in":36891,"tokens_out":15159,"duration_ms":161560,"concrete_test":"Provide a complete proof of Theorem 4.3 for singular Σ1,Σ2, for instance by showing that the positive-definite construction O^*Σ1^{−1/2}(Σ1^{1/2}Σ2Σ1^{1/2})^{1/2}Σ1^{−1/2}O = Λ passes to the limit and yields a common correlation matrix, and verify the trace formula on the limiting pair. As a computational falsification check, draw 10^4 pairs (Σµ,Σν) with Σν rank 1 in d=3, obtain O from the ε-perturbation limit, and compare Theorem 4.1's I2 and J2 formulas against a projected-gradient solve of the convex problems in §4.4; a mismatch would point to a counterexample.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing step is not Proposition 4.7, whose proof is self-contained, but the existence input supplied by Theorem 4.3: for arbitrary Σ1,Σ2 ∈ S_d^+ there must exist O ∈ O_d and C ∈ C_d such that O^*Σ1O and O^*Σ2O share the correlation matrix C, and then bw2(Σ1,Σ2) = Σ_i (√(O^*Σ1O)_ii − √(O^*Σ2O)_ii)^2. Proposition 4.9 constructs the O featured in Theorem 4.1 by applying Theorem 4.3 to (Σν, J2(Σν,Σµ)); the J2 formula in (4.1) and formula (4.10) likewise invoke Theorem 4.3. Section 4.1 only sketches the result: the positive-definite case is indicated via (4.2), and the singular case is explicitly left to the companion paper [5] ('we show (see [5])'). The singular case is not cosmetic: it is exactly the regime in which the J2 formula is needed and where the preprint claims new uniqueness results. Without an independent proof of Theorem 4.3, the central characterization is conditional rather than fully established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies Wasserstein projections onto sets of probability measures ordered in the convex order. It proves continuity of the projections when they are unique, a 1-Lipschitz (non-expansive) dependence of I_2(μ,ν) on μ, a local Hölder-1/2 dependence on ν, and then analyzes the Gaussian case. For Gaussian laws, it proves that I_2 is Gaussian, studies uniqueness of J_2 in the singular-covariance regime, gives a characterization of when non-Gaussian projections with the same covariance exist, and states an explicit orthogonal-diagonalization formula for the covariance matrices of both projections. The main technical engine for the Gaussian characterization is Theorem 4.3, a common-correlation reduction for the Bures-Wasserstein distance, whose proof is only sketched and partly deferred to a companion paper.","tokens_in":37144,"tokens_out":4591,"duration_ms":50957,"significance":"If fully established, the results are significant. The regularity results in Sections 1–2 are clean and useful, and the Gaussian analysis appears to settle a previously open uniqueness question for J_2 in the singular Gaussian case. The explicit formulas in Theorem 4.1, modulo Theorem 4.3, give a striking reduction to an axis-aligned computation and also support a practical projected-gradient algorithm. The proofs of Propositions 1.1, 2.2, 2.9, and 4.7 are careful and largely self-contained, and the paper includes explicit examples and a concrete numerical procedure. The main weakness is that the central Gaussian characterization depends on a theorem whose proof is not contained in the manuscript.","major_comments":[{"comment":"Theorem 4.3 is load-bearing: Proposition 4.9 invokes it to construct the orthogonal matrix O for positive-definite matrices and then extends by compactness to semidefinite ones; Theorem 4.1 uses (4.10) via Theorem 4.3; Corollary 4.11 and Proposition 3.3 both rely on Theorem 4.1. However, the proof of Theorem 4.3 is only sketched in the positive-definite case via (4.2), and the singular case is explicitly deferred with the sentence \"we show (see [5])\". The singular case is not a cosmetic edge case: it is exactly the regime needed for the J_2 formula (4.1) and for the new uniqueness results. As it stands, the Gaussian characterization is conditional on an unproved theorem, and the manuscript should either include a complete proof of Theorem 4.3 or clearly restate it with a proof in an appendix.","section":"Section 4.1 (Theorem 4.3)"},{"comment":"Proposition 3.3, which characterizes when the Gaussian projection is the unique W_2-projection in the singular case, is proved using Corollary 4.11 and Theorem 4.1, and therefore inherits the dependency on Theorem 4.3. Since Theorem 4.3 is not proved in this preprint, the uniqueness characterization is also conditional. This should be fixed together with the gap in Theorem 4.3; alternatively, these results should be moved to or explicitly attributed to the companion paper with a full proof provided.","section":"Section 3.3 (Proposition 3.3)"}],"minor_comments":[{"comment":"In the display after equation (2.5), the upper bound is written as \"(W_2(μ,I_2(μ,ν)) + W_2(μ,I_2(μ,ν))) W_2(ν,\\tilde ν)\", but the second term should be W_2(μ,I_2(μ,\\tilde ν)) to match the statement (2.4). This is a typo, but it makes the displayed estimate weaker than claimed.","section":"Section 2, proof of Proposition 2.9"},{"comment":"In the statement of Corollary 4.14, the assumption reads \"(O^*Σ_ν O)_{ii} > 0 and (O^*Σ_ν O)_{ii} > 0 for all i\", where the second inequality should refer to (O^*Σ_μ O)_{ii}. Please correct this typo.","section":"Section 4.3 (Corollary 4.14)"},{"comment":"The symbol W_2 is overloaded: it denotes the usual Wasserstein distance, and after Proposition 2.1 it is also used for the centered version W_2(μ,ν) = (W_2^2(μ,ν) − |m(μ)−m(ν)|^2)^{1/2}. The authors do warn the reader, but the reuse of the same symbol in equations such as (2.4) and its proof is a source of confusion; introducing a different symbol for the centered distance would improve readability.","section":"Notation and Section 2"}],"recommendation":"major_revision","confidential_remarks":"The decisive issue is the deferral of Theorem 4.3 to the companion paper [5], which is by the same authors. If the companion paper is available and accepted, the dependency may be acceptable for a journal, but the present preprint is not self-contained on its central claim. I would recommend asking the authors to either prove Theorem 4.3 in this paper or to state explicitly that the theorem is proven in [5] and to provide the full proof in an appendix if the journal insists on self-containedness. The regularity results in Sections 1–2 appear sound and could be published independently even if the Gaussian section is shortened."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Plain take: the regularity half of this paper is solid and self-contained; the Gaussian half is elegant but built on a theorem whose proof is in a companion paper. If you work in this area, read Sections 1–2 now; treat Section 4 as conditional until [5] is available.\n\nWhat's new and good: Proposition 2.2 (non-expansiveness of I2 in the first marginal) is a clean extension of the Hilbert projection argument, and Proposition 2.9's local 1/2-Hölder estimate in ν improves the known Lipschitz bounds in [22] and [3]. The proofs in Sections 1–2 are careful and do not cut corners. The Gaussian covariance characterization in Theorem 4.1 is a genuine first-principles derivation: if Theorem 4.3 holds, the formulas for I2 and J2 are explicit and likely to become standard tools. The singularity discussion in Section 3 is also new and worth attention.\n\nSoft spots: Theorem 4.3 is the load-bearing wall. It says any two PSD matrices can be simultaneously diagonalized in a certain correlation sense, and it is used to construct the orthogonal O that makes Theorem 4.1 work. In the positive definite case the paper sketches the construction; the singular case—exactly where the authors claim new uniqueness results—is deferred to the companion paper [5]. That is not a cosmetic gap: if Theorem 4.3 fails for singular matrices, the J2 formula and the uniqueness characterization lose their foundation. The abstract also says \"Hölder continuous\" without the qualifier \"locally\" that appears in the main text; minor, but worth fixing. These are the two things I would want addressed.\n\nWho this is for: people working on convex order constrained optimal transport, martingale optimal transport, and sampling of ordered measures. The regularity results are immediately useful; the Gaussian formulas will be useful if the companion proof is solid.\n\nRecommendation: send it to a serious referee. The regularity part is strong enough on its own, and the Gaussian part is important enough that the referee should push the authors to make the proof of Theorem 4.3 available, or at least to state clearly what remains assumed. Not a desk reject.","headline":"The regularity half is solid and self-contained; the Gaussian characterization is elegant but rests on a companion-paper theorem, so the paper is conditionally ready for serious refereeing.","tokens_in":37670,"tokens_out":2769,"would_cite":true,"duration_ms":29658,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60B10","60E05","49Q22"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proves that, for Gaussian measures, the two quadratic Wasserstein projections onto convex-order sets have explicitly characterized Gaussian covariance matrices, obtained by rotating both covariances into a common…","keywords":["optimal transport","Wasserstein projection","convex order","Gaussian measures","Bures-Wasserstein distance","correlation matrix","weak optimal transport"],"falsifier":"Take two non-commuting singular covariance matrices, compute their Bures-Wasserstein distance directly from the formula $\\operatorname{tr}(\\Sigma_1+\\Sigma_2-2(\\Sigma_1^{1/2}\\Sigma_2\\Sigma_1^{1/2})^{1/2})$, and check whether any orthogonal matrix $O$ makes this equal to the sum of squared differences of the square roots of the diagonal entries of $O^*\\Sigma_1O$ and $O^*\\Sigma_2O$; a pair where no such $O$ exists would falsify Theorem 4.3 and the formulas built on it.","tokens_in":36675,"feed_emoji":"📐","tokens_out":5401,"duration_ms":63193,"temperature":0.7,"pith_summary":"This paper establishes continuity of the Wasserstein projections in the convex order whenever they are unique, and then gives a complete Gaussian analysis. For the lower projection, the map sending a measure to its projection is non-expansive in the first argument and locally H\\\"older continuous with exponent 1/2 in the second. When the two measures are Gaussian, the lower projection is Gaussian, and even when the upper measure has singular covariance there is a unique Gaussian upper projection. The central structural result is that an orthogonal change of variables makes the two covariance computations behave like the diagonal case, so the projection covariance matrices can be written explicitly. This matters because these projections are used to restore convex order in numerical approximations, and explicit formulas turn the Gaussian case into a tractable matrix computation.","feed_headline":"Gaussian Wasserstein projections get explicit formulas","feed_subtitle":"One coordinate rotation reduces both convex-order projections on Gaussians to entrywise rescaling, including singular cases.","key_machinery":"The load-bearing tool is the common-correlation-matrix representation of the Bures-Wasserstein distance. Two matrices $\\Sigma_1,\\Sigma_2$ share a correlation matrix $C\\in C_d$ when $\\Sigma_i=\\operatorname{dg}(\\Sigma_i)^{1/2}C\\operatorname{dg}(\\Sigma_i)^{1/2}$; the paper states that some orthogonal $O$ always makes $O^*\\Sigma_1O$ and $O^*\\Sigma_2O$ share a correlation matrix, and in that frame the Bures-Wasserstein distance equals $\\sum_i(\\sqrt{(O^*\\Sigma_1O)_{ii}}-\\sqrt{(O^*\\Sigma_2O)_{ii}})^2$. Together with the diagonal rescaling matrix $D$, this reduces the projection formulas to one-dimensional comparisons of diagonal entries and supplies the optimal transport maps $x\\mapsto ODO^*x$ and their inverses.","core_discovery":"The main discovery is an explicit characterization of the covariance matrices of the two quadratic Wasserstein projections for Gaussian measures. The paper proves that for any covariance matrices $\\Sigma_\\mu, \\Sigma_\\nu\\in S_d^+$ there exists an orthogonal matrix $O$ such that, with $D=\\operatorname{diag}\\left(1\\wedge \\sqrt{(O^*\\Sigma_\\nu O)_{ii}/(O^*\\Sigma_\\mu O)_{ii}},\\ldots\\right)$, the inequality $DO^*\\Sigma_\\mu OD\\le O^*\\Sigma_\\nu O$ holds, and for any such $O$, the projected covariances are $I_2(\\Sigma_\\mu,\\Sigma_\\nu)=ODO^*\\Sigma_\\mu ODO^*$ and $J_2(\\Sigma_\\nu,\\Sigma_\\mu)=O\\tilde\\Sigma_J O^*$ with $\\tilde\\Sigma_J$ given explicitly in terms of $O^*\\Sigma_\\mu O$, $O^*\\Sigma_\\nu O$, and $D$. The paper also shows that when $\\Sigma_\\nu$ is singular and $d\\ge 2$, non-Gaussian upper projections can exist, but they all share the same covariance and the Gaussian representative is unique; a necessary and sufficient condition is given for the Gaussian projection to be the only projection at all.","pith_inferences":["This reader's inference: the explicit orthogonal frame turns the Gaussian projection problem into a diagonal semidefinite program, so the same formulas could accelerate the projected-gradient scheme in Section 4.4 and make high-dimensional Gaussian projections numerically routine.","This reader's inference: the construction of non-Gaussian projections for singular $\\nu$ uses martingale couplings between Gaussian conditionals; a testable extension is whether the same mechanism describes non-Gaussian projections for elliptically contoured or conditionally Gaussian inputs.","This reader's inference: the fact that the optimal transport maps between the original measures and the two projections are mutual inverses suggests that convex-order projection could be used to build explicit one-sided martingale couplings with exact covariance control in martingale optimal transport problems."],"forward_implications":["For Gaussian inputs, both projection covariance matrices can now be computed by an orthogonal diagonalization plus an entrywise rescaling, with no iterative optimization needed beyond finding the orthogonal matrix.","When $\\Sigma_\\nu$ is singular, the paper settles the previously open uniqueness question: the Gaussian upper projection is always unique, and all non-Gaussian projections, when they exist, have the same covariance matrix and the same lower-dimensional Gaussian marginals.","The inverse optimal transport maps between the lower and upper projections give concrete convex contraction and expansion maps, linking the Gaussian result to known backward and forward projection structure in convex order.","The projected-gradient algorithm in the final section computes $J_2$ first and then derives $I_2$ from the explicit formula, giving a practical route for numerical calculation in dimension higher than one.","Trace identities from the formulas show that the second moments of the two projections sum to the second moments of the original measures, a multidimensional Gaussian analogue of the one-dimensional identity for the projections on the line."],"supporting_citations":[{"why":"Introduces the Wasserstein projections in the convex order, gives the uniqueness conditions, and the identity $W_\\rho(\\mu,I_\\rho(\\mu,\\nu))=W_\\rho(\\nu,J_\\rho(\\nu,\\mu))$ used throughout the paper.","marker":"[3]"},{"why":"Companion paper that proves the common-correlation-matrix theorem, the key structural result on which Theorem 4.1 and the explicit projection formulas rely.","marker":"[5]"},{"why":"Derivation of the closed-form quadratic Wasserstein distance between Gaussian laws that the paper uses to identify the projection covariance matrices.","marker":"[16, 18, 26]"},{"why":"Gives the backward and forward Wasserstein projections in stochastic order and the convex contraction and expansion structure of the optimal transport maps.","marker":"[25]"},{"why":"Provides the barycentric weak optimal transport representation of the projection used in the regularity arguments of Section 2.","marker":"[1]"}],"fun_headline_variants":["Explicit covariance formulas for Gaussian Wasserstein projections","One rotation simplifies Gaussian projection formulas","Gaussian projections: explicit covariance even for singular cases","Unique Gaussian projections with explicit covariance"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire Gaussian characterization rests on the claim that one can always rotate both covariance matrices into a common correlation-matrix frame and that the Bures-Wasserstein distance then decomposes as a sum of one-dimensional squared differences of square roots of diagonal entries.","fun_headline_variants_meta":{"raw":{"variants":["Explicit covariance formulas for Gaussian Wasserstein projections","One rotation simplifies Gaussian projection formulas","Gaussian projections: explicit covariance even for singular cases","Unique Gaussian projections with explicit covariance"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000661,"raw_usage":{"total_tokens":3061,"prompt_tokens":1025,"completion_tokens":2036,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":641,"completion_tokens_details":{"reasoning_tokens":1982}},"tokens_in":641,"tokens_out":2036,"duration_ms":17385,"temperature":1.0,"reasoning_tokens":1982,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T21:27:20.391841+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take two non-commuting singular covariance matrices, compute their Bures-Wasserstein distance directly from the formula $\\operatorname{tr}(\\Sigma_1+\\Sigma_2-2(\\Sigma_1^{1/2}\\Sigma_2\\Sigma_1^{1/2})^{1/2})$, and check whether any orthogonal matrix $O$ makes this equal to the sum of squared differences of the square roots of the diagonal entries of $O^*\\Sigma_1O$ and $O^*\\Sigma_2O$; a pair where no such $O$ exists would falsify Theorem 4.3 and the formulas built on it.","supporting_citations":[{"cited_title":"Alfonsi, J","cited_arxiv_id":null,"evidence_quote":"Introduces the Wasserstein projections in the convex order, gives the uniqueness conditions, and the identity $W_\\rho(\\mu,I_\\rho(\\mu,\\nu))=W_\\rho(\\nu,J_\\rho(\\nu,\\mu))$ used throughout the paper."},{"cited_title":"Alfonsi and B","cited_arxiv_id":null,"evidence_quote":"Companion paper that proves the common-correlation-matrix theorem, the key structural result on which Theorem 4.1 and the explicit projection formulas rely."},{"cited_title":"Kim and Y","cited_arxiv_id":null,"evidence_quote":"Gives the backward and forward Wasserstein projections in stochastic order and the convex contraction and expansion structure of the optimal transport maps."},{"cited_title":"Alfonsi, J","cited_arxiv_id":null,"evidence_quote":"Provides the barycentric weak optimal transport representation of the projection used in the regularity arguments of Section 2."}],"review_version":1}