{"id":"db3f7c41-10c4-491b-a096-b0d32d9312a4","arxiv_id":"2412.09698","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"IPLA gives KL and Wasserstein convergence guarantees for inexact proximal Langevin sampling of convex potentials with polynomial growth beyond global gradient Lipschitz smoothness.","lead":"This paper introduces IPLA, a sampling algorithm that keeps Langevin Monte Carlo stable and provably accurate for potentials with super-quadratic growth, where gradients are not Lipschitz. It provides convergence rates and moment bounds, and demonstrates the method on light-tailed, Ginzburg-Landau, and Bayesian image deconvolution problems.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central rates in Theorems 5.5 and 5.7 are outsourced to Benko et al. (2024), an unreviewed companion by the same authors; if Lemma A.6 or the exact-PLMC contraction needs stronger assumptions than (V), the main claims are unsupported.","rationale":"The paper's algorithmic idea is new and the proof strategy is plausible: split the Wasserstein gradient flow into an inexact proximal step and a Gaussian smoothing step, then combine an energy-difference inequality with a telescoping Wasserstein argument. The moment bound in Appendix B is detailed, and the telescoping step in Proposition 5.4 is standard once the imported lemmas are granted. However, the decisive estimates are not supplied in this paper: Lemma A.6 is quoted as a black box, and the strongly convex Wasserstein result is delegated entirely to a companion preprint by the same authors. A single incorrect or over-general inequality in that companion would invalidate both the KL and Wasserstein convergence claims and the advertised d^{(qV+1)/2} epsilon^{-2} complexity. This is a genuine external-dependency risk, not a manufactured objection. The apparent non-convexity of the Ginzburg-Landau experiment and the loose epsilon-scaling in Corollary 5.8 are secondary: they concern the scope of a demonstration and the tightness of an upper bound, not the validity of the theorem under (V). The conditional verdict is appropriate: accept only after the companion inequalities are independently verified or supplied in full.","tokens_in":21740,"tokens_out":22657,"duration_ms":229916,"concrete_test":"Independently re-derive Lemma A.6(i)-(iii) and the exact-PLMC contraction claimed in Benko et al. (2024, Theorems 2 and 6) starting only from assumption (V), without consulting the companion. In particular, check whether Lemma A.6(ii) is valid for arbitrary nu in P_2 under (V) with R_V > 0, as Proposition 5.4 requires, and whether the companion's contraction holds without global lambda_V-convexity. If the derivation fails, or if it requires R_V = 0 or an extra global smoothness condition, then the KL and Wasserstein rates of Theorems 5.5 and 5.7 are not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing point is not an internal contradiction but the fact that the two main convergence theorems rest on inequalities taken verbatim from a companion preprint by the same four authors. Proposition 5.4, which drives Theorem 5.5, invokes Lemma A.6(i)-(iii), stated as Benko et al. (2024, Lemma 4.5); Theorem 5.7's Wasserstein bound is not proved here at all, only referred to via 'the same reasoning as in Benko et al. (2024, Proof of Theorem 2)' and 'same lines as ... Theorem 6'. These imported estimates are precisely the proximal-descent inequality, the Gaussian-smoothing bias K(tau), and the lambda-convex contraction that determine the tau^{1+alpha} error tolerance and the d^{(qV+1)/2} epsilon^{-2} complexity. If any of them requires global lambda_V-convexity (R_V = 0) or smoothness not stated in (V), then Theorem 5.5's scope is narrower than claimed. The Section 6 Ginzburg-Landau example has a negative quadratic term, so it is not globally convex and cannot be used to validate the (V) regime.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces the Inexact Proximal Langevin Algorithm (IPLA), a splitting scheme in which the exact proximal step of the proximal Langevin algorithm is replaced by an approximation with error bounded by δ, followed by an additive Gaussian step. Under assumption (V) (global convexity, λ_V-convexity outside a ball, and a polynomial-growth majorant of order q_V+1) and a finite initial moment of order q_V+1, the paper proves uniform moment bounds (Theorem 5.1), an averaged KL error bound (Theorem 5.5), and a Wasserstein contraction bound under global λ_V-convexity (Theorem 5.7). Complexity corollaries claim d^{(q_V+1)/2} O(ε^{-2}) iteration counts in both KL and Wasserstein settings. The numerical section tests IPLA on a quartic-potential example, a Ginzburg–Landau model, and a Bayesian image deconvolution task, with a code repository provided.","tokens_in":22017,"tokens_out":15327,"duration_ms":128804,"significance":"If the main results hold, the paper makes a useful contribution by extending Langevin-type sampling to potentials with non-Lipschitz, super-quadratic growth while keeping dimension dependence comparable to the best-known LMC rates when q_V=1. The moment bounds and the inexact-proximal error propagation analysis are nontrivial and of independent interest. The paper also provides reproducible code and clearly structured experiments. However, the central convergence theorems currently rest on inequalities imported from a same-author companion preprint, so the contribution is conditional unless those dependencies are resolved.","major_comments":[{"comment":"The main convergence results depend critically on Lemma A.6, which is stated as Benko et al. (2024, Lemma 4.5) and not proved in this manuscript. Proposition 5.4 and Theorem 5.5 invoke Lemma A.6(i)–(iii), and Theorem 5.7 additionally invokes the exact-PLMC contraction 'by the same reasoning as in Benko et al. (2024, Proof of Theorem 2)' and 'the same lines as ... Theorem 6'. These imported estimates provide the proximal-descent inequality, the Gaussian-smoothing bias K(τ), and the contraction that determine the allowable δ and the reported complexity. Since the companion preprint is by the same four authors and is not independently reviewed, the paper should either give complete proofs of these inequalities in the appendix or state and verify the precise hypotheses under which they hold; as written, the main claims are conditional on an external unpublished source.","section":"Sec. 5 and App. A, Lemma A.6"},{"comment":"The stated target is W_2^2(ρ_{n_ε}, μ_*) ≤ ε, but the reported iteration counts O(ε^{-2}) and O(ε^{-α^{-1}}) correspond to a target of the form W_2 ≤ ε (i.e., W_2^2 ≤ ε^2). Substituting the corollary's own conditions, Remark 5.3 gives K(τ_ε) ≤ C τ_ε d^{(q_V+1)/2}, so K(τ_ε) ≤ λ_V ε / 12 forces τ_ε ≲ ε d^{-(q_V+1)/2}; with n_ε ≍ τ_ε^{-1}, this yields d^{(q_V+1)/2} O(ε^{-1}) iterations when α ≥ 1/2. For α < 1/2, the condition τ_ε^{2α} ≤ λ_V^2 ε/(96κ^2 log^2(...)) dominates and gives τ_ε ≤ ε^{1/(2α)}, hence d^{(q_V+1)/2} O(ε^{-1/(2α)}) iterations, not d^{(q_V+1)/2} O(ε^{-α^{-1}}). The target and the complexity exponents must be made consistent.","section":"Corollary 5.8"},{"comment":"In bounding the term I_2 = F_V[ρ_{k+2/3}] - F_V[ρ_{k+1/3}], the proof applies Lemma A.8 with the arbitrary measure ν. However, Lemma A.8 applies to the measure being smoothed, which here is ρ_{k+1/3}, not ν. Thus the displayed bound on I_2 does not follow from Lemma A.8 as written. The constant C(ν) in Appendix C also does not reflect moments of ρ_{k+1/3} in that first part. This is likely fixable using Theorem 5.1, but the proof needs correction before the KL bound can be considered established.","section":"App. B, proof of Proposition 5.4"},{"comment":"The Ginzburg–Landau potential with the stated parameters υ=2, κ=0.1, ς=0.5 has a negative quadratic contribution (1-υ)/2 = -0.5, so along constant configurations the Hessian at zero is negative and V is not convex on R^d. This violates the global convexity requirement in assumption (V). The example is presented as a demonstration of IPLA, but the proven guarantees do not cover it as parameterized. The authors should either extend the theory to this setting, change the parameters, or explicitly state that Example 2 is outside the theorem's scope, as they already do for the image-deconvolution example.","section":"Sec. 6, Example 2"}],"minor_comments":[{"comment":"The first displayed inequality in Lemma A.1 appears to be missing a square: the term should be |prox_τ^V(x) - z|^2, not |prox_τ^V(x) - z|.","section":"Lemma A.1"},{"comment":"The statement of Lemma A.6(ii) writes W_2^2(ρ_{k+1/3}, ρ_k) as the second Wasserstein term, but the proof of Proposition 5.4 uses W_2^2(ρ_{k+1/3}, ν); the lemma statement is likely a typo and should be corrected.","section":"Lemma A.6(ii)"},{"comment":"The text says 'dimension d = 10 3'; this should read d = 10^3.","section":"Sec. 6, Example 1"},{"comment":"The bound (1 - e^{-λ_V τ(n_ε-1)})/(1 - e^{-λ_V τ}) ≤ n_ε is very loose and is stated without justification; since τ λ_V < 1, the denominator is bounded below by a constant, so a constant bound also holds. The proof should be clarified.","section":"Corollary 5.8 proof"}],"recommendation":"major_revision","confidential_remarks":"The main editorial risk is the heavy reliance on the same-author companion preprint for the core estimates. I would recommend requesting that the authors include full proofs or precise verifiable statements of Lemma A.6 and the exact-PLMC contraction used in Theorem 5.7, and that they fix the Wasserstein complexity statement. The experiments and code are a strength, but the theoretical claims need to be self-contained enough for a journal referee to verify them."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Mate, here's my read on 2412.09698. The core idea is genuinely new: an inexact-proximal Langevin algorithm (IPLA) that handles convex potentials with polynomial growth beyond global Lipschitz gradients, with explicit error control on the proximal step and bounds on all moments. The complexity statement — d^{(qV+1)/2} O(ε^{-2}) for KL error, with the qV=1 case recovering the best-known LMC rate — is worth taking seriously. The moment-bound proof (Theorem 5.1) looks self-contained and solid. They also ship code, which helps.\n\nThe soft spots are real but reworkable. The two main convergence theorems are propped up by Lemma A.6 and an exact-PLMC contraction taken from Benko et al. (2024), a companion preprint by the same four authors that hasn't been peer-reviewed. The lemma is actually stated in the appendix, so the assumptions are visible — they match (V) — but the proof is elsewhere, and Theorem 5.7's Wasserstein bound is delegated with 'same reasoning as' references. A referee can't fully verify this without chasing an unreviewed manuscript. If any of those imported inequalities secretly need global λV-convexity or extra smoothness, the main claims lose their advertised scope. That's the biggest risk.\n\nSecond, Corollary 5.8 contains what looks like a genuine arithmetic slip: it targets W_2^2 ≤ ε and then claims O(ε^{-2}) iterations, but plugging their own conditions on τ and n yields O(ε^{-1/(2α)}) — for α ≥ 1/2 that's O(ε^{-1}), up to logs. So the stated complexity is pessimistic rather than wrong, but it's inconsistent as written.\n\nThird, the Ginzburg-Landau experiment in Section 6 uses (1-υ)/2 = -0.5, a negative quadratic term. That makes V non-convex globally, so the example doesn't satisfy the convexity bullet in (V). The theory may still hold under local strong convexity in the tails, but then the paper should say so and adjust the assumptions.\n\nOverall, I'd rate this as a qualified accept: the algorithmic contribution is novel and useful, the moment bounds are a nice standalone result, and the flaws are fixable. The paper belongs on the desk of someone working on MCMC for super-quadratic or non-smooth targets. Send it to a serious venue with a request to move the key lemmas into the paper, fix the corollary, and redo or relabel the G-L experiment.","headline":"New inexact-proximal LMC scheme with real convergence claims; the key estimates are outsourced to a companion preprint and one experiment overreaches, but the core is worth refereeing.","tokens_in":22556,"tokens_out":6360,"would_cite":true,"duration_ms":57443,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["65C05","60J22"],"pacs":[],"model":"deepseek-v4-flash","headline":"Inexact Proximal Langevin Algorithm samples from super-quadratic, non-Lipschitz-gradient potentials by keeping the proximal-step error below $\\kappa\\tau^{1+\\alpha}$, with iteration cost $d^{(q_V+1)/2} O(\\varepsilon^{-2})$.","keywords":["Langevin Monte Carlo","proximal sampling","non-smooth potentials","super-quadratic growth","inexact proximal operator","KL divergence","Wasserstein distance","moment bounds"],"falsifier":"Run IPLA on $V(x)=|x|^4/4$ in dimension $d=1000$ with a step $\\tau < 1/\\lambda_V$, an initial measure with finite fourth moment, and proximal error $\\delta \\le \\kappa \\tau^2$, then estimate the KL divergence between the averaged empirical chain and the target $\\mu_* \\propto e^{-|x|^4/4}$. Theorem 5.5 predicts $KL(\\nu_n^0 | \\mu_*) \\le (2n\\tau)^{-1} W_2^2(\\rho_0,\\mu_*) + C\\kappa\\tau + K(\\tau)$; if the measured KL consistently exceeds that right-hand side, or if any moment $\\mathbb{E}|X_k|^m$ grows without bound over $n$, the central bound is wrong.","tokens_in":21537,"feed_emoji":"🎲","tokens_out":14767,"duration_ms":121107,"temperature":0.7,"pith_summary":"The paper introduces IPLA, an Inexact Proximal Langevin Algorithm, and claims it can sample from densities proportional to $\\exp(-V)$ for potentials $V$ that are convex, strongly convex outside a ball, and have polynomial growth of order $q_V+1 \\ge 2$, without requiring the gradient $\\nabla V$ to be globally Lipschitz. This is the regime where ordinary Langevin Monte Carlo can become transient and blow up, and tamed versions remain stable but move sluggishly in the tails. The main theorems state that if the proximal step is computed to error $\\delta \\le \\kappa \\tau^{1+\\alpha}$ with $\\tau < 1/\\lambda_V$, then the averaged chain converges in KL divergence at the rate of Theorem 5.5, and in the globally strongly convex case the chain converges in Wasserstein distance at the rate of Theorem 5.7. As a consequence, generating one sample with accuracy $\\varepsilon$ costs $d^{(q_V+1)/2} O(\\varepsilon^{-2})$ iterations, which for $q_V=1$ matches the best-known dimension dependence of Langevin Monte Carlo. The paper also proves uniform bounds on all moments of the chain and demonstrates the method on a quartic light-tailed target, a Ginzburg-Landau model, and a 360000-dimensional Bayesian image deconvolution problem.","feed_headline":"Inexact proximal steps tame super-quadratic sampling","feed_subtitle":"A Langevin variant samples convex, faster-than-quadratic targets while keeping the best-known dimension dependence.","key_machinery":"The object that carries the argument is IPLA itself, Algorithm 1: each iteration applies an approximate proximal map $x \\mapsto \\operatorname{prox}_{\\tau V}(x)$, the minimizer of $V(y) + |y-x|^2/(2\\tau)$, with output error bounded by $\\delta$, then adds Gaussian noise $Z \\sim N(0,2\\tau I_d)$. The proof uses the variational splitting of the free energy $F = F_V + F_E$ into potential and entropy terms, whose Wasserstein gradient flows are respectively the proximal map and Brownian motion. Its core estimate is Proposition 5.4, a Wasserstein-space analogue of the classical convex-optimization descent inequality $2\\tau(f(x_{k+1}) - f(x_*)) \\le |x_k - x_*|^2 - |x_{k+1} - x_*|^2 + C\\tau^2$, with the additional term $C(\\nu)\\delta + K(\\tau)\\tau$ measuring the cost of the inexact step. The constant $K(\\tau)$ is explicit in equation (5), and the moment bound of Theorem 5.1, proved by induction with Lemmas A.2 and A.3, controls all moments of the chain so that the super-quadratic tails cannot push the chain to infinity.","core_discovery":"The central discovery is that the exact proximal map in the Proximal Langevin Algorithm can be replaced by an approximate one without breaking convergence, as long as the approximation error is tied to the step size by $\\delta = \\kappa \\tau^{1+\\alpha}$. Under assumption (V), Theorem 5.5 bounds the Kullback-Leibler divergence of the averaged chain by $$KL(\\nu_N^n | \\mu_*) \\le \\frac{1}{2n\\tau}\\bigl($W_2^{2}$(\\rho_N,\\mu_*) - $W_2^{2}$(\\rho_{N+n},\\mu_*)\\bigr) + C(\\mu_*) \\kappa \\tau^\\$\\alpha$ + K(\\tau),$$ and Theorem 5.7, for globally $\\lambda_V$-convex $V$, gives $$$W_2^{2}$(\\rho_k,\\mu_*) \\le 2\\left(1 - \\frac{\\tau\\lambda_V}{2}\\right)^k $W_2^{2}$(\\rho_0,\\mu_*) + \\frac{4}{\\lambda_V}K(\\tau) + 2\\$kappa^{2}$ \\$tau^{{2+2\\alpha}}$\\left(\\frac{1-$e^{{-\\lambda_V\\tau(k-1)}}$}{1-$e^{{-\\lambda_V\\tau}}$}\\right)^2.$$ These bounds imply the claimed $d^{(q_V+1)/2} O(\\varepsilon^{-2})$ iteration complexity for one sample, with the proximal accuracy requirement being $\\delta \\le \\kappa \\tau^2$ for the KL result and $\\delta \\le \\kappa \\tau^{3/2}$ in the strongly convex case. The authors present these bounds as the proof that IPLA extends Langevin Monte Carlo to potentials whose gradients are not Lipschitz, and they support the theory with experiments in which ULA blows up while IPLA remains stable.","pith_inferences":["A natural extension the paper does not pursue is a stochastic model of the proximal error; a Markov-noise version would clarify whether the deterministic bound $\\delta \\le \\kappa \\tau^{1+\\alpha}$ is necessary or merely sufficient.","The dimension exponent is controlled by $q_V$, so the explicit constant in $K(\\tau)$ could be tested directly by measuring KL against $\\tau$ for $V(x)=|x|^4/4$ and comparing with Remark 5.3.","If the Ginzburg-Landau experiment lies outside the global-convexity assumption, its stability hints that the contraction mechanism may tolerate local non-convexity, and a dissipativity or local-convexity version of Theorems 5.5 and 5.7 would be the next testable step."],"forward_implications":["IPLA remains stable on targets with faster-than-quadratic polynomial tails, where ULA diverges and TULA takes unnecessarily small steps, as demonstrated on the quartic and Ginzburg-Landau experiments.","With proximal error $\\delta \\le \\kappa \\tau^2$, KL accuracy $\\varepsilon$ is reached in $O(\\varepsilon^{-2})$ iterations with dimension factor $d^{(q_V+1)/2}$, and $q_V=1$ recovers the best-known dimension scaling for Langevin Monte Carlo.","When $V$ is globally strongly convex, the proximal step may be solved less accurately ($\\delta \\le \\kappa \\tau^{3/2}$) and the Wasserstein error still converges at $O(\\varepsilon^{-2})$ order.","All moments of the IPLA chain are finite uniformly in time for every $\\tau<1/\\lambda_V$, so the chain does not explode even though the gradient is not Lipschitz."],"supporting_citations":[{"why":"Supplies Lemma A.6 and the exact-PLMC contraction estimates that Proposition 5.4 and Theorem 5.7 call as black boxes.","marker":"Benko et al. (2024)"},{"why":"Provides the Wasserstein gradient-flow theory and the formula for the proximal operator that defines PLMC.","marker":"Ambrosio, Gigli, and Savaré (2008)"},{"why":"Establishes the variational formulation of the Fokker-Planck equation that identifies $\\mu_*$ as the minimizer of the free energy $F$.","marker":"Jordan, Kinderlehrer, and Otto (1998)"},{"why":"Gives the convex-analysis framework for LMC and the best-known dimension dependence that IPLA's $q_V=1$ case matches.","marker":"Durmus, Majewski, and Miasojedow (2019)"},{"why":"Defines TULA, the main baseline for super-quadratic potentials, and supplies the Ginzburg-Landau convexity facts used in Section 6.","marker":"Brosse et al. (2019)"},{"why":"Provides the Unadjusted Barker algorithm whose complexity IPLA matches while avoiding its one-sided Lipschitz assumption.","marker":"Livingstone et al. (2024)"},{"why":"Contains the first-order optimization inequality that Proposition 5.4 transplants to the Wasserstein space.","marker":"Beck (2017)"},{"why":"Justifies the convexity of KL divergence used to pass from individual measures to the averaged chain in Theorem 5.5.","marker":"Cover and Thomas (2012)"}],"fun_headline_variants":["Langevin sampling without Lipschitz gradients: now possible","Approximate proximal steps lift Langevin's Lipschitz limit","IPLA: sampling from super-quadratic targets with ease","Beyond Lipschitz: a proximal trick for Langevin Monte Carlo","Inexact proximal maps make Langevin work for steep targets"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Two load-bearing premises hold the proof together: the contraction estimates for the exact proximal chain are imported as a black box from a companion preprint by the same authors, and the potential $V$ is assumed convex on all of $\\mathbb{R}^d$, an assumption that the Ginzburg-Landau experiment in Section 6, as parameterized, appears to violate; if either fails, the stated rates no longer follow.","fun_headline_variants_meta":{"raw":{"variants":["Langevin sampling without Lipschitz gradients: now possible","Approximate proximal steps lift Langevin's Lipschitz limit","IPLA: sampling from super-quadratic targets with ease","Beyond Lipschitz: a proximal trick for Langevin Monte Carlo","Inexact proximal maps make Langevin work for steep targets"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000583,"raw_usage":{"total_tokens":2785,"prompt_tokens":1026,"completion_tokens":1759,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":642,"completion_tokens_details":{"reasoning_tokens":1672}},"tokens_in":642,"tokens_out":1759,"duration_ms":13431,"temperature":1.0,"reasoning_tokens":1672,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T16:50:37.772944+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run IPLA on $V(x)=|x|^4/4$ in dimension $d=1000$ with a step $\\tau < 1/\\lambda_V$, an initial measure with finite fourth moment, and proximal error $\\delta \\le \\kappa \\tau^2$, then estimate the KL divergence between the averaged empirical chain and the target $\\mu_* \\propto e^{-|x|^4/4}$. Theorem 5.5 predicts $KL(\\nu_n^0 | \\mu_*) \\le (2n\\tau)^{-1} W_2^2(\\rho_0,\\mu_*) + C\\kappa\\tau + K(\\tau)$; if the measured KL consistently exceeds that right-hand side, or if any moment $\\mathbb{E}|X_k|^m$ grows without bound over $n$, the central bound is wrong.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies Lemma A.6 and the exact-PLMC contraction estimates that Proposition 5.4 and Theorem 5.7 call as black boxes."},{"cited_title":"urich. Birkh\\","cited_arxiv_id":null,"evidence_quote":"Provides the Wasserstein gradient-flow theory and the formula for the proximal operator that defines PLMC."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the variational formulation of the Fokker-Planck equation that identifies $\\mu_*$ as the minimizer of the free energy $F$."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines TULA, the main baseline for super-quadratic potentials, and supplies the Ginzburg-Landau convexity facts used in Section 6."},{"cited_title":"Skew-symmetric schemes for stochastic differential equations with non-Lipschitz drift: an unadjusted Barker algorithm","cited_arxiv_id":"2405.14373","evidence_quote":"Provides the Unadjusted Barker algorithm whose complexity IPLA matches while avoiding its one-sided Lipschitz assumption."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Contains the first-order optimization inequality that Proposition 5.4 transplants to the Wasserstein space."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Justifies the convexity of KL divergence used to pass from individual measures to the averaged chain in Theorem 5.5."}],"review_version":1}