{"id":"a40f0e78-1f09-40a1-a90e-8bd79427d75d","arxiv_id":"2504.20400","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The HK-Boltzmann gradient flow preserves Gaussianity, and the reduced equations for mean, covariance, and mass admit exponential convergence rates with explicit dependence on the geometry parameters.","lead":"This paper derives the exact equations that govern how a Gaussian distribution changes when driven by a gradient flow mixing optimal transport and Hellinger dynamics, and proves exponential convergence to a Gaussian target. The explicit convergence rates matter for Bayesian inference, where Gaussian approximations are the standard workhorse.","discovery_kind":"first_principles","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Slope comparison in Theorem 5.11 is inverted: the correct bound is |∂OttoE|² ≤ σmin^{-1}|∂SHeE|², not ≤ σmin|∂SHeE|², so the stated non-Gaussian decay rate is unsupported.","rationale":"The core reduction is solid: Lemma 3.6 computes the reduced Onsager operators explicitly, and Theorem 3.8 verifies by direct substitution that the ODEs (2.4) are exactly ẋ=-K^{red}DE. Thus the central Gaussian-invariance/reduction claim is well supported. The PL estimates for Gaussian targets in Sections 5.1–5.2 are mostly routine, although Theorem 5.8's displayed bound drops the prefactor appearing in (5.9)–(5.10), so that statement also needs a correction. The most load-bearing issue is the slope comparison in Theorem 5.11: the displayed inequality |∂OttoE|² ≤ σmin|∂SHeE|² is algebraically inverted. Because this comparison is the step that converts energy dissipation into exponential decay for non-Gaussian targets, the theorem's stated rate is unsupported. However, the corrected inequality still gives a positive rate with σmin in place of σmin^{-1}, so the theorem is repairable rather than fundamentally false. This confirms the reader's condition: the paper should be published only after the decay statements in Theorem 5.8 and Theorem 5.11 are fixed. I therefore keep the reader's conditional verdict unchanged.","tokens_in":29043,"tokens_out":18609,"duration_ms":180472,"concrete_test":"Evaluate the two slope formulas in Section 5.3 for d=1, target V(x)=λx²/2, m=0, and Σ=ε<1 with λ=1. Then Γ^{-1}=1, g=0, and the paper's claimed inequality reduces to (1-ε)²/ε ≤ ε(1-ε)², which is false for ε<1. Independently re-derive the comparison using tr(A²Σ^{-1}) ≤ σmin^{-1}tr(A²) and |g|² ≤ σmin^{-1}gᵀΣg; this gives the opposite direction. Finally, rerun the Gronwall argument with |∂SHeE|² ≥ σmin|∂OttoE|² and σmin ≥ r_E^{-1} to see that Theorem 5.11's decay holds with a corrected rate constant involving r_E^{-1}, whereas the stated γ=λ(2α+βr_E^{-1}) does not follow.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 5.3 derives the two slopes for a log-λ-concave target π=e^{-V}dx: with A=I-Γ^{-1}Σ and g=∫∇V G(Σ,m)dx, it claims |∂OttoE|² ≤ σmin|∂SHeE|², where σmin is the smallest eigenvalue of Σ. This inequality is backwards. Since Σ^{-1} ≤ σmin^{-1}I and |g|² ≤ σmin^{-1}gᵀΣg, the trace and inner-product terms give |∂OttoE|² ≤ σmin^{-1}|∂SHeE|². The displayed claim already fails for d=1, Γ=1, Σ=ε<1, g=0: the left side is (1-ε)²/ε while the right side is ε(1-ε)², which is false for 0<ε<1. Consequently the Gronwall chain just below it, and in particular the rate γ=λ(2α+βr_E^{-1}) in Theorem 5.11, is not established as stated. The good news is that the corrected direction, |∂SHeE|² ≥ σmin|∂OttoE|², together with σmin ≥ r_E^{-1} from Proposition 5.10, yields an exponential decay estimate with a positive rate of the form λ(2α+2βr_E^{-1}) (up to the slope normalization). So the qualitative claim of Theorem 5.11 is repairable, but the theorem as printed rests on an inverted inequality.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the gradient flow of the relative Boltzmann entropy with respect to Gaussian reference measures in Hellinger-Kantorovich (HK) geometry. It proves that the class of scaled Gaussians is invariant under the HK-Boltzmann flow, derives the reduced Onsager operators on the parameter space P=(Sigma,m,kappa), obtains the explicit ODE system (2.4)/(3.6), and shows that this parameter-space system is the exact reduced gradient flow. The paper then analyzes geodesic convexity of the normalized Gaussian system, establishes Polyak-Lojasiewicz-type inequalities and refined eigenvalue-based decay estimates for Gaussian targets, extends the results to strongly log-lambda-concave targets, and reports numerical experiments including Bayesian logistic regression.","tokens_in":29438,"tokens_out":18921,"duration_ms":187054,"significance":"The reduction results are a genuine contribution: the closed-form reduced Onsager operators in Theorem 3.8 and the explicit ODEs in (2.4) give a parameter-free, computationally transparent description of HK-Boltzmann flow on Gaussians. The geodesic-convexity analysis clarifies the loss of global convexity for beta>0 while retaining sublevel semi-convexity, and the PL estimates provide explicit rates in alpha, beta, and Gamma. The numerical experiments complement the theory. However, the non-Gaussian decay theorem in Section 5.3 rests on an inverted slope comparison, so that part of the paper needs correction before the stated results can be accepted.","major_comments":[{"comment":"The displayed inequality |\\partial_Otto E_1|^2 <= sigma_min |\\partial_SHe E_1|^2 is reversed. Writing A=I-Gamma^{-1}Sigma and g=integral nabla V G dx, the two slopes are |\\partial_Otto E_1|^2=tr(A^2 Sigma^{-1})+|g|^2 and |\\partial_SHe E_1|^2=tr(A^2)+g^T Sigma g. Since Sigma^{-1} <= sigma_min^{-1} I and |g|^2 <= sigma_min^{-1} g^T Sigma g, the correct comparison is |\\partial_Otto E_1|^2 <= sigma_min^{-1}|\\partial_SHe E_1|^2. The printed claim already fails for d=1, Gamma=1, Sigma=epsilon in (0,1), g=0, where it would assert (1-epsilon)^2/epsilon <= epsilon(1-epsilon)^2. The Gronwall chain immediately below the display therefore does not follow as written: the valid argument gives dE_1/dt <= -(alpha+beta sigma_min)|\\partial_Otto E_1|^2, and with (5.12) this yields a rate of the form lambda(2alpha+2beta r_E^{-1}), not lambda(2alpha+beta r_E^{-1}). The qualitative statement of Theorem 5.11 is repairable, but the theorem as stated is not supported by the displayed computation.","section":"Section 5.3, slope comparison preceding Theorem 5.11"},{"comment":"The statement of Theorem 5.8 omits the multiplicative prefactors that are part of the estimates proved in (5.9) and (5.10). The theorem displays E_1(p(t)) <= e^{-nu_cov t}H_cov(Sigma(0)) + e^{-nu_m t}H_m(m(0)), whereas (5.9) and (5.10) contain the constants [min{1,b_min(0)}]^{-beta/nu_cov} and [min{1,b_min(0)}]^{-2beta/nu_m}, which can exceed 1 when b_min(0)<1. As printed the theorem is false; either the prefactors must be included in the display or the statement should explicitly say that the bounds hold up to those constants.","section":"Theorem 5.8 and Eqs. (5.9)-(5.10)"}],"minor_comments":[{"comment":"The formula DE_1(p)=(2(Gamma^{-1}-Sigma^{-1}), Gamma^{-1}(m-n)) is inconsistent with E_1(p)=H(p|Gamma,n), whose covariance derivative is (1/2)(Gamma^{-1}-Sigma^{-1}); it is also inconsistent with the vector field V(p) displayed immediately below. The downstream formulas in (5.4) appear to use the correct prefactor, so please correct the display.","section":"Section 4.2, displayed DE1"},{"comment":"For the Otto part at kappa=1, the sign of the first term in \\dot S_s is opposite to Hamilton's equations (4.3); in the scalar case it should be -2 alpha S_s^2, not +2 alpha S_s^2. Since the remark is declared unused this is not load-bearing, but it should be corrected or the remark removed.","section":"Remark 4.1, Eq. (4.4c)"},{"comment":"The parenthetical 'cf. Remark' is a dangling reference; please supply the intended number or delete the parenthetical.","section":"Remark 5.9"},{"comment":"The symbol b_i(t) is used both for the evolving eigenvalue and for the explicit upper/lower bound 1+(b_i(0)-1)e^{-nu_cov t} in the same display; please use a different symbol, such as \\bar b_i(t), for the bound.","section":"Lemma 5.7, Eq. (5.8)"}],"recommendation":"major_revision","confidential_remarks":"For the editor: the core reduction in Theorems 3.1, 3.2, and 3.8 is careful and sound, and the main issue is a localized, repairable error in Section 5.3. The prefactor omission in Theorem 5.8 and the notation/display issues are also easily fixed. I do not see grounds for rejection; the paper would be suitable for publication after the stated theorem is corrected and the remaining presentation issues are addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe thing to know: this paper does the reduction of the Hellinger-Kantorovich-Boltzmann gradient flow to the Gaussian manifold properly. The reduced Onsager operator (3.3) is explicit and additive in alpha and beta, the ODEs (2.4) follow by direct computation, and the PL-based decay estimates for Gaussian targets are a real step beyond the Wasserstein-only [LC*22] and Fisher-Rao-only [CH*23] treatments. The convexity dichotomy (global geodesic convexity iff beta=0, with sublevel semi-convexity otherwise) is cleanly derived. The numerical experiments are appropriate illustrations, not overclaimed.\n\nThe central reduction is solid: Proposition 2.1 and Theorem 3.8 check out. The paper is honest about which parts rest on prior work ([MaM20] for the reduction recipe, [MiZ25] for the shape-mass relation). No fitted parameters, no circular reasoning.\n\nNow the soft spots, in proportion.\n\nFirst, Section 5.3: the inequality |dOttoE1|^2 <= sigma_min |dSHeE1|^2 is backwards. Direct computation gives |dOttoE1|^2 <= sigma_min^{-1} |dSHeE1|^2, since Sigma^{-1} <= sigma_min^{-1} I and |g|^2 <= sigma_min^{-1} g^T Sigma g. The displayed claim already fails for d=1, Gamma=1, Sigma=eps<1, g=0. Consequently the Gronwall chain in Theorem 5.11 is not established as stated. That said, the corrected inequality still gives exponential decay, with rate lambda(2alpha+2beta r_E^{-1}) rather than the printed lambda(2alpha+beta r_E^{-1}), so the qualitative theorem is repairable and the printed rate is, if anything, conservative.\n\nSecond, Theorem 5.8 as stated omits the prefactors that appear in (5.9)-(5.10): the factors min{1,b_min(0)}^{-beta/nu_cov} and min{1,b_min(0)}^{-2beta/nu_m}. When b_min(0)<1 these exceed 1, so the clean inequality E1(p(t)) <= e^{-nu_cov t} Hcov + e^{-nu_m t} Hm is false. The rate statement survives with an initial-condition-dependent constant, but the theorem overstates.\n\nMinor: Remark B.1's gradient structure for (2.3a) is informal; not a problem for the main results.\n\nOverall: the reduction machinery and the Gaussian-target decay analysis are worth engaging with. The errors are localized to Section 5 statements, not the core derivation. A serious referee should ask for a corrected Theorem 5.11 (or a caveat) and a restated Theorem 5.8 with prefactors. I would send this to peer review rather than desk reject.","headline":"Solid reduction of HK-Boltzmann flow to Gaussian ODEs, but Section 5 overstates two results: the slope comparison in Theorem 5.11 is inverted and Theorem 5.8 drops initial-data prefactors.","tokens_in":29915,"tokens_out":6182,"would_cite":true,"duration_ms":50764,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["49Q22","35Q49"],"pacs":[],"model":"deepseek-v4-flash","headline":"The Hellinger–Kantorovich–Boltzmann gradient flow preserves the class of Gaussian measures, and the paper derives the exact finite-dimensional ODEs for their mean, covariance, and mass.","keywords":["Gaussian measures","Hellinger-Kantorovich distance","gradient-flow equation","relative Boltzmann entropy","Kullback-Leibler divergence","reduced Onsager operator","Polyak-Lojasiewicz inequality","log-concave target"],"falsifier":"Take $d=1$, target $\\pi=\\mathcal{N}(0,1)$, covariance $\\Sigma=s^2$ with $s<1$, and any mean $m\\ne0$. Write $g=\\int \\nabla V\\,G\\,dx$. The transport slope is $(1-s^2)^2/s^2+g^2$ and the Hellinger slope is $(1-s^2)^2+s^2g^2$. The inequality used in Theorem 5.11 would require $(1-s^2)^2/s^2+g^2 \\le s^2((1-s^2)^2+s^2g^2)$, which simplifies to $(1-s^4)((1-s^2)^2+s^2g^2)\\le0$, false for every $s<1$. This one-dimensional check falsifies the slope comparison and hence the stated decay rate.","tokens_in":28884,"feed_emoji":"📉","tokens_out":22585,"duration_ms":183716,"temperature":0.7,"pith_summary":"The paper establishes that, when the target measure is a (scaled) Gaussian, the gradient flow of the relative Boltzmann entropy in the Hellinger–Kantorovich (HK) geometry — the metric that combines spatial transport with mass growth — leaves the class of Gaussian measures invariant. Because of this invariance, the infinite-dimensional flow reduces exactly to a finite-dimensional gradient system on the parameters $(\\Sigma,m,\\kappa)$, with explicit ODEs for the covariance, mean, and mass. The reduction is carried out through a reduced Onsager operator that is a linear combination of the transport and Hellinger parts, so the additive structure of the HK metric survives in the parameter space. Using this reduced structure, the paper proves exponential convergence to equilibrium for Gaussian targets, with decay rates refined by tracking the eigenvalues of the normalized covariance, and it extends the analysis to strongly log-$\\lambda$-concave targets. Together, these results give a tractable family of dynamics for Gaussian variational inference, with explicit convergence rates in the Gaussian-target case and numerical support in non-Gaussian applications.","feed_headline":"Gaussians stay Gaussian under the Hellinger-Kantorovich flow","feed_subtitle":"Exact finite-dimensional ODEs give exponential convergence rates that depend only on α, β, and the target covariance.","key_machinery":"The load-bearing object is the reduced Onsager operator on the parameter space $P=\\mathbb{R}^{d\\times d}_{\\mathrm{spd}}\\times\\mathbb{R}^d\\times(0,\\infty)$. It is obtained by restricting the full HK Onsager operator $K_{\\alpha,\\beta}(\\rho)\\xi=-\\alpha\\,\\mathrm{div}(\\rho\\nabla\\xi)+\\beta\\rho\\xi$ (the object that pairs entropic gradients with velocities in the HK metric) to the Gaussian manifold, using the general reduction formula for Onsager operators, which amounts to a Schur complement after Lagrange-multiplier elimination. The reduction is additive, $K^{\\mathrm{red}}_{\\alpha,\\beta}(p)=\\alpha K^{\\mathrm{red}}_{\\mathrm{tr}}(p)+\\beta K^{\\mathrm{red}}_{\\mathrm{He}}(p)$ as in (3.3), and inserting it into $\\dot p=-K^{\\mathrm{red}}DE(p)$ yields the explicit ODEs (3.6). The argument also uses the natural parametrization $\\rho(x)=\\exp(c+b\\cdot x-\\tfrac12 x\\cdot Ax)$, in which the flow equations become polynomial, and the evolution of the normalized covariance $B=\\Gamma^{-1/2}\\Sigma\\Gamma^{-1/2}$, whose eigenvalues obey the scalar equations $\\dot b_i=-(2\\alpha\\langle v_i|\\Gamma^{-1}v_i\\rangle+\\beta b_i)(b_i-1)$ obtained by the standard eigenvalue perturbation formula.","core_discovery":"The central claim is that the relative Boltzmann entropy flow in the $(\\alpha,\\beta)$-HK geometry preserves scaled Gaussians whenever the reference measure is a scaled Gaussian, and that the restricted flow is exactly the gradient flow of the reduced entropy under the reduced Onsager operator $K^{\\mathrm{red}}_{\\alpha,\\beta}(p)=\\alpha K^{\\mathrm{red}}_{\\mathrm{tr}}(p)+\\beta K^{\\mathrm{red}}_{\\mathrm{He}}(p)$. Consequently the ODEs (3.6) fully describe the infinite-dimensional flow on the Gaussian manifold, not merely a projected approximation. On normalized Gaussians the paper shows that global geodesic $\\lambda$-convexity holds only in the pure transport case ($\\beta=0$ with $\\lambda\\le\\alpha\\nu_{\\min}(\\Gamma^{-1})$), while a sublevel version of semi-convexity holds in general. For Gaussian targets it proves exponential decay of the relative Boltzmann entropy with explicit rates, the sharpest being $\\nu_{\\mathrm{cov}}=2\\alpha\\nu_{\\min}(\\Gamma^{-1})+\\beta$ for the covariance part and $\\nu_{\\mathrm{m}}=2\\alpha\\nu_{\\min}(\\Gamma^{-1})+2\\beta$ for the mean part. For non-Gaussian targets that are strongly log-$\\lambda$-concave, it claims exponential convergence with rate $\\lambda(2\\alpha+\\beta r_E^{-1})$ obtained from a slope comparison between the transport and Hellinger dissipations.","pith_inferences":["Beyond the paper, the same reduction should apply to any exponential family whose log-density is quadratic: if the target lies in the family, the HK–Boltzmann flow stays in it, and the finite-dimensional ODEs can be derived by the same algebraic reduction.","Beyond the paper, if the slope comparison behind Theorem 5.11 is indeed reversed, the exponential-decay conclusion may survive with a different rate; a natural repair is to work with $\\sigma_{\\min}^{-1}$ or a different sublevel bound, which would change the rate constant but not the qualitative behavior.","Beyond the paper, the explicit covariance formula in Remark 2.3 offers a sharp numerical test of the PL rates: computing the exact covariance at time $t$ and comparing the exponent with $\\nu_{\\mathrm{cov}}$ would show whether the prefactor in (5.9) is an artifact of the proof.","Beyond the paper, a promising extension is a variational (minimizing-movement) discretization of the reduced flow rather than the Euler splitting of Algorithm 1; the additive Onsager structure suggests exact substeps for transport and Hellinger would preserve Gaussianity at every iteration."],"forward_implications":["In the pure transport case ($\\beta=0$) and the pure Hellinger case ($\\alpha=0$) the covariance equation has closed-form solutions, so the HK flow's decay can be compared exactly with its two limiting geometries.","The additive form of the reduced Onsager operator means the HK flow interpolates between mass-preserving transport and pure mass growth without introducing coupling terms at the metric level, which is what makes the parameter-space analysis tractable.","For Gaussian targets the refined decay rates are independent of the initial energy level, so the long-time behavior is governed only by $\\alpha$, $\\beta$, and the smallest eigenvalue of $\\Gamma^{-1}$.","The discrete Algorithm 1 provides a practical Gaussian variational-inference scheme for non-Gaussian targets such as Bayesian logistic regression, alternating transport and Hellinger updates with Monte Carlo estimation; the experiments indicate transport steps dominate far from equilibrium while Hellinger steps converge faster near it."],"supporting_citations":[{"why":"Defines the Hellinger–Kantorovich distance and the Onsager operator $K_{\\alpha,\\beta}(\\rho)\\xi=-\\alpha\\,\\mathrm{div}(\\rho\\nabla\\xi)+\\beta\\rho\\xi$ that sets up the gradient system.","marker":"[LMS16, LMS18]"},{"why":"Supplies the general reduction formula for Onsager operators (a Schur complement after constraint elimination) used to build the reduced operator on the Gaussian manifold.","marker":"[MaM20]"},{"why":"Gives the shape–mass relation between the HK flow and the spherical HK flow, which justifies studying normalized Gaussians separately from mass.","marker":"[MiZ25]"},{"why":"Provides the quadratic-transport gradient flow over Gaussians and its ODEs, the limiting case the HK flow extends.","marker":"[LC∗22]"},{"why":"Provides the Hellinger (reaction-only) Gaussian flow and the energy-dissipation/PL method used for strongly log-concave targets.","marker":"[CH∗23]"},{"why":"Supplies the metric gradient-flow framework, displacement convexity, and the fact that quadratic-transport geodesics between Gaussians remain Gaussian.","marker":"[AGS05]"},{"why":"Gives the eigenvalue perturbation formula that turns the covariance evolution into scalar ODEs for the eigenvalues of $B=\\Gamma^{-1/2}\\Sigma\\Gamma^{-1/2}$.","marker":"[Hel15]"},{"why":"Supplies the Polyak–Lojasiewicz (gradient dominance) framework used to turn dissipation lower bounds into exponential decay.","marker":"[KNS20]"}],"fun_headline_variants":["Gaussians preserved: exact ODEs and sharp convergence rates","Exact Gaussian ODEs from HK-Boltzmann flow with exponential decay","Sharp rates for Gaussian-preserving HK-Boltzmann flow","Reduced HK-gradient flow: Gaussian ODEs with exponential rates","Gaussian invariance yields exact reduced dynamics and exponential rates"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise for the non-Gaussian decay theorem is the slope comparison $|\\partial_{\\mathrm{Otto}}E_1|^2 \\le \\sigma_{\\min}|\\partial_{\\mathrm{SHe}}E_1|^2$, which direct computation seems to reverse to $|\\partial_{\\mathrm{Otto}}E_1|^2 \\le \\sigma_{\\min}^{-1}|\\partial_{\\mathrm{SHe}}E_1|^2$, so the stated rate $\\lambda(2\\alpha+\\beta r_E^{-1})$ is unsupported unless that inequality is repaired.","fun_headline_variants_meta":{"raw":{"variants":["Gaussians preserved: exact ODEs and sharp convergence rates","Exact Gaussian ODEs from HK-Boltzmann flow with exponential decay","Sharp rates for Gaussian-preserving HK-Boltzmann flow","Reduced HK-gradient flow: Gaussian ODEs with exponential rates","Gaussian invariance yields exact reduced dynamics and exponential rates"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001194,"raw_usage":{"total_tokens":4980,"prompt_tokens":1058,"completion_tokens":3922,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":674,"completion_tokens_details":{"reasoning_tokens":3834}},"tokens_in":674,"tokens_out":3922,"duration_ms":25393,"temperature":1.0,"reasoning_tokens":3834,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:30:13.845923+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take $d=1$, target $\\pi=\\mathcal{N}(0,1)$, covariance $\\Sigma=s^2$ with $s<1$, and any mean $m\\ne0$. Write $g=\\int \\nabla V\\,G\\,dx$. The transport slope is $(1-s^2)^2/s^2+g^2$ and the Hellinger slope is $(1-s^2)^2+s^2g^2$. The inequality used in Theorem 5.11 would require $(1-s^2)^2/s^2+g^2 \\le s^2((1-s^2)^2+s^2g^2)$, which simplifies to $(1-s^4)((1-s^2)^2+s^2g^2)\\le0$, false for every $s<1$. This one-dimensional check falsifies the slope comparison and hence the stated decay rate.","supporting_citations":[],"review_version":1}