Pith. sign in

REVIEW 2 major objections 3 minor 1 cited by

The Stein-log-Sobolev inequality and the exponential rate of convergence for the continuous Stein variational gradient descent method

T0 review · 2 major / 3 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read A target-weighted kernel class satisfies the Stein-log-Sobolev inequality, giving the first rigorous exponential convergence for continuous Stein variational gradient descent.

desk verdict Real result: a first Stein-log-Sobolev inequality with exponential KL convergence for continuous SVGD, but Lemma 3.1 has a displayed factor-4π² typo that must be fixed before refereeing. read the letter →

arxiv 2412.10295 v1 pith:6JLYOAZP submitted 2024-12-13 math.AP cs.NAmath.NAmath.PRmath.STstat.TH

classification math.APcs.NAmath.NAmath.PRmath.STstat.TH MSC 35Q6235Q6835B4062-0862D05
keywords SteinvariationalgradientdescentStein-log-Sobolevinequalityexponentialconvergencemean-fieldlimitlog-SobolevFouriertransformSobolevspacesBayesianinference
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proves the first Stein-log-Sobolev inequality for any kernel: for a target Gibbs density whose potential lies above a Gaussian, and for any convolution kernel whose Fourier transform decays quadratically, multiplying the kernel by target-dependent exponential weights yields $\lambda\, KL(\rho||\rho_8) \le D^2_k(\rho||\rho_8)$. Because the mean-field SVGD equation dissipates Kullback-Leibler divergence at exactly the rate $D^2$, this yields weak solutions with exponential KL decay, $KL(\rho_t||\rho_8) \le e^{-\lambda t} KL(\rho_0||\rho_8)$, in every space dimension. The proof works by viewing the Stein-Fisher information as a duality pairing between $H^{-1}(\mathbb{R}^d)$ and $H^1(\mathbb{R}^d)$ and passing to Fourier variables, where the weighted ansatz cancels the potential. Negative examples show that kernels with faster Fourier decay, or naive polynomial weights, fail the inequality, so the assumptions are partially necessary.

What carries the argument

The central object is the weighted kernel ansatz $K(x,y)=e^{(V(x)-V_0(x))/2} k(x-y) e^{(V(y)-V_0(y))/2}$, where $V_0$ is the quadratic lower envelope of $V$. This form makes the target potential cancel when the Stein-Fisher information, also called the squared Stein discrepancy, is rewritten as a duality pairing between $H^{-1}(\mathbb{R}^d)$ and $H^1(\mathbb{R}^d)$ and passed to Fourier variables. The proof reconstructs admissible kernels from a chosen radial weight $q$ by solving the ODE $(4\pi^2 r^2 - d/2)\hat{k}(r) - (r/2)\hat{k}'(r) = q(r)$, then uses a Poincar\'e-Wirtinger inequality on a ball to compensate the regions where $q$ is negative, yielding the lower bound (3.18).

What would settle it

Take $d=1$, $V(x)=x^2/2$, $k(x)=e^{-|x|}$, and evaluate the inequality (B) for the Gaussian family $\rho_\sigma = N(0,\sigma^2)$ as $\sigma\to 0$; if the optimal constant $\lambda(\sigma):=\inf D^2(\rho_\sigma||\rho_8)/KL(\rho_\sigma||\rho_8)$ drops below the value in Remark 1.2, the stated rate is wrong. For $d\ge 2$, test the constructed kernel $k_{0,d}$ of (3.29) on Gaussian targets; if $D^2/KL$ can be made arbitrarily small, the theorem fails.

Watch

Extended reading notes

Core claim

For every dimension $d$, any target density $\rho_8 = Z^{-1}e^{-V}$ with $V \ge C + \tfrac12 (x-\mu)\cdot\Sigma^{-1}(x-\mu)$, and any $k \in L^1(\mathbb{R}^d)+L^2(\mathbb{R}^d)$ whose Fourier transform satisfies $(1.10)$ (quadratic decay and local boundedness away from zero and infinity), the kernel $K(x,y)=e^{(V(x)-V_0(x))/2} k(x-y) e^{(V(y)-V_0(y))/2}$ satisfies $\lambda\, KL(\rho||\rho_8) \le D^2_k(\rho||\rho_8)$ for all densities with the regularity (H), with an explicit constant $\lambda$. The paper then constructs weak solutions of the mean-field SVGD equation (MF SVGD) for these kernels and shows $KL(\rho_t||\rho_8) \le e^{-\lambda t} KL(\rho_0||\rho_8)$. This is the first proof of a Stein-log-Sobolev inequality for any kernel, and the failure conditions (Theorem 1.4) show that the quadratic-decay and exponential-weight assumptions are close to necessary.

Load-bearing premise

The proof relies on the ansatz $K(x,y)=e^{(V(x)-V_0(x))/2} k(x-y) e^{(V(y)-V_0(y))/2}$, which attaches the kernel to the target's quadratic lower envelope; if this form is not used, the Fourier cancellation that controls the Kullback-Leibler divergence is lost and the inequality is not proven.

Editorial extensions

If this is right

  • The mean-field SVGD flow with these kernels converges to the target in Kullback-Leibler divergence at rate $e^{-\lambda t}$, with the rate constant $\lambda$ given explicitly in Remark 1.2.
  • The admissible kernels must be adapted to the target through the quadratic envelope $V_0$; the common choice of a fixed translation-invariant kernel is outside the theorem's scope.
  • Kernels whose Fourier transform decays faster than quadratically fail the inequality for Gaussian targets, so the quadratic-decay condition in (1.10) is nearly necessary.
  • In one dimension the classical Mat\'ern kernel $e^{-|x|}$ is admissible; in higher dimensions the paper constructs explicit admissible kernels whose frequency profile is given in (3.29).
  • The proof also yields an $L^2$ and $H^{-1}$ bound on $\rho e^{(V-V_0)/2}$ in terms of the dissipation, giving quantitative control beyond the KL decay.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A practical consequence the paper leaves implicit: to obtain exponential convergence one should design kernels from the quadratic lower envelope of the log-target rather than reusing a universal kernel, and the computational cost of this adaptation is a testable question for numerical inference.
  • The regularity gap between the working Fourier-decay range $s\in[0,1]$ and the failure regime $s>1+d/2$ suggests a sharp threshold; if it exists, it would characterise exactly which reproducing-kernel Hilbert spaces admit exponential SVGD convergence.
  • The $H^{-1}$--$H^1$ duality formulation may extend to finite-particle SVGD, potentially transferring these exponential rates to the discrete algorithm, although the paper does not address that.
  • The failure results for polynomial weights hint that only exponential target weights can save the inequality; verifying the conjecture stated in case (F2) would complete the picture.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 3 minor

Summary. The paper claims a resolution of the long-open question of exponential convergence for the continuous Stein variational gradient descent method. For target densities of the form ρ_∞ = e^{-V}/Z satisfying the quadratic lower bound (1.5), and for kernels of the ansatz (1.6) whose Fourier transform satisfies the two-sided quadratic bound (1.10), Theorem 1.1 proves the Stein-log-Sobolev inequality (SLSI) with an explicit constant λ. The proof rewrites the squared Stein discrepancy as an H^{-1}–H^1 duality pairing, passes to Fourier variables in Lemma 3.1, compensates the negative part of the weight q by a Poincaré–Wirtinger argument, and constructs explicit kernels k_{0,d} ∈ L^1+L^2. Theorem 1.3 then constructs weak solutions to the mean-field SVGD equation with exponential KL-decay, and Theorem 1.4 collects several conditions under which algebraic versions of the inequality fail. The paper includes full proofs in appendices and explicit dependencies for the constants.

Significance. If the technical issues are repaired, this is a substantial contribution: it provides the first proof of the Stein-log-Sobolev inequality for genuinely nontrivial kernels, with an explicit rate, and it supplies a rigorous weak-solution framework for the continuous SVGD equation. The Fourier-duality interpretation and the explicit kernel construction are valuable ideas that will likely be reused. The paper is also unusually careful about the constants: the Stein-log-Sobolev constant in Remark 1.2 is tracked through all dependencies, and the negative results in Theorem 1.4 give falsifiable, concrete limitations of the approach. The target-dependent kernel ansatz (1.6) is a real restriction, but the failure examples partially justify it, so this restriction is not an internal inconsistency.

major comments (2)
  1. [Lemma 3.1, Eq. (3.2), (3.5), (3.7)] There is a concrete factor-4π² error in the Fourier representation of the drift term. Under the paper's convention F(x_j f) = (i/(2π)) ∂_j f̂, the correct transform of (1/2)Σ^{-1}(x−μ)g_0 carries the coefficient i/(4π) after the change of variables, not (i/2)(2π) = iπ as displayed in (3.5) and (3.7). With the printed coefficient, the expansion of |(2πiξ)ĝ + iπ∇ĝ|² gives a gradient term π²|∇ĝ|² and a divergent cross term −4π²div(ξk̂), whereas the displayed equality in (3.2) uses the gradient coefficient 1/(16π²) and the q defined by q = 4π²|ξ|²k̂ − (1/2)div(ξk̂). Those two expressions are inconsistent, so the displayed chain (3.2) cannot be reproduced as written. Since Lemma 3.1 is the foundation of both Theorem 1.1 and Theorem 1.3, this is a load-bearing issue. The later 1D computation in Section 3.2 and the numerical check of (3.14) use the corrected coefficient, so the error appears to be a typo rather than a conceptual flaw, but it must be corrected and the proof of Lemma 3.1 rechecked.
  2. [Theorem 1.4, Case (F2)] The admissible range for β in the statement is incompatible with the proof. The statement claims failure for β ∈ [0, (2 − 1/(2r))d + 1), but the proof requires β < d + 1 − d/p with p = 2r/(2r−1), i.e. β < 1 + d/(2r). For d ≥ 2 the printed range contains values such as β ≈ 2d + 1 for which the proof gives no bound on the dissipation terms and for which the constructed density has infinite ∫ρV with V = |x|². As stated, Theorem 1.4(F2) therefore overclaims; either the upper bound must be corrected to 1 + d/(2r) or a proof covering the printed range must be supplied.
minor comments (3)
  1. [Section 3, first paragraph] The text says 'We will assume that k is radially symmetric', but Theorem 1.1 quantifies over all kernels satisfying (1.10). The later replacement argument only uses the two-sided bound (1.10), so radial symmetry appears unnecessary; if it is truly not needed, the sentence should be removed, and if it is needed, it should be stated in Theorem 1.1.
  2. [Eq. (3.2) and surrounding notation] After the change of variables, the notation confuses g_0 and g in several displayed formulas. Since the Fourier variable is renaming anyway, a brief note that g_0 = g ∘ z and that subsequent integrals are written in the η variable would improve readability.
  3. [General typography] There are several typographical slips, for example 'probabiliy', 'Lipshitz', and 'Schwarz space' for the Schwartz space; these should be cleaned up in revision.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the kernel-existence proof is self-contained and does not reduce the target inequality to its own assumptions.

full rationale

The derivation of Theorem 1.1 is self-contained. The sharp inequality (B) is proved by passing to Fourier variables, bounding the Stein-Fisher information from below via Lemma 3.1, and then compensating negative parts of the weight q with a Poincare-Wirtinger argument; no step assumes the target inequality (SLSI). The kernel ansatz (1.6) is an explicit hypothesis, not a fitted output, and the kernel k0,d is constructed from a chosen weight q through the ODE (3.20), with condition (3.17) checked directly. The centering parameter tau in (1.9) is a construction that enforces the zero-mean condition used in the Poincare-Wirtinger step, not a fitted parameter disguised as a prediction. The only self-citations are auxiliary technical results used in the compactness and bootstrapping parts of Theorem 1.3, and these are not load-bearing for the main SLSI proof. The apparent factor 4*pi^2 inconsistency in the display of Lemma 3.1 is a correctness/reproducibility issue, not a circular reduction: nothing in the argument defines the dissipation or the kernel in terms of the final KL bound, so the claimed inequality is not equivalent to its input by construction.

Assumptions & free parameters 3 free parameters · 7 assumptions · 0 invented entities

The construction introduces no new physical or probabilistic entities. The free parameters alpha, epsilon, and delta are internal to the kernel construction and do not represent fitted data. The axioms are standard mathematical tools and explicit domain restrictions of the theorem statements; the kernel ansatz is the strongest modeling assumption.

free parameters (3)
  • alpha = arbitrary positive, e.g. 1 (Section 3.5, eq. (3.27))
    Scales the piecewise-constant weight q; the kernel k0,d is proportional to alpha and the inequality constant lambda is proportional to alpha; any alpha>0 works, so it is a normalization choice, not fitted to data.
  • epsilon = chosen sufficiently small so that (3.17) holds; explicit optimal in (3.36) for d>=2
    Radius of the ball where q is negative; the proof requires epsilon small enough to satisfy the Poincare-Wirtinger condition (3.17); the final constant lambda0,d is optimized over epsilon. It is a construction parameter, not fitted to data.
  • delta (in 1D) = any small positive number such that (3.14) holds, then optimized in (3.15)
    Used in the 1D proof to enlarge the negative region to B_{epsilon+delta}; lambda0,1 is the supremum over delta; a proof artifact.
assumptions (7)
  • standard math Fourier transform and Plancherel formula extend to the H^{-s}/H^s duality pairing (Lemma B.1)
    Used throughout Section 3 to rewrite the dissipation in Fourier variables; proved in Appendix B using density arguments.
  • standard math Optimal Poincare-Wirtinger constant on balls is (R/p(d/2,1))^2 (Lemma D.3)
    Central to compensating negative values of q in Lemma 3.4; cited from [33,61].
  • standard math Bochner's theorem: nonnegative Fourier transform implies positive definiteness of k
    Used in Section 3.5 to guarantee that the constructed k0,d is a valid positive definite kernel; cited [53].
  • standard math Standard parabolic existence, Leray-Schauder fixed point, Schauder estimates, Aubin-Lions lemma (Appendix C)
    Used to construct weak solutions to the mean-field PDE via the regularized problem (4.2)-(4.4).
  • domain assumption Target potential V satisfies (1.5): V >= C + V0 with V0 quadratic and positive definite; when needed, V in H^m_loc, m>d/2, and KL(rho0||rho8) finite
    The theorem statements assume these conditions; the proofs require the Gaussian lower bound for the L2 controls and KL bounds.
  • domain assumption Kernel ansatz K(x,y)=e^{(V(x)-V0(x))/2} k(x-y) e^{(V(y)-V0(y))/2} with k in L1+L2 satisfying (1.10)
    The theorem is stated for kernels of this form; this is a modeling restriction, not derived from SVGD itself.
  • domain assumption The dissipation D2 and KL are finite for the densities considered; densities satisfy regularity (H)
    The inequality (B) is proven for rho satisfying (H), which is stated as a standing hypothesis.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Stein-log-Sobolev inequality and the exponential rate of convergence for the continuous Stein variational gradient descent method." pith.science (2026). https://pith.science/paper/6JLYOAZP

@misc{pith2026241210295,
  author       = {Pith},
  title        = {Pith review of: The Stein-log-Sobolev inequality and the exponential rate of convergence for the continuous Stein variational gradient descent method},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6JLYOAZP}},
  note         = {Machine review of arXiv:2412.10295}
}
abstract

The Stein Variational Gradient Descent method is a variational inference method in statistics that has recently received a lot of attention. The method provides a deterministic approximation of the target distribution, by introducing a nonlocal interaction with a kernel. Despite the significant interest, the exponential rate of convergence for the continuous method has remained an open problem, due to the difficulty of establishing the related so-called Stein-log-Sobolev inequality. Here, we prove that the inequality is satisfied for each space dimension and every kernel whose Fourier transform has a quadratic decay at infinity and is locally bounded away from zero and infinity. Moreover, we construct weak solutions to the related PDE satisfying exponential rate of decay towards the equilibrium. The main novelty in our approach is to interpret the Stein-Fisher information, also called the squared Stein discrepancy, as a duality pairing between $H^{-1}(\mathbb{R}^d)$ and $H^{1}(\mathbb{R}^d)$, which allows us to employ the Fourier transform. We also provide several examples of kernels for which the Stein-log-Sobolev inequality fails, partially showing the necessity of our assumptions.

Figures

Figures reproduced from arXiv: 2412.10295 by the authors.

Figure 1
Figure 1. The value α kˆ0,1pεq ´ 4πε pd{2,1 ¯2 (y-axis, log-scale) vs the dimension d (x-axis). The constraint (3.14) is only satisfied for d “ 1. Lemma 3.4 (sufficient conditions on ˆk and q). Let k and q be as in Lemma 3.1. Further￾more, assume that k and q are radially symmetric and that there exists constants α, β, ε, θ ą 0 such that qprq ě $ ’& ’% ´α if r ă ε, β if r ě ε for r P r0, 8q and ˆkprq ě θ for r P r0, εs. If th… view at source ↗
Figure 2
Figure 2. The shape of the kernel frequency ˆk0,d with α “ 1 and ε “ 0.05, 0.1 for different dimensions d. On the x axis we have the radius r, and on the y axis we have the value ˆk0,dprq. Contrary to the kernel which satisfies (B) in one dimension ˆk0,dpξq “ 1 1`C |ξ| 2 , the ones constructed for higher dimensions first increase and then decrease to 0 at infinity. We conclude by following the same steps as in the proof of Th… view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Riesz-Kernel Stein Variational Gradient Descent: Renormalized Entropy and Long-Time Particle Limits

    math.AP 2026-07 conditional novelty 7.0 of 10

    For singular Riesz-kernel SVGD with self-interaction removed, the time-averaged empirical measure converges weakly to the target as particle number and averaging horizon diverge, with an explicit N^{-1+σ/d} correction...

Reference graph

Works this paper leans on

66 extracted references · 56 canonical work pages · cited by 1 Pith paper

  1. [1]

    Alsup, T

    T. Alsup, T. Hartland, B. Peherstorfer, and N. Petra. Fur ther analysis of multilevel Stein variational gradient descent with an application to the Bayesian infere nce of glacier ice models. Adv. Comput. Math., 50(4):Paper No. 65, 29, 2024

  2. [2]

    Alsup, L

    T. Alsup, L. Venturi, and B. Peherstorfer. Multilevel St ein variational gradient descent with applications to bayesian inverse problems. Proceedings of the 2nd Mathematical and Scientific Machine L earning Conference, 145:93–117, 2022

  3. [3]

    Ambrosio, N

    L. Ambrosio, N. Fusco, and D. Pallara. Functions of bounded variation and free discontinuity prob lems. Oxford Mathematical Monographs. The Clarendon Press, Oxfo rd University Press, New York, 2000

  4. [4]

    Ambrosio, N

    L. Ambrosio, N. Gigli, and G. Savaré. Gradient flows in metric spaces and in the space of probabilit y measures. Lectures in Mathematics ETH Zürich. Birkhäuser Verlag, Ba sel, second edition, 2008

  5. [5]

    Aronszajn and K

    N. Aronszajn and K. T. Smith. Theory of Bessel potentials . I. Ann. Inst. Fourier (Grenoble), 11:385–475, 1961

  6. [6]

    Balasubramanian, S

    K. Balasubramanian, S. Banerjee, and P. Ghosal. Improve d finite-particle convergence rates for Stein variational gradient descent. arXiv:2409.08469v2, 2024

  7. [7]

    Benilan, H

    P. Benilan, H. Brezis, and M. G. Crandall. A semilinear eq uation in L1pRN q. Ann. Scuola Norm. Sup. Pisa Cl. Sci. (4) , 2(4):523–555, 1975

  8. [8]

    J. A. Carrillo, A. Esposito, J. Skrzeczkowski, and J. S.- H. Wu. Nonlocal particle approximation for linear and fast diffusion equations. arXiv:2408.02345v1, 2024

Show all 66 references
  1. [9]

    J. A. Carrillo, Y. Salmaniw, and J. Skrzeczkowski. Well- posedness of aggregation-diffusion systems with irregular kernels. arXiv:2406.09227v1, 2024

  2. [10]

    J. A. Carrillo and J. Skrzeczkowski. Convergence and st ability results for the particle system in the Stein gradient descent method. arXiv preprint arXiv:2312.16344; to appear in Math. Comp. , 2023. 62 JOSÉ A. CARRILLO, JAKUB SKRZECZKOWSKI, AND JETHRO W ARNET T

  3. [11]

    Chen and O

    P. Chen and O. Ghattas. Projected Stein variational gra dient descent. Advances in Neural Information Processing Systems, 33:1947–1958, 2020

  4. [12]

    P. Chen, K. Wu, J. Chen, T. O'Leary-Roseberry, and O. Gha ttas. Projected Stein variational Newton: A fast and scalable Bayesian inference method in high dimens ions. In Advances in Neural Information Processing Systems, volume 32, 2019

  5. [13]

    Chewi, T

    S. Chewi, T. Le Gouic, C. Lu, T. Maunu, and P. Rigollet. SV GD as a kernelized Wasserstein gradient flow of the chi-squared divergence. In Advances in Neural Information Processing Systems , volume 33, pages 2098–2109, 2020

  6. [14]

    Detommaso, T

    G. Detommaso, T. Cui, Y. Marzouk, A. Spantini, and R. Sch eichl. A Stein variational Newton method. In Advances in Neural Information Processing Systems , volume 31, 2018

  7. [15]

    Doumic, S

    M. Doumic, S. Hecht, B. Perthame, and D. Peurichard. Mul tispecies cross-diffusions: from a nonlocal mean-field to a porous medium system without self-diffusion. Journal of Differential Equations , 389:228– 256, 2024

  8. [16]

    Duncan, N

    A. Duncan, N. Nüsken, and L. Szpruch. On the geometry of S tein variational gradient descent. J. Mach. Learn. Res., 24:Paper No. [56], 39, 2023

  9. [17]

    L. C. Evans. Weak convergence methods for nonlinear partial differentia l equations, volume 74 of CBMS Regional Conference Series in Mathematics . Published for the Conference Board of the Mathematical Sciences, Washington, DC; by the American Mathematical Soc iety, Providence,...

  10. [18]

    L. C. Evans. Partial differential equations , volume 19 of Graduate Studies in Mathematics . American Mathematical Society, Providence, RI, second edition, 201 0

  11. [19]

    G. B. Folland. Real analysis. Pure and Applied Mathematics (New York). John Wiley & Sons, Inc., New York, second edition, 1999. Modern techniques and their app lications, A Wiley-Interscience Publication

  12. [20]

    Gallego and D

    V. Gallego and D. R. Insua. Stochastic gradient MCMC wit h repulsive forces. arXiv:1812.00071v2, 2020

  13. [21]

    Geman and D

    S. Geman and D. Geman. Stochastic relaxation, Gibbs dis tributions, and the Bayesian restoration of images. IEEE Transactions on pattern analysis and machine intellig ence, (6):721–741, 1984

  14. [22]

    Gilbarg and N

    D. Gilbarg and N. S. Trudinger. Elliptic partial differential equations of second order . Classics in Math- ematics. Springer-Verlag, Berlin, 2001. Reprint of the 199 8 edition

  15. [23]

    Grafakos

    L. Grafakos. Modern Fourier analysis . Number 250 in Graduate texts in mathematics. Springer, New York, 2nd ed. edition, 2009

  16. [24]

    L. Gross. Logarithmic Sobolev inequalities. Amer. J. Math. , 97(4):1061–1083, 1975

  17. [25]

    Han and Q

    J. Han and Q. Liu. Stein variational gradient descent wi thout gradient. Proceedings of the 35th Inter- national Conference on Machine Learning , 80:1900–1908, 10–15 Jul 2018

  18. [26]

    W. K. Hastings. Monte Carlo sampling methods using Mark ov chains and their applications. Biometrika, 57(1):97–109, 1970. THE STEIN-LOG-SOBOLEV INEQUALITY 63

  19. [27]

    Y. He, K. Balasubramanian, B. K. Sriperumbudur, and J. L u. Regularized Stein variational gradient flow. Found. Comput. Math. (in press), preprint arXiv:2211.0786 1v2, 2024

  20. [28]

    Hieber and J

    M. Hieber and J. Prüss. Heat kernels and maximal Lp-Lq estimates for parabolic evolution equations. Comm. Partial Differential Equations , 22(9-10):1647–1669, 1997

  21. [29]

    M. D. Hoffman, D. M. Blei, C. Wang, and J. Paisley. Stochas tic variational inference. J. Mach. Learn. Res., 14:1303–1347, 2013

  22. [30]

    Jordan, D

    R. Jordan, D. Kinderlehrer, and F. Otto. The variationa l formulation of the Fokker-Planck equation. SIAM J. Math. Anal. , 29(1):1–17, 1998

  23. [31]

    Korba, P.-C

    A. Korba, P.-C. Aubin-Frankowski, S. Majewski, and P. A blin. Kernel Stein discrepancy descent. Pro- ceedings of the 38th International Conference on Machine Le arning, pages 5719–5730, 2021

  24. [32]

    Korba, A

    A. Korba, A. Salim, M. Arbel, G. Luise, and A. Gretton. A n on-asymptotic analysis for Stein variational gradient descent. Advances in Neural Information Processing Systems , 33:4672–4682, 2020

  25. [33]

    Kuznetsov and A

    N. Kuznetsov and A. Nazarov. Sharp constants in the Poin caré, Steklov and related inequalities (a survey). Mathematika, 61(2):328–344, 2015

  26. [34]

    L. Li, Y. Li, J.-G. Liu, Z. Liu, and J. Lu. A stochastic ver sion of Stein variational gradient descent for efficient sampling. Commun. Appl. Math. Comput. Sci. , 15(1):37–63, 2020

  27. [35]

    Lions and C

    P.-L. Lions and C. Villani. Régularité optimale de raci nes carrées. Comptes rendus de l’Académie des sciences. Série 1, Mathématique , 321(12):1537–1541, 1995

  28. [36]

    Q. Liu. Stein variational gradient descent as gradient flow. Advances in Neural Information Processing Systems, 30, 2017

  29. [37]

    Q. Liu, J. Lee, and M. Jordan. A Kernelized Stein Discrep ancy for Goodness-of-fit Tests. In Proceedings of The 33rd International Conference on Machine Learning , 48:276–284, 2016

  30. [38]

    Liu and D

    Q. Liu and D. Wang. Stein variational gradient descent: a general purpose Bayesian inference algorithm. Proc. 30th Int. Conf. Neural Inf. Proc. Syst. , page 2378–2386, 2016

  31. [39]

    Liu and D

    Q. Liu and D. Wang. Stein variational gradient descent a s moment matching. Advances in Neural Information Processing Systems , 31, 2018

  32. [40]

    T. Liu, P. Ghosal, K. Balasubramanian, and N. Pillai. To wards Understanding the Dynamics of Gaussian-Stein Variational Gradient Descent. Advances in Neural Information Processing Systems , 36, 2024

  33. [41]

    J. Lu, Y. Lu, and J. Nolen. Scaling limit of the Stein vari ational gradient descent: the mean field regime. SIAM J. Math. Anal. , 51(2):648–671, 2019

  34. [42]

    A. Lunardi. Analytic semigroups and optimal regularity in parabolic pr oblems. Progress in Nonlinear Differential Equations and their Applications, 16. Birkhäu ser Verlag, Basel, 1995

  35. [43]

    Metropolis, A

    N. Metropolis, A. W. Rosenbluth, M. N. Rosenbluth, A. H. Teller, and E. Teller. Equation of state calculations by fast computing machines. J. Chem. Phys. , 21(6):1087–1092, 1953. 64 JOSÉ A. CARRILLO, JAKUB SKRZECZKOWSKI, AND JETHRO W ARNET T

  36. [44]

    Nüsken and D

    N. Nüsken and D. R. M. Renger. Stein variational gradien t descent: many-particle and long-time asymptotics. Found. Data Sci. , 5(3):286–320, 2023

  37. [45]

    F. W. J. Olver. Asymptotics and special functions . AKP Classics. A K Peters, Ltd., Wellesley, MA,

  38. [46]

    F. Otto. The geometry of dissipative evolution equatio ns: the porous medium equation. Comm. Partial Differential Equations , 26(1-2):101–174, 2001

  39. [47]

    Otto and C

    F. Otto and C. Villani. Generalization of an inequality by Talagrand and links with the logarithmic Sobolev inequality. J. Funct. Anal. , 173(2):361–400, 2000

  40. [48]

    Otto and M

    F. Otto and M. Westdickenberg. Eulerian calculus for th e contraction in the Wasserstein distance. SIAM J. Math. Anal. , 37(4):1227–1255, 2005

  41. [49]

    Perthame and N

    B. Perthame and N. Vauchelet. Incompressible limit of a mechanical model of tumour growth with viscosity. Philos. Trans. Roy. Soc. A , 373(2050):20140283, 16, 2015

  42. [50]

    Priser, P

    V. Priser, P. Bianchi, and A. Salim. Long-time asymptot ics of noisy SVGD outside the population limit. arXiv preprint arXiv:2406.11929 , 2024

  43. [51]

    Ranganath, S

    R. Ranganath, S. Gerrish, and D. Blei. Black box variati onal inference. In Artificial intelligence and statistics, pages 814–822. PMLR, 2014

  44. [52]

    G. O. Roberts and J. S. Rosenthal. Optimal scaling of dis crete approximations to langevin diffusions. Journal of the Royal Statistical Society: Series B (Statist ical Methodology), 60(1):255–268, 1998

  45. [53]

    W. Rudin. Fourier analysis on groups , volume No. 12 of Interscience Tracts in Pure and Applied Mathematics. Interscience Publishers (a division of John Wiley & Sons, I nc.), New York-London, 1962

  46. [54]

    Saitoh and Y

    S. Saitoh and Y. Sawano. Theory of reproducing kernels and applications . Springer, 2016

  47. [55]

    Shi and L

    J. Shi and L. Mackey. A finite-particle convergence rate for Stein variational gradient descent. Advances in Neural Information Processing Systems , 36, 2024

  48. [56]

    G. Toscani. Entropy production and the rate of converge nce to equilibrium for the Fokker-Planck equation. Quart. Appl. Math. , 57(3):521–541, 1999

  49. [57]

    C. Villani. Topics in optimal transportation , volume 58 of Graduate Studies in Mathematics . American Mathematical Society, Providence, RI, 2003

  50. [58]

    C. Villani. Optimal transport, volume 338 of Grundlehren der mathematischen Wissenschaften . Springer- Verlag, Berlin, 2009. Old and new

  51. [59]

    D. Wang, Z. Tang, C. Bajaj, and Q. Liu. Stein variational gradient descent with matrix-valued kernels. Advances in neural information processing systems , 32, 2019

  52. [60]

    Y. Wang, P. Chen, and W. Li. Projected Wasserstein gradi ent descent for high-dimensional Bayesian inference. SIAM/ASA J. Uncertain. Quantif. , 10(4):1513–1532, 2022

  53. [61]

    H. F. Weinberger. An isoperimetric inequality for the N -dimensional free membrane problem. J. Rational Mech. Anal., 5:633–636, 1956. THE STEIN-LOG-SOBOLEV INEQUALITY 65

  54. [62]

    Welling and Y

    M. Welling and Y. W. Teh. Bayesian learning via stochast ic gradient Langevin dynamics. In Proceedings of the 28th international conference on machine learning , pages 681–688, 2011

  55. [63]

    Z. Wu, J. Yin, and C. Wang. Elliptic & parabolic equations . World Scientific Publishing Co. Pte. Ltd., Hackensack, NJ, 2006

  56. [64]

    Zhu and A

    J.-J. Zhu and A. Mielke. Kernel approximation of Fisher -Rao gradient flows. arXiv:2410.20622v1, 2024

  57. [65]

    J. Zhuo, C. Liu, J. Shi, J. Zhu, N. Chen, and B. Zhang. Mess age passing Stein variational gradient descent. In International Conference on Machine Learning , pages 6018–6027. PMLR, 2018. José A. Carrillo: Ma thema tical Institute, University of Oxford, Woodstock R oad, Oxford...

  58. [1997]

    Reprint of the 1974 original [Academic Press, New York ; MR0435697 (55 #8655)]

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.