REVIEW 2 major objections 3 minor 1 cited by
The Stein-log-Sobolev inequality and the exponential rate of convergence for the continuous Stein variational gradient descent method
T0 review · 2 major / 3 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A target-weighted kernel class satisfies the Stein-log-Sobolev inequality, giving the first rigorous exponential convergence for continuous Stein variational gradient descent.
desk verdict Real result: a first Stein-log-Sobolev inequality with exponential KL convergence for continuous SVGD, but Lemma 3.1 has a displayed factor-4π² typo that must be fixed before refereeing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the weighted kernel ansatz $K(x,y)=e^{(V(x)-V_0(x))/2} k(x-y) e^{(V(y)-V_0(y))/2}$, where $V_0$ is the quadratic lower envelope of $V$. This form makes the target potential cancel when the Stein-Fisher information, also called the squared Stein discrepancy, is rewritten as a duality pairing between $H^{-1}(\mathbb{R}^d)$ and $H^1(\mathbb{R}^d)$ and passed to Fourier variables. The proof reconstructs admissible kernels from a chosen radial weight $q$ by solving the ODE $(4\pi^2 r^2 - d/2)\hat{k}(r) - (r/2)\hat{k}'(r) = q(r)$, then uses a Poincar\'e-Wirtinger inequality on a ball to compensate the regions where $q$ is negative, yielding the lower bound (3.18).
What would settle it
Take $d=1$, $V(x)=x^2/2$, $k(x)=e^{-|x|}$, and evaluate the inequality (B) for the Gaussian family $\rho_\sigma = N(0,\sigma^2)$ as $\sigma\to 0$; if the optimal constant $\lambda(\sigma):=\inf D^2(\rho_\sigma||\rho_8)/KL(\rho_\sigma||\rho_8)$ drops below the value in Remark 1.2, the stated rate is wrong. For $d\ge 2$, test the constructed kernel $k_{0,d}$ of (3.29) on Gaussian targets; if $D^2/KL$ can be made arbitrarily small, the theorem fails.
Extended reading notes
Core claim
For every dimension $d$, any target density $\rho_8 = Z^{-1}e^{-V}$ with $V \ge C + \tfrac12 (x-\mu)\cdot\Sigma^{-1}(x-\mu)$, and any $k \in L^1(\mathbb{R}^d)+L^2(\mathbb{R}^d)$ whose Fourier transform satisfies $(1.10)$ (quadratic decay and local boundedness away from zero and infinity), the kernel $K(x,y)=e^{(V(x)-V_0(x))/2} k(x-y) e^{(V(y)-V_0(y))/2}$ satisfies $\lambda\, KL(\rho||\rho_8) \le D^2_k(\rho||\rho_8)$ for all densities with the regularity (H), with an explicit constant $\lambda$. The paper then constructs weak solutions of the mean-field SVGD equation (MF SVGD) for these kernels and shows $KL(\rho_t||\rho_8) \le e^{-\lambda t} KL(\rho_0||\rho_8)$. This is the first proof of a Stein-log-Sobolev inequality for any kernel, and the failure conditions (Theorem 1.4) show that the quadratic-decay and exponential-weight assumptions are close to necessary.
Load-bearing premise
The proof relies on the ansatz $K(x,y)=e^{(V(x)-V_0(x))/2} k(x-y) e^{(V(y)-V_0(y))/2}$, which attaches the kernel to the target's quadratic lower envelope; if this form is not used, the Fourier cancellation that controls the Kullback-Leibler divergence is lost and the inequality is not proven.
Editorial extensions
If this is right
- The mean-field SVGD flow with these kernels converges to the target in Kullback-Leibler divergence at rate $e^{-\lambda t}$, with the rate constant $\lambda$ given explicitly in Remark 1.2.
- The admissible kernels must be adapted to the target through the quadratic envelope $V_0$; the common choice of a fixed translation-invariant kernel is outside the theorem's scope.
- Kernels whose Fourier transform decays faster than quadratically fail the inequality for Gaussian targets, so the quadratic-decay condition in (1.10) is nearly necessary.
- In one dimension the classical Mat\'ern kernel $e^{-|x|}$ is admissible; in higher dimensions the paper constructs explicit admissible kernels whose frequency profile is given in (3.29).
- The proof also yields an $L^2$ and $H^{-1}$ bound on $\rho e^{(V-V_0)/2}$ in terms of the dissipation, giving quantitative control beyond the KL decay.
Reading between the lines
- A practical consequence the paper leaves implicit: to obtain exponential convergence one should design kernels from the quadratic lower envelope of the log-target rather than reusing a universal kernel, and the computational cost of this adaptation is a testable question for numerical inference.
- The regularity gap between the working Fourier-decay range $s\in[0,1]$ and the failure regime $s>1+d/2$ suggests a sharp threshold; if it exists, it would characterise exactly which reproducing-kernel Hilbert spaces admit exponential SVGD convergence.
- The $H^{-1}$--$H^1$ duality formulation may extend to finite-particle SVGD, potentially transferring these exponential rates to the discrete algorithm, although the paper does not address that.
- The failure results for polynomial weights hint that only exponential target weights can save the inequality; verifying the conjecture stated in case (F2) would complete the picture.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper claims a resolution of the long-open question of exponential convergence for the continuous Stein variational gradient descent method. For target densities of the form ρ_∞ = e^{-V}/Z satisfying the quadratic lower bound (1.5), and for kernels of the ansatz (1.6) whose Fourier transform satisfies the two-sided quadratic bound (1.10), Theorem 1.1 proves the Stein-log-Sobolev inequality (SLSI) with an explicit constant λ. The proof rewrites the squared Stein discrepancy as an H^{-1}–H^1 duality pairing, passes to Fourier variables in Lemma 3.1, compensates the negative part of the weight q by a Poincaré–Wirtinger argument, and constructs explicit kernels k_{0,d} ∈ L^1+L^2. Theorem 1.3 then constructs weak solutions to the mean-field SVGD equation with exponential KL-decay, and Theorem 1.4 collects several conditions under which algebraic versions of the inequality fail. The paper includes full proofs in appendices and explicit dependencies for the constants.
Significance. If the technical issues are repaired, this is a substantial contribution: it provides the first proof of the Stein-log-Sobolev inequality for genuinely nontrivial kernels, with an explicit rate, and it supplies a rigorous weak-solution framework for the continuous SVGD equation. The Fourier-duality interpretation and the explicit kernel construction are valuable ideas that will likely be reused. The paper is also unusually careful about the constants: the Stein-log-Sobolev constant in Remark 1.2 is tracked through all dependencies, and the negative results in Theorem 1.4 give falsifiable, concrete limitations of the approach. The target-dependent kernel ansatz (1.6) is a real restriction, but the failure examples partially justify it, so this restriction is not an internal inconsistency.
major comments (2)
- [Lemma 3.1, Eq. (3.2), (3.5), (3.7)] There is a concrete factor-4π² error in the Fourier representation of the drift term. Under the paper's convention F(x_j f) = (i/(2π)) ∂_j f̂, the correct transform of (1/2)Σ^{-1}(x−μ)g_0 carries the coefficient i/(4π) after the change of variables, not (i/2)(2π) = iπ as displayed in (3.5) and (3.7). With the printed coefficient, the expansion of |(2πiξ)ĝ + iπ∇ĝ|² gives a gradient term π²|∇ĝ|² and a divergent cross term −4π²div(ξk̂), whereas the displayed equality in (3.2) uses the gradient coefficient 1/(16π²) and the q defined by q = 4π²|ξ|²k̂ − (1/2)div(ξk̂). Those two expressions are inconsistent, so the displayed chain (3.2) cannot be reproduced as written. Since Lemma 3.1 is the foundation of both Theorem 1.1 and Theorem 1.3, this is a load-bearing issue. The later 1D computation in Section 3.2 and the numerical check of (3.14) use the corrected coefficient, so the error appears to be a typo rather than a conceptual flaw, but it must be corrected and the proof of Lemma 3.1 rechecked.
- [Theorem 1.4, Case (F2)] The admissible range for β in the statement is incompatible with the proof. The statement claims failure for β ∈ [0, (2 − 1/(2r))d + 1), but the proof requires β < d + 1 − d/p with p = 2r/(2r−1), i.e. β < 1 + d/(2r). For d ≥ 2 the printed range contains values such as β ≈ 2d + 1 for which the proof gives no bound on the dissipation terms and for which the constructed density has infinite ∫ρV with V = |x|². As stated, Theorem 1.4(F2) therefore overclaims; either the upper bound must be corrected to 1 + d/(2r) or a proof covering the printed range must be supplied.
minor comments (3)
- [Section 3, first paragraph] The text says 'We will assume that k is radially symmetric', but Theorem 1.1 quantifies over all kernels satisfying (1.10). The later replacement argument only uses the two-sided bound (1.10), so radial symmetry appears unnecessary; if it is truly not needed, the sentence should be removed, and if it is needed, it should be stated in Theorem 1.1.
- [Eq. (3.2) and surrounding notation] After the change of variables, the notation confuses g_0 and g in several displayed formulas. Since the Fourier variable is renaming anyway, a brief note that g_0 = g ∘ z and that subsequent integrals are written in the η variable would improve readability.
- [General typography] There are several typographical slips, for example 'probabiliy', 'Lipshitz', and 'Schwarz space' for the Schwartz space; these should be cleaned up in revision.
Circularity Check
No significant circularity: the kernel-existence proof is self-contained and does not reduce the target inequality to its own assumptions.
full rationale
The derivation of Theorem 1.1 is self-contained. The sharp inequality (B) is proved by passing to Fourier variables, bounding the Stein-Fisher information from below via Lemma 3.1, and then compensating negative parts of the weight q with a Poincare-Wirtinger argument; no step assumes the target inequality (SLSI). The kernel ansatz (1.6) is an explicit hypothesis, not a fitted output, and the kernel k0,d is constructed from a chosen weight q through the ODE (3.20), with condition (3.17) checked directly. The centering parameter tau in (1.9) is a construction that enforces the zero-mean condition used in the Poincare-Wirtinger step, not a fitted parameter disguised as a prediction. The only self-citations are auxiliary technical results used in the compactness and bootstrapping parts of Theorem 1.3, and these are not load-bearing for the main SLSI proof. The apparent factor 4*pi^2 inconsistency in the display of Lemma 3.1 is a correctness/reproducibility issue, not a circular reduction: nothing in the argument defines the dissipation or the kernel in terms of the final KL bound, so the claimed inequality is not equivalent to its input by construction.
Assumptions & free parameters
free parameters (3)
- alpha =
arbitrary positive, e.g. 1 (Section 3.5, eq. (3.27))
- epsilon =
chosen sufficiently small so that (3.17) holds; explicit optimal in (3.36) for d>=2
- delta (in 1D) =
any small positive number such that (3.14) holds, then optimized in (3.15)
assumptions (7)
- standard math Fourier transform and Plancherel formula extend to the H^{-s}/H^s duality pairing (Lemma B.1)
- standard math Optimal Poincare-Wirtinger constant on balls is (R/p(d/2,1))^2 (Lemma D.3)
- standard math Bochner's theorem: nonnegative Fourier transform implies positive definiteness of k
- standard math Standard parabolic existence, Leray-Schauder fixed point, Schauder estimates, Aubin-Lions lemma (Appendix C)
- domain assumption Target potential V satisfies (1.5): V >= C + V0 with V0 quadratic and positive definite; when needed, V in H^m_loc, m>d/2, and KL(rho0||rho8) finite
- domain assumption Kernel ansatz K(x,y)=e^{(V(x)-V0(x))/2} k(x-y) e^{(V(y)-V0(y))/2} with k in L1+L2 satisfying (1.10)
- domain assumption The dissipation D2 and KL are finite for the densities considered; densities satisfy regularity (H)
Cite this review
Pith. "Pith review of The Stein-log-Sobolev inequality and the exponential rate of convergence for the continuous Stein variational gradient descent method." pith.science (2026). https://pith.science/paper/6JLYOAZP
@misc{pith2026241210295,
author = {Pith},
title = {Pith review of: The Stein-log-Sobolev inequality and the exponential rate of convergence for the continuous Stein variational gradient descent method},
year = {2026},
howpublished = {\url{https://pith.science/paper/6JLYOAZP}},
note = {Machine review of arXiv:2412.10295}
}
abstract
The Stein Variational Gradient Descent method is a variational inference method in statistics that has recently received a lot of attention. The method provides a deterministic approximation of the target distribution, by introducing a nonlocal interaction with a kernel. Despite the significant interest, the exponential rate of convergence for the continuous method has remained an open problem, due to the difficulty of establishing the related so-called Stein-log-Sobolev inequality. Here, we prove that the inequality is satisfied for each space dimension and every kernel whose Fourier transform has a quadratic decay at infinity and is locally bounded away from zero and infinity. Moreover, we construct weak solutions to the related PDE satisfying exponential rate of decay towards the equilibrium. The main novelty in our approach is to interpret the Stein-Fisher information, also called the squared Stein discrepancy, as a duality pairing between $H^{-1}(\mathbb{R}^d)$ and $H^{1}(\mathbb{R}^d)$, which allows us to employ the Fourier transform. We also provide several examples of kernels for which the Stein-log-Sobolev inequality fails, partially showing the necessity of our assumptions.
Figures
Forward citations
Cited by 1 Pith paper
-
Riesz-Kernel Stein Variational Gradient Descent: Renormalized Entropy and Long-Time Particle Limits
For singular Riesz-kernel SVGD with self-interaction removed, the time-averaged empirical measure converges weakly to the target as particle number and averaging horizon diverge, with an explicit N^{-1+σ/d} correction...
Reference graph
Works this paper leans on
- [1]
- [2]
-
[3]
Ambrosio, N
L. Ambrosio, N. Fusco, and D. Pallara. Functions of bounded variation and free discontinuity prob lems. Oxford Mathematical Monographs. The Clarendon Press, Oxfo rd University Press, New York, 2000
2000
-
[4]
L. Ambrosio, N. Gigli, and G. Savaré. Gradient flows in metric spaces and in the space of probabilit y measures. Lectures in Mathematics ETH Zürich. Birkhäuser Verlag, Ba sel, second edition, 2008
work page 2008
-
[5]
N. Aronszajn and K. T. Smith. Theory of Bessel potentials . I. Ann. Inst. Fourier (Grenoble), 11:385–475, 1961
work page 1961
-
[6]
K. Balasubramanian, S. Banerjee, and P. Ghosal. Improve d finite-particle convergence rates for Stein variational gradient descent. arXiv:2409.08469v2, 2024
arXiv 2024
-
[7]
P. Benilan, H. Brezis, and M. G. Crandall. A semilinear eq uation in L1pRN q. Ann. Scuola Norm. Sup. Pisa Cl. Sci. (4) , 2(4):523–555, 1975
work page 1975
-
[8]
J. A. Carrillo, A. Esposito, J. Skrzeczkowski, and J. S.- H. Wu. Nonlocal particle approximation for linear and fast diffusion equations. arXiv:2408.02345v1, 2024
arXiv 2024
Show all 66 references
-
[9]
J. A. Carrillo, Y. Salmaniw, and J. Skrzeczkowski. Well- posedness of aggregation-diffusion systems with irregular kernels. arXiv:2406.09227v1, 2024
2024
-
[10]
J. A. Carrillo and J. Skrzeczkowski. Convergence and st ability results for the particle system in the Stein gradient descent method. arXiv preprint arXiv:2312.16344; to appear in Math. Comp. , 2023. 62 JOSÉ A. CARRILLO, JAKUB SKRZECZKOWSKI, AND JETHRO W ARNET T
2023 arXiv
-
[11]
Chen and O
P. Chen and O. Ghattas. Projected Stein variational gra dient descent. Advances in Neural Information Processing Systems, 33:1947–1958, 2020
1947
-
[12]
P. Chen, K. Wu, J. Chen, T. O'Leary-Roseberry, and O. Gha ttas. Projected Stein variational Newton: A fast and scalable Bayesian inference method in high dimens ions. In Advances in Neural Information Processing Systems, volume 32, 2019
2019
-
[13]
Chewi, T
S. Chewi, T. Le Gouic, C. Lu, T. Maunu, and P. Rigollet. SV GD as a kernelized Wasserstein gradient flow of the chi-squared divergence. In Advances in Neural Information Processing Systems , volume 33, pages 2098–2109, 2020
2020
-
[14]
Detommaso, T
G. Detommaso, T. Cui, Y. Marzouk, A. Spantini, and R. Sch eichl. A Stein variational Newton method. In Advances in Neural Information Processing Systems , volume 31, 2018
2018
-
[15]
Doumic, S
M. Doumic, S. Hecht, B. Perthame, and D. Peurichard. Mul tispecies cross-diffusions: from a nonlocal mean-field to a porous medium system without self-diffusion. Journal of Differential Equations , 389:228– 256, 2024
2024
-
[16]
Duncan, N
A. Duncan, N. Nüsken, and L. Szpruch. On the geometry of S tein variational gradient descent. J. Mach. Learn. Res., 24:Paper No. [56], 39, 2023
2023
-
[17]
L. C. Evans. Weak convergence methods for nonlinear partial differentia l equations, volume 74 of CBMS Regional Conference Series in Mathematics . Published for the Conference Board of the Mathematical Sciences, Washington, DC; by the American Mathematical Soc iety, Providence,...
1990
-
[18]
L. C. Evans. Partial differential equations , volume 19 of Graduate Studies in Mathematics . American Mathematical Society, Providence, RI, second edition, 201 0
-
[19]
G. B. Folland. Real analysis. Pure and Applied Mathematics (New York). John Wiley & Sons, Inc., New York, second edition, 1999. Modern techniques and their app lications, A Wiley-Interscience Publication
1999
-
[20]
Gallego and D
V. Gallego and D. R. Insua. Stochastic gradient MCMC wit h repulsive forces. arXiv:1812.00071v2, 2020
2020 arXiv
-
[21]
Geman and D
S. Geman and D. Geman. Stochastic relaxation, Gibbs dis tributions, and the Bayesian restoration of images. IEEE Transactions on pattern analysis and machine intellig ence, (6):721–741, 1984
1984
-
[22]
Gilbarg and N
D. Gilbarg and N. S. Trudinger. Elliptic partial differential equations of second order . Classics in Math- ematics. Springer-Verlag, Berlin, 2001. Reprint of the 199 8 edition
2001
-
[23]
Grafakos
L. Grafakos. Modern Fourier analysis . Number 250 in Graduate texts in mathematics. Springer, New York, 2nd ed. edition, 2009
2009
-
[24]
L. Gross. Logarithmic Sobolev inequalities. Amer. J. Math. , 97(4):1061–1083, 1975
1975
-
[25]
Han and Q
J. Han and Q. Liu. Stein variational gradient descent wi thout gradient. Proceedings of the 35th Inter- national Conference on Machine Learning , 80:1900–1908, 10–15 Jul 2018
1900
-
[26]
W. K. Hastings. Monte Carlo sampling methods using Mark ov chains and their applications. Biometrika, 57(1):97–109, 1970. THE STEIN-LOG-SOBOLEV INEQUALITY 63
1970
-
[27]
Y. He, K. Balasubramanian, B. K. Sriperumbudur, and J. L u. Regularized Stein variational gradient flow. Found. Comput. Math. (in press), preprint arXiv:2211.0786 1v2, 2024
2024
-
[28]
Hieber and J
M. Hieber and J. Prüss. Heat kernels and maximal Lp-Lq estimates for parabolic evolution equations. Comm. Partial Differential Equations , 22(9-10):1647–1669, 1997
1997
-
[29]
M. D. Hoffman, D. M. Blei, C. Wang, and J. Paisley. Stochas tic variational inference. J. Mach. Learn. Res., 14:1303–1347, 2013
2013
-
[30]
Jordan, D
R. Jordan, D. Kinderlehrer, and F. Otto. The variationa l formulation of the Fokker-Planck equation. SIAM J. Math. Anal. , 29(1):1–17, 1998
1998
-
[31]
Korba, P.-C
A. Korba, P.-C. Aubin-Frankowski, S. Majewski, and P. A blin. Kernel Stein discrepancy descent. Pro- ceedings of the 38th International Conference on Machine Le arning, pages 5719–5730, 2021
2021
-
[32]
Korba, A
A. Korba, A. Salim, M. Arbel, G. Luise, and A. Gretton. A n on-asymptotic analysis for Stein variational gradient descent. Advances in Neural Information Processing Systems , 33:4672–4682, 2020
2020
-
[33]
Kuznetsov and A
N. Kuznetsov and A. Nazarov. Sharp constants in the Poin caré, Steklov and related inequalities (a survey). Mathematika, 61(2):328–344, 2015
2015
-
[34]
L. Li, Y. Li, J.-G. Liu, Z. Liu, and J. Lu. A stochastic ver sion of Stein variational gradient descent for efficient sampling. Commun. Appl. Math. Comput. Sci. , 15(1):37–63, 2020
2020
-
[35]
Lions and C
P.-L. Lions and C. Villani. Régularité optimale de raci nes carrées. Comptes rendus de l’Académie des sciences. Série 1, Mathématique , 321(12):1537–1541, 1995
1995
-
[36]
Q. Liu. Stein variational gradient descent as gradient flow. Advances in Neural Information Processing Systems, 30, 2017
2017
-
[37]
Q. Liu, J. Lee, and M. Jordan. A Kernelized Stein Discrep ancy for Goodness-of-fit Tests. In Proceedings of The 33rd International Conference on Machine Learning , 48:276–284, 2016
2016
-
[38]
Liu and D
Q. Liu and D. Wang. Stein variational gradient descent: a general purpose Bayesian inference algorithm. Proc. 30th Int. Conf. Neural Inf. Proc. Syst. , page 2378–2386, 2016
2016
-
[39]
Liu and D
Q. Liu and D. Wang. Stein variational gradient descent a s moment matching. Advances in Neural Information Processing Systems , 31, 2018
2018
-
[40]
T. Liu, P. Ghosal, K. Balasubramanian, and N. Pillai. To wards Understanding the Dynamics of Gaussian-Stein Variational Gradient Descent. Advances in Neural Information Processing Systems , 36, 2024
2024
-
[41]
J. Lu, Y. Lu, and J. Nolen. Scaling limit of the Stein vari ational gradient descent: the mean field regime. SIAM J. Math. Anal. , 51(2):648–671, 2019
2019
-
[42]
A. Lunardi. Analytic semigroups and optimal regularity in parabolic pr oblems. Progress in Nonlinear Differential Equations and their Applications, 16. Birkhäu ser Verlag, Basel, 1995
1995
-
[43]
Metropolis, A
N. Metropolis, A. W. Rosenbluth, M. N. Rosenbluth, A. H. Teller, and E. Teller. Equation of state calculations by fast computing machines. J. Chem. Phys. , 21(6):1087–1092, 1953. 64 JOSÉ A. CARRILLO, JAKUB SKRZECZKOWSKI, AND JETHRO W ARNET T
1953
-
[44]
Nüsken and D
N. Nüsken and D. R. M. Renger. Stein variational gradien t descent: many-particle and long-time asymptotics. Found. Data Sci. , 5(3):286–320, 2023
2023
-
[45]
F. W. J. Olver. Asymptotics and special functions . AKP Classics. A K Peters, Ltd., Wellesley, MA,
-
[46]
F. Otto. The geometry of dissipative evolution equatio ns: the porous medium equation. Comm. Partial Differential Equations , 26(1-2):101–174, 2001
2001
-
[47]
Otto and C
F. Otto and C. Villani. Generalization of an inequality by Talagrand and links with the logarithmic Sobolev inequality. J. Funct. Anal. , 173(2):361–400, 2000
2000
-
[48]
Otto and M
F. Otto and M. Westdickenberg. Eulerian calculus for th e contraction in the Wasserstein distance. SIAM J. Math. Anal. , 37(4):1227–1255, 2005
2005
-
[49]
Perthame and N
B. Perthame and N. Vauchelet. Incompressible limit of a mechanical model of tumour growth with viscosity. Philos. Trans. Roy. Soc. A , 373(2050):20140283, 16, 2015
2015
-
[50]
Priser, P
V. Priser, P. Bianchi, and A. Salim. Long-time asymptot ics of noisy SVGD outside the population limit. arXiv preprint arXiv:2406.11929 , 2024
2024
-
[51]
Ranganath, S
R. Ranganath, S. Gerrish, and D. Blei. Black box variati onal inference. In Artificial intelligence and statistics, pages 814–822. PMLR, 2014
2014
-
[52]
G. O. Roberts and J. S. Rosenthal. Optimal scaling of dis crete approximations to langevin diffusions. Journal of the Royal Statistical Society: Series B (Statist ical Methodology), 60(1):255–268, 1998
1998
-
[53]
W. Rudin. Fourier analysis on groups , volume No. 12 of Interscience Tracts in Pure and Applied Mathematics. Interscience Publishers (a division of John Wiley & Sons, I nc.), New York-London, 1962
1962
-
[54]
Saitoh and Y
S. Saitoh and Y. Sawano. Theory of reproducing kernels and applications . Springer, 2016
2016
-
[55]
Shi and L
J. Shi and L. Mackey. A finite-particle convergence rate for Stein variational gradient descent. Advances in Neural Information Processing Systems , 36, 2024
2024
-
[56]
G. Toscani. Entropy production and the rate of converge nce to equilibrium for the Fokker-Planck equation. Quart. Appl. Math. , 57(3):521–541, 1999
1999
-
[57]
C. Villani. Topics in optimal transportation , volume 58 of Graduate Studies in Mathematics . American Mathematical Society, Providence, RI, 2003
2003
-
[58]
C. Villani. Optimal transport, volume 338 of Grundlehren der mathematischen Wissenschaften . Springer- Verlag, Berlin, 2009. Old and new
2009
-
[59]
D. Wang, Z. Tang, C. Bajaj, and Q. Liu. Stein variational gradient descent with matrix-valued kernels. Advances in neural information processing systems , 32, 2019
2019
-
[60]
Y. Wang, P. Chen, and W. Li. Projected Wasserstein gradi ent descent for high-dimensional Bayesian inference. SIAM/ASA J. Uncertain. Quantif. , 10(4):1513–1532, 2022
2022
-
[61]
H. F. Weinberger. An isoperimetric inequality for the N -dimensional free membrane problem. J. Rational Mech. Anal., 5:633–636, 1956. THE STEIN-LOG-SOBOLEV INEQUALITY 65
1956
-
[62]
Welling and Y
M. Welling and Y. W. Teh. Bayesian learning via stochast ic gradient Langevin dynamics. In Proceedings of the 28th international conference on machine learning , pages 681–688, 2011
2011
-
[63]
Z. Wu, J. Yin, and C. Wang. Elliptic & parabolic equations . World Scientific Publishing Co. Pte. Ltd., Hackensack, NJ, 2006
2006
-
[64]
Zhu and A
J.-J. Zhu and A. Mielke. Kernel approximation of Fisher -Rao gradient flows. arXiv:2410.20622v1, 2024
2024 arXiv
-
[65]
J. Zhuo, C. Liu, J. Shi, J. Zhu, N. Chen, and B. Zhang. Mess age passing Stein variational gradient descent. In International Conference on Machine Learning , pages 6018–6027. PMLR, 2018. José A. Carrillo: Ma thema tical Institute, University of Oxford, Woodstock R oad, Oxford...
2018
-
[1997]
Reprint of the 1974 original [Academic Press, New York ; MR0435697 (55 #8655)]
1974
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.