REVIEW 2 cited by
Efficient Gradient Flows in Sliced-Wasserstein Space
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Minimizing functionals in the space of probability distributions can be done with Wasserstein gradient flows. To solve them numerically, a possible approach is to rely on the Jordan-Kinderlehrer-Otto (JKO) scheme which is analogous to the proximal scheme in Euclidean spaces. However, it requires solving a nested optimization problem at each iteration, and is known for its computational challenges, especially in high dimension. To alleviate it, very recent works propose to approximate the JKO scheme leveraging Brenier's theorem, and using gradients of Input Convex Neural Networks to parameterize the density (JKO-ICNN). However, this method comes with a high computational cost and stability issues. Instead, this work proposes to use gradient flows in the space of probability measures endowed with the sliced-Wasserstein (SW) distance. We argue that this method is more flexible than JKO-ICNN, since SW enjoys a closed-form differentiable approximation. Thus, the density at each step can be parameterized by any generative model which alleviates the computational burden and makes it tractable in higher dimensions.
Forward citations
Cited by 2 Pith papers
-
Distributional Determinantal Point Process for Repulsive Clustering of Distributions
A dDPP prior built on a sliced Wasserstein kernel provides a repulsive distribution-valued point process that, in a generalized Bayesian mixture model, clusters distributions into better-separated groups than a Dirich...
-
Understanding Learning with Sliced-Wasserstein Requires Rethinking Informative Slices
Under a low-dimensional subspace assumption, rescaling informative slices of the sliced-Wasserstein distance reduces to one global constant, so the classical SWD with a tuned learning rate is competitive with speciali...
Discussion (0). Continue with ORCID to comment.