REVIEW 3 cited by
Wasserstein Gradient Flows for Moreau Envelopes of f-Divergences in Reproducing Kernel Hilbert Spaces
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
abstract
Commonly used $f$-divergences of measures, e.g., the Kullback-Leibler divergence, are subject to limitations regarding the support of the involved measures. A remedy is regularizing the $f$-divergence by a squared maximum mean discrepancy (MMD) associated with a characteristic kernel $K$. We use the kernel mean embedding to show that this regularization can be rewritten as the Moreau envelope of some function on the associated reproducing kernel Hilbert space. Then, we exploit well-known results on Moreau envelopes in Hilbert spaces to analyze the MMD-regularized $f$-divergences, particularly their gradients. Subsequently, we use our findings to analyze Wasserstein gradient flows of MMD-regularized $f$-divergences. We provide proof-of-the-concept numerical examples for flows starting from empirical measures. Here, we cover $f$-divergences with infinite and finite recession constants. Lastly, we extend our results to the tight variational formulation of $f$-divergences and numerically compare the resulting flows.
Forward citations
Cited by 3 Pith papers
-
Ultra-fast feature learning for the training of two-layer neural networks in the two-timescale regime
In the teacher-student setting, variable-projection training of two-layer networks is shown to match a weighted ultra-fast diffusion in the zero-regularization limit, giving linear convergence of the learned feature d...
-
Kernel Trace Distance: Quantum Statistical Metric between Measures through RKHS Density Operators
A new integral probability metric, the kernel trace distance, compares distributions via the Schatten 1-norm of kernel covariance operators and admits dimension-free sample rates.
-
Flowing Datasets with Wasserstein over Wasserstein Gradient Flows
A new gradient flow framework on the space of probability distributions over probability distributions is introduced and applied to flowing labeled datasets between domains.
Discussion (0). Continue with ORCID to comment.