Pith. sign in

Pathwise Derivatives Beyond the Reparameterization Trick

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

We observe that gradients computed via the reparameterization trick are in direct correspondence with solutions of the transport equation in the formalism of optimal transport. We use this perspective to compute (approximate) pathwise gradients for probability distributions not directly amenable to the reparameterization trick: Gamma, Beta, and Dirichlet. We further observe that when the reparameterization trick is applied to the Cholesky-factorized multivariate Normal distribution, the resulting gradients are suboptimal in the sense of optimal transport. We derive the optimal gradients and show that they have reduced variance in a Gaussian Process regression task. We demonstrate with a variety of synthetic experiments and stochastic variational inference tasks that our pathwise gradients are competitive with other methods.

fields

cs.LG 1

years

2019 1

verdicts

REJECT 1

representative citing papers

PAC-Bayes with Backprop

cs.LG · 2019-08-19 · reject · novelty 5.0

Training neural networks with PAC-Bayes objectives yields MNIST test error of 1.4% and a non-vacuous risk bound of 2.3%, much tighter than prior PAC-Bayes certificates.

citing papers explorer

Showing 1 of 1 citing paper.

  • PAC-Bayes with Backprop cs.LG · 2019-08-19 · reject · none · ref 10 · internal anchor

    Training neural networks with PAC-Bayes objectives yields MNIST test error of 1.4% and a non-vacuous risk bound of 2.3%, much tighter than prior PAC-Bayes certificates.