Pith. sign in

REVIEW 2 cited by

Stochastic Variational Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1809.04855 v1 pith:ELNVAOEX submitted 2018-09-13 stat.ML cs.LG

classification stat.MLcs.LG
keywords optimizationvariationalapproachesderivativesdifferentiabledirectionalgaussiangradient
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Variational Optimization forms a differentiable upper bound on an objective. We show that approaches such as Natural Evolution Strategies and Gaussian Perturbation, are special cases of Variational Optimization in which the expectations are approximated by Gaussian sampling. These approaches are of particular interest because they are parallelizable. We calculate the approximate bias and variance of the corresponding gradient estimators and demonstrate that using antithetic sampling or a baseline is crucial to mitigate their problems. We contrast these methods with an alternative parallelizable method, namely Directional Derivatives. We conclude that, for differentiable objectives, using Directional Derivatives is preferable to using Variational Optimization to perform parallel Stochastic Gradient Descent.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Neural Image Compression and Explanation

    cs.CV 2019-08 conditional novelty 6.0 of 10

    NICE trains a stochastic binary mask that marks decision-relevant pixels and turns the rest into a low-resolution background, giving both an explanation and about 1.6x PNG compression with a small accuracy drop.

  2. Neural Plasticity Networks

    cs.NE 2019-08 reject novelty 5.0 of 10

    A single parameter k in a binary-gate network training method interpolates between dropout, standard training, and sparse or expanded architectures.

Pith tools