Pith. sign in

REVIEW 3 cited by

Diagnosing and Enhancing VAE Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1903.05789 v2 pith:LWLHYHCR submitted 2019-03-14 cs.LG cs.CVstat.ML

classification cs.LGcs.CVstat.ML
keywords actuallymodelmodelssamplesvaesadditionalalthoughanalyze
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Although variational autoencoders (VAEs) represent a widely influential deep generative model, many aspects of the underlying energy function remain poorly understood. In particular, it is commonly believed that Gaussian encoder/decoder assumptions reduce the effectiveness of VAEs in generating realistic samples. In this regard, we rigorously analyze the VAE objective, differentiating situations where this belief is and is not actually true. We then leverage the corresponding insights to develop a simple VAE enhancement that requires no additional hyperparameters or sensitive tuning. Quantitatively, this proposal produces crisp samples and stable FID scores that are actually competitive with a variety of GAN models, all while retaining desirable attributes of the original VAE architecture. A shorter version of this work will appear in the ICLR 2019 conference proceedings (Dai and Wipf, 2019). The code for our model is available at https://github.com/daib13/ TwoStageVAE.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Latent Space Consistency for Sparse-View CT Reconstruction

    eess.IV 2025-07 reject novelty 6.0 of 10

    CLS-DM adds a contrastive-learning alignment stage and a reconstruction constraint to a latent diffusion model for sparse-view 3D CT reconstruction.

  2. TokBench: Evaluating Your Visual Tokenizer before Visual Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    TokBench measures text recognition accuracy and face similarity on reconstructed images and videos across 16 tokenizers, showing small text and faces are poorly preserved and often missed by traditional metrics.

  3. Machine-Learning-Assisted Photonic Device Development: A Multiscale Approach from Theory to Characterization

    physics.optics 2025-06 accept novelty 4.0 of 10

    This review organizes machine-learning-assisted photonic device development into a five-step Bayesian framework spanning theory, simulation, design, fabrication, and characterization.

Pith tools