Pith. sign in

REVIEW 2 cited by

Explicitly Minimizing the Blur Error of Variational Autoencoders

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.05939 v1 pith:FIJQXLTN submitted 2023-04-12 cs.CV cs.LGeess.IV

classification cs.CVcs.LGeess.IV
keywords reconstructionautoencodersbeenblurrycostdistributionelbogenerative
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Variational autoencoders (VAEs) are powerful generative modelling methods, however they suffer from blurry generated samples and reconstructions compared to the images they have been trained on. Significant research effort has been spent to increase the generative capabilities by creating more flexible models but often flexibility comes at the cost of higher complexity and computational cost. Several works have focused on altering the reconstruction term of the evidence lower bound (ELBO), however, often at the expense of losing the mathematical link to maximizing the likelihood of the samples under the modeled distribution. Here we propose a new formulation of the reconstruction term for the VAE that specifically penalizes the generation of blurry images while at the same time still maximizing the ELBO under the modeled distribution. We show the potential of the proposed loss on three different data sets, where it outperforms several recently proposed reconstruction losses for VAEs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Diffusion Prior Interpolation for Flexibility Real-World Face Super-Resolution

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A diffusion-based face super-resolution method using fixed and random masks plus a trained corrector network reports state-of-the-art perceptual quality and face recognition consistency on common benchmarks.

  2. Motion Free B-frame Coding for Neural Video Compression

    eess.IV 2024-11 conditional novelty 5.0 of 10

    A kernel-based, motion-free autoencoder for B-frame coding that synthesizes frames from two reconstructed references and an interpolated frame.

Pith tools