Pith. sign in

REVIEW 4 cited by

Epsilon-VAE: Denoising as Visual Decoding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.04081 v4 pith:QKLXDLAM submitted 2024-10-05 cs.CV cs.AIeess.IV

classification cs.CVcs.AIeess.IV
keywords generationreconstructioncompressiondataiterativequalityvisualautoencoder
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In generative modeling, tokenization simplifies complex data into compact, structured representations, creating a more efficient, learnable space. For high-dimensional visual data, it reduces redundancy and emphasizes key features for high-quality generation. Current visual tokenization methods rely on a traditional autoencoder framework, where the encoder compresses data into latent representations, and the decoder reconstructs the original input. In this work, we offer a new perspective by proposing denoising as decoding, shifting from single-step reconstruction to iterative refinement. Specifically, we replace the decoder with a diffusion process that iteratively refines noise to recover the original image, guided by the latents provided by the encoder. We evaluate our approach by assessing both reconstruction (rFID) and generation quality (FID), comparing it to state-of-the-art autoencoding approaches. By adopting iterative reconstruction through diffusion, our autoencoder, namely Epsilon-VAE, achieves high reconstruction quality, which in turn enhances downstream generation quality by 22% at the same compression rates or provides 2.3x inference speedup through increasing compression rates. We hope this work offers new insights into integrating iterative generation and autoencoding for improved compression and generation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion

    cs.CV 2025-06 conditional novelty 7.0 of 10

    A hierarchical motion autoencoder with a conditional diffusion decoder reconstructs 16-frame videos from latents as small as 0.07% of the input size while maintaining competitive PSNR and perceptual scores.

  2. D-AR: Diffusion via Autoregressive Models

    cs.CV 2025-05 conditional novelty 7.0 of 10

    D-AR recasts pixel-space diffusion as vanilla autoregressive next-token prediction using a diffusion-ordered discrete tokenizer, reaching 2.09 FID on ImageNet 256x256 with a 775M Llama backbone.

  3. Diffusion Autoencoders are Scalable Image Tokenizers

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A single diffusion L2 loss can train scalable image tokenizers that match or outperform GAN-LPIPS tokenizers for reconstruction and downstream generation.

  4. DisCoRD: Discrete Tokens to Continuous Motion via Rectified Flow Decoding

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Replacing the decoder of discrete motion generation models with a rectified flow decoder improves naturalness and FID in text-to-motion, co-speech gesture, and music-to-dance generation.

Pith tools