Pith. sign in

REVIEW 3 cited by

Improved Vector Quantized Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.16007 v2 pith:RZSRSVZ6 submitted 2022-05-31 cs.CV

classification cs.CV
keywords vq-diffusiondiffusionscoreclassifier-freeguidanceimproveimprovedmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vector quantized diffusion (VQ-Diffusion) is a powerful generative model for text-to-image synthesis, but sometimes can still generate low-quality samples or weakly correlated images with text input. We find these issues are mainly due to the flawed sampling strategy. In this paper, we propose two important techniques to further improve the sample quality of VQ-Diffusion. 1) We explore classifier-free guidance sampling for discrete denoising diffusion model and propose a more general and effective implementation of classifier-free guidance. 2) We present a high-quality inference strategy to alleviate the joint distribution issue in VQ-Diffusion. Finally, we conduct experiments on various datasets to validate their effectiveness and show that the improved VQ-Diffusion suppresses the vanilla version by large margins. We achieve an 8.44 FID score on MSCOCO, surpassing VQ-Diffusion by 5.42 FID score. When trained on ImageNet, we dramatically improve the FID score from 11.89 to 4.83, demonstrating the superiority of our proposed techniques.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Dissecting Bit-Level Scaling Laws in Quantizing Vision Generative Models

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Token-based language-style vision models (VAR, LlamaGen) tolerate quantization better than diffusion models, and a custom TopKLD distillation loss pushes their low-bit scaling roughly one precision level higher.

  2. PersonaMagic: Stage-Regulated High-Fidelity Face Customization with Tandem Equilibrium

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A single-image face personalization method that trains timestep-dependent face embeddings only in an intermediate denoising stage and balances text-encoder attention to improve both identity and editability.

  3. Black box behavioural modelling: Predicting human activity schedules with a deep conditional generative approach

    cs.LG 2025-12 conditional novelty 5.0 of 10

    A conditional variational autoencoder generates diverse, realistic daily activity schedules from demographic labels, outperforming both a most-likely-schedule baseline and a label-blind generative baseline on density ...

Pith tools