REVIEW 11 cited by
Variational Diffusion Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Diffusion-based generative models have demonstrated a capacity for perceptually impressive synthesis, but can they also be great likelihood-based models? We answer this in the affirmative, and introduce a family of diffusion-based generative models that obtain state-of-the-art likelihoods on standard image density estimation benchmarks. Unlike other diffusion-based models, our method allows for efficient optimization of the noise schedule jointly with the rest of the model. We show that the variational lower bound (VLB) simplifies to a remarkably short expression in terms of the signal-to-noise ratio of the diffused data, thereby improving our theoretical understanding of this model class. Using this insight, we prove an equivalence between several models proposed in the literature. In addition, we show that the continuous-time VLB is invariant to the noise schedule, except for the signal-to-noise ratio at its endpoints. This enables us to learn a noise schedule that minimizes the variance of the resulting VLB estimator, leading to faster optimization. Combining these advances with architectural improvements, we obtain state-of-the-art likelihoods on image density estimation benchmarks, outperforming autoregressive models that have dominated these benchmarks for many years, with often significantly faster optimization. In addition, we show how to use the model as part of a bits-back compression scheme, and demonstrate lossless compression rates close to the theoretical optimum. Code is available at https://github.com/google-research/vdm .
Forward citations
Cited by 11 Pith papers
-
Flow Matching Policy Gradients
FPO trains flow-based policies with PPO by replacing the likelihood ratio with an exponentiated flow matching loss difference.
-
Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion
PPAD injects MLLM semantic feedback into diffusion denoising via lookahead sketches and ping-pong-ahead resampling, improving text-to-image alignment.
-
Flow Map Learning via Nongradient Vector Flow
SGFlow learns the integral map of a probability-flow ODE via a stop-gradient loss whose only stationary point is the true flow map, and it reaches the best-in-comparison FID at 10 steps on CIFAR-10.
-
D$e^+e^-$ffusion: Capturing the Beam-Beam Physics of $e^+e^-$ Collisions with Diffusion Models
A diffusion model trained on GuineaPig++ reproduces FCC-ee beam-induced pair-production distributions at particle and detector level, about 10^4 times faster.
-
DecNefSimulator: A Modular, Interpretable Framework for Decoded Neurofeedback Simulation Using Generative Models
A generative-model simulator of decoded neurofeedback shows that alternative-class choice, initial cognitive state, and random exploration jointly determine whether simulated participants learn or appear as non-responders.
-
Exploring Probabilistic Modeling Beyond Domain Generalization for Semantic Segmentation
PDAF estimates a latent domain prior with a lightweight diffusion model and uses it to condition segmentation features, improving domain-generalized semantic segmentation on four unseen urban datasets.
-
Transition Matching: Scalable and Flexible Generative Modeling
Transition Matching unifies flow matching and continuous autoregressive generation as discrete-time Markov processes, with three variants that improve text-to-image quality and speed.
-
Discrete Markov Bridge
Discrete Markov Bridge learns the forward rate matrix and the reverse score in a continuous-time Markov chain, achieving BPC 1.38 on Text8 and FID 11.63 on CIFAR-10.
-
Spatio-Temporal Conditional Diffusion Models for Forecasting Future Multiple Sclerosis Lesion Masks Conditioned on Treatments
A treatment-conditioned diffusion model generates future multiple sclerosis lesion masks from baseline MRI and better predicts lesion activity than population-level baselines.
-
On Designing Diffusion Autoencoders for Efficient Generation and Representation Learning
Small binary latents conditioned via cross-attention let a diffusion autoencoder generate from a uniform Bernoulli prior with fewer steps while keeping representation quality.
-
DiffPR: Diffusion-Based Phase Reconstruction via Frequency-Decoupled Learning
DiffPR couples a quarter-resolution U-Net phase predictor with an unconditional diffusion refiner and reports solid but modest gains over U-Net baselines on four QPI datasets, with the claimed spectral-bias mechanism ...
Discussion (0). Sign in to comment.