REVIEW 6 cited by
Cascaded Diffusion Models for High Fidelity Image Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We show that cascaded diffusion models are capable of generating high fidelity images on the class-conditional ImageNet generation benchmark, without any assistance from auxiliary image classifiers to boost sample quality. A cascaded diffusion model comprises a pipeline of multiple diffusion models that generate images of increasing resolution, beginning with a standard diffusion model at the lowest resolution, followed by one or more super-resolution diffusion models that successively upsample the image and add higher resolution details. We find that the sample quality of a cascading pipeline relies crucially on conditioning augmentation, our proposed method of data augmentation of the lower resolution conditioning inputs to the super-resolution models. Our experiments show that conditioning augmentation prevents compounding error during sampling in a cascaded model, helping us to train cascading pipelines achieving FID scores of 1.48 at 64x64, 3.52 at 128x128 and 4.88 at 256x256 resolutions, outperforming BigGAN-deep, and classification accuracy scores of 63.02% (top-1) and 84.06% (top-5) at 256x256, outperforming VQ-VAE-2.
Forward citations
Cited by 6 Pith papers
-
UltraZoom: Generating Gigapixel Images from Regular Photos
UltraZoom generates coherent gigapixel imagery from a regular full view and sparse close-ups by per-instance fine-tuning of a pretrained generative model with video-based registration.
-
FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation
FlowMo reduces temporal artifacts in video generation by guiding the denoising process to lower the maximum patch-wise variance of consecutive-frame differences in the latent space.
-
Hume: Introducing System-2 Thinking in Visual-Language-Action Model
A dual-system vision-language-action model that improves robot control by ranking multiple sampled action chunks with a learned value function before fast execution.
-
From Sound to Sight: Towards AI-authored Music Videos
This paper presents two off-the-shelf model pipelines (CLAP or LALM, an LLM, and a text-to-video model) for generating music videos from arbitrary songs, validated by a preliminary five-participant user study with mod...
-
CD-TVD: Contrastive Diffusion for 3D Super-Resolution with Scarce High-Resolution Time-Varying Data
CD-TVD pretrains a contrastive encoder and diffusion super-resolution network on historical simulation data, then fine-tunes with a single high-resolution timestep to reconstruct all low-resolution timesteps in a new ...
-
Retrieve-Augmented Generation for Speeding up Diffusion Policy without Additional Training
RAGDP accelerates pretrained diffusion policies by initializing denoising from the nearest retrieved expert demonstration action, improving accuracy-versus-speed trade-offs without extra training.
Discussion (0). Continue with ORCID to comment.