Pith. sign in

REVIEW 6 cited by

Cascaded Diffusion Models for High Fidelity Image Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.15282 v3 pith:F2WM47RK submitted 2021-05-30 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords diffusionmodelscascadedresolutionaugmentationconditioningimagemodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We show that cascaded diffusion models are capable of generating high fidelity images on the class-conditional ImageNet generation benchmark, without any assistance from auxiliary image classifiers to boost sample quality. A cascaded diffusion model comprises a pipeline of multiple diffusion models that generate images of increasing resolution, beginning with a standard diffusion model at the lowest resolution, followed by one or more super-resolution diffusion models that successively upsample the image and add higher resolution details. We find that the sample quality of a cascading pipeline relies crucially on conditioning augmentation, our proposed method of data augmentation of the lower resolution conditioning inputs to the super-resolution models. Our experiments show that conditioning augmentation prevents compounding error during sampling in a cascaded model, helping us to train cascading pipelines achieving FID scores of 1.48 at 64x64, 3.52 at 128x128 and 4.88 at 256x256 resolutions, outperforming BigGAN-deep, and classification accuracy scores of 63.02% (top-1) and 84.06% (top-5) at 256x256, outperforming VQ-VAE-2.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. UltraZoom: Generating Gigapixel Images from Regular Photos

    cs.CV 2025-06 conditional novelty 6.0 of 10

    UltraZoom generates coherent gigapixel imagery from a regular full view and sparse close-ups by per-instance fine-tuning of a pretrained generative model with video-based registration.

  2. FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    FlowMo reduces temporal artifacts in video generation by guiding the denoising process to lower the maximum patch-wise variance of consecutive-frame differences in the latent space.

  3. Hume: Introducing System-2 Thinking in Visual-Language-Action Model

    cs.RO 2025-05 conditional novelty 6.0 of 10

    A dual-system vision-language-action model that improves robot control by ranking multiple sampled action chunks with a learned value function before fast execution.

  4. From Sound to Sight: Towards AI-authored Music Videos

    cs.SD 2025-08 conditional novelty 5.0 of 10

    This paper presents two off-the-shelf model pipelines (CLAP or LALM, an LLM, and a text-to-video model) for generating music videos from arbitrary songs, validated by a preliminary five-participant user study with mod...

  5. CD-TVD: Contrastive Diffusion for 3D Super-Resolution with Scarce High-Resolution Time-Varying Data

    cs.CV 2025-08 conditional novelty 5.0 of 10

    CD-TVD pretrains a contrastive encoder and diffusion super-resolution network on historical simulation data, then fine-tunes with a single high-resolution timestep to reconstruct all low-resolution timesteps in a new ...

  6. Retrieve-Augmented Generation for Speeding up Diffusion Policy without Additional Training

    cs.LG 2025-07 conditional novelty 4.0 of 10

    RAGDP accelerates pretrained diffusion policies by initializing denoising from the nearest retrieved expert demonstration action, improving accuracy-versus-speed trade-offs without extra training.

Pith tools