Pith. sign in

REVIEW 3 cited by

PaGoDA: Progressive Growing of a One-Step Generator from a Low-Resolution Diffusion Teacher

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.14822 v2 pith:MOZIAAFR submitted 2024-05-23 cs.CV cs.AIcs.LGstat.ML

classification cs.CVcs.AIcs.LGstat.ML
keywords diffusionpagodatrainingpipelineprogressiveautoencoderdatadownsampled
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

The diffusion model performs remarkable in generating high-dimensional content but is computationally intensive, especially during training. We propose Progressive Growing of Diffusion Autoencoder (PaGoDA), a novel pipeline that reduces the training costs through three stages: training diffusion on downsampled data, distilling the pretrained diffusion, and progressive super-resolution. With the proposed pipeline, PaGoDA achieves a $64\times$ reduced cost in training its diffusion model on 8x downsampled data; while at the inference, with the single-step, it performs state-of-the-art on ImageNet across all resolutions from 64x64 to 512x512, and text-to-image. PaGoDA's pipeline can be applied directly in the latent space, adding compression alongside the pre-trained autoencoder in Latent Diffusion Models (e.g., Stable Diffusion). The code is available at https://github.com/sony/pagoda.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Geometry-Preserving Encoder/Decoder in Latent Generative Models

    math.NA 2025-01 conditional novelty 6.0 of 10

    A geometry-preserving encoder/decoder trained by minimizing a logarithmic Gromov-Monge distance converges provably and reconstructs images faster than VAE-based latent generative models.

  2. Adversarial Diffusion Compression for Real-World Image Super-Resolution

    eess.IV 2024-11 conditional novelty 6.0 of 10

    AdcSR distills OSEDiff into a pruned diffusion-GAN that cuts inference time 3.7x and parameters 74% while achieving comparable super-resolution quality.

  3. Inconsistencies In Consistency Models: Better ODE Solving Does Not Imply Better Samples

    cs.LG 2024-11 conditional novelty 6.0 of 10

    Directly supervising a consistency model against an ODE solver lowers ODE solving error yet degrades image quality, so better ODE solving does not imply better samples.

Pith tools