REVIEW 4 cited by
Accelerating Image Generation with Sub-path Linear Approximation Model
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Diffusion models have significantly advanced the state of the art in image, audio, and video generation tasks. However, their applications in practical scenarios are hindered by slow inference speed. Drawing inspiration from the approximation strategies utilized in consistency models, we propose the Sub-path Linear Approximation Model (SLAM), which accelerates diffusion models while maintaining high-quality image generation. SLAM treats the PF-ODE trajectory as a series of PF-ODE sub-paths divided by sampled points, and harnesses sub-path linear (SL) ODEs to form a progressive and continuous error estimation along each individual PF-ODE sub-path. The optimization on such SL-ODEs allows SLAM to construct denoising mappings with smaller cumulative approximated errors. An efficient distillation method is also developed to facilitate the incorporation of more advanced diffusion models, such as latent diffusion models. Our extensive experimental results demonstrate that SLAM achieves an efficient training regimen, requiring only 6 A100 GPU days to produce a high-quality generative model capable of 2 to 4-step generation with high performance. Comprehensive evaluations on LAION, MS COCO 2014, and MS COCO 2017 datasets also illustrate that SLAM surpasses existing acceleration methods in few-step generation tasks, achieving state-of-the-art performance both on FID and the quality of the generated images.
Forward citations
Cited by 4 Pith papers
-
TurboClear: One-Step Object-Effect Removal via Region-Calibrated Distribution Matching and Fusion
TurboClear combines region-calibrated distribution matching and learnable spatial fusion to turn a multi-step object-removal diffusion model into a one-step student with comparable or better quality and up to 665x low...
-
DMM: Building a Versatile Image Generation Model via Distillation-Based Model Merging
DMM trains a single style-promptable diffusion model to reproduce the outputs of multiple teacher models, achieving a merged-model FIDt of 77.51 versus a reference of 74.91.
-
Adversarial Diffusion Compression for Real-World Image Super-Resolution
AdcSR distills OSEDiff into a pruned diffusion-GAN that cuts inference time 3.7x and parameters 74% while achieving comparable super-resolution quality.
-
NitroFusion: High-Fidelity Single-Step Diffusion through Dynamic Adversarial Training
NitroFusion trains one-step diffusion models using a dynamic pool of noise-specialized discriminator heads, achieving better aesthetic scores but mixed FID results compared with prior one-step distillation.
Discussion (0). Continue with ORCID to comment.