REVIEW 6 cited by
Optimizing DDPM Sampling with Shortcut Fine-Tuning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this study, we propose Shortcut Fine-Tuning (SFT), a new approach for addressing the challenge of fast sampling of pretrained Denoising Diffusion Probabilistic Models (DDPMs). SFT advocates for the fine-tuning of DDPM samplers through the direct minimization of Integral Probability Metrics (IPM), instead of learning the backward diffusion process. This enables samplers to discover an alternative and more efficient sampling shortcut, deviating from the backward diffusion process. Inspired by a control perspective, we propose a new algorithm SFT-PG: Shortcut Fine-Tuning with Policy Gradient, and prove that under certain assumptions, gradient descent of diffusion models with respect to IPM is equivalent to performing policy gradient. To our best knowledge, this is the first attempt to utilize reinforcement learning (RL) methods to train diffusion models. Through empirical evaluation, we demonstrate that our fine-tuning method can further enhance existing fast DDPM samplers, resulting in sample quality comparable to or even surpassing that of the full-step model across various datasets.
Forward citations
Cited by 6 Pith papers
-
Curvature-Adaptive Consistency Flow Matching: Autonomous Trajectory Optimization via Reinforcement Learning
CACFM applies RL to adaptively select critical regions in probability flow ODE trajectories for consistency distillation, yielding SOTA few-step results on FLUX and SDXL.
-
NaP-Control: Navigating Diffusion Prior for Versatile and Fast Character Control
NaP-Control uses RL to directly predict optimized diffusion noise from a task-agnostic prior, enabling fast inference and higher success rates for versatile whole-body character control while preserving motion quality.
-
Conditional Diffusion Guidance under Hard Constraint: A Stochastic Analysis Approach
By adding drift g(t)^2 ∇log h(t,y) with h estimated via martingale and covariation losses, diffusion samples can be hard-conditioned on an event.
-
Directly Aligning the Full Diffusion Trajectory with Fine-Grained Human Preference
Direct-Align and SRPO fine-tune FLUX using ground-truth-noise recovery and text-conditional relative rewards, improving human-evaluated realism and aesthetics roughly 3x.
-
Text2Stereo: Repurposing Stable Diffusion for Stereo Generation with Consistency Rewards
Text2Stereo adapts Stable Diffusion to generate wide-baseline stereo image pairs from text by fine-tuning with LoRA and a disparity-correlation consistency reward.
-
Constraints-Guided Diffusion Reasoner for Neuro-Symbolic Learning
A masked diffusion model fine-tuned with group-reward reinforcement learning solves Sudoku and Maze puzzles more accurately than the same model trained by supervised learning alone.
Discussion (0). Continue with ORCID to comment.