Pith. sign in

REVIEW 6 cited by

Optimizing DDPM Sampling with Shortcut Fine-Tuning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.13362 v4 pith:PYG2Z6WY submitted 2023-01-31 cs.LG

classification cs.LG
keywords diffusionfine-tuningshortcutddpmgradientmodelssamplerssampling
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this study, we propose Shortcut Fine-Tuning (SFT), a new approach for addressing the challenge of fast sampling of pretrained Denoising Diffusion Probabilistic Models (DDPMs). SFT advocates for the fine-tuning of DDPM samplers through the direct minimization of Integral Probability Metrics (IPM), instead of learning the backward diffusion process. This enables samplers to discover an alternative and more efficient sampling shortcut, deviating from the backward diffusion process. Inspired by a control perspective, we propose a new algorithm SFT-PG: Shortcut Fine-Tuning with Policy Gradient, and prove that under certain assumptions, gradient descent of diffusion models with respect to IPM is equivalent to performing policy gradient. To our best knowledge, this is the first attempt to utilize reinforcement learning (RL) methods to train diffusion models. Through empirical evaluation, we demonstrate that our fine-tuning method can further enhance existing fast DDPM samplers, resulting in sample quality comparable to or even surpassing that of the full-step model across various datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Curvature-Adaptive Consistency Flow Matching: Autonomous Trajectory Optimization via Reinforcement Learning

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    CACFM applies RL to adaptively select critical regions in probability flow ODE trajectories for consistency distillation, yielding SOTA few-step results on FLUX and SDXL.

  2. NaP-Control: Navigating Diffusion Prior for Versatile and Fast Character Control

    cs.GR 2026-04 unverdicted novelty 6.0 of 10

    NaP-Control uses RL to directly predict optimized diffusion noise from a task-agnostic prior, enabling fast inference and higher success rates for versatile whole-body character control while preserving motion quality.

  3. Conditional Diffusion Guidance under Hard Constraint: A Stochastic Analysis Approach

    cs.AI 2026-02 conditional novelty 6.0 of 10

    By adding drift g(t)^2 ∇log h(t,y) with h estimated via martingale and covariation losses, diffusion samples can be hard-conditioned on an event.

  4. Directly Aligning the Full Diffusion Trajectory with Fine-Grained Human Preference

    cs.AI 2025-09 conditional novelty 6.0 of 10

    Direct-Align and SRPO fine-tune FLUX using ground-truth-noise recovery and text-conditional relative rewards, improving human-evaluated realism and aesthetics roughly 3x.

  5. Text2Stereo: Repurposing Stable Diffusion for Stereo Generation with Consistency Rewards

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Text2Stereo adapts Stable Diffusion to generate wide-baseline stereo image pairs from text by fine-tuning with LoRA and a disparity-correlation consistency reward.

  6. Constraints-Guided Diffusion Reasoner for Neuro-Symbolic Learning

    cs.AI 2025-08 conditional novelty 4.0 of 10

    A masked diffusion model fine-tuned with group-reward reinforcement learning solves Sudoku and Maze puzzles more accurately than the same model trained by supervised learning alone.

Pith tools