Pith. sign in

REVIEW 4 cited by

Towards Controllable Diffusion Models via Reward-Guided Exploration

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.07132 v1 pith:TRAQS3SZ submitted 2023-04-14 cs.LG q-bio.BM

classification cs.LGq-bio.BM
keywords diffusionmodelsconditionalframeworksamplesclassifiercontrollingenables
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

By formulating data samples' formation as a Markov denoising process, diffusion models achieve state-of-the-art performances in a collection of tasks. Recently, many variants of diffusion models have been proposed to enable controlled sample generation. Most of these existing methods either formulate the controlling information as an input (i.e.,: conditional representation) for the noise approximator, or introduce a pre-trained classifier in the test-phase to guide the Langevin dynamic towards the conditional goal. However, the former line of methods only work when the controlling information can be formulated as conditional representations, while the latter requires the pre-trained guidance classifier to be differentiable. In this paper, we propose a novel framework named RGDM (Reward-Guided Diffusion Model) that guides the training-phase of diffusion models via reinforcement learning (RL). The proposed training framework bridges the objective of weighted log-likelihood and maximum entropy RL, which enables calculating policy gradients via samples from a pay-off distribution proportional to exponential scaled rewards, rather than from policies themselves. Such a framework alleviates the high gradient variances and enables diffusion models to explore for highly rewarded samples in the reverse process. Experiments on 3D shape and molecule generation tasks show significant improvements over existing conditional diffusion models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Modular Diffusion Policy Training: Decoupling and Recombining Guidance and Diffusion for Offline RL

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Guidance-first diffusion training, which trains and freezes the value guidance before the policy, improves sample efficiency and enables cross-algorithm reuse of guidance modules in offline RL.

  2. Lyapunov Guidance: A Unified Framework for Stabilizing Generative Flows

    cs.LG 2026-07 reject novelty 5.0 of 10

    Flow guidance is framed as Lyapunov control with a pseudo-projection for stability, but the projected flow is not shown to sample the target conditional distribution.

  3. A Reward-Directed Diffusion Framework for Generative Design Optimization

    cs.LG 2025-08 conditional novelty 5.0 of 10

    The paper reports a reward-directed diffusion framework that fine-tunes a DDPM with reward-weighted likelihood and then samples with soft-value importance weighting, claiming 25% resistance reduction in ship hulls and...

  4. VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL

    cs.CV 2025-05 conditional novelty 5.0 of 10

    VARD fine-tunes diffusion models by backpropagating through a learned value function that assigns dense, differentiable reward estimates to every intermediate denoising step, with KL regularization keeping the model n...

Pith tools