Pith. sign in

REVIEW 9 cited by

Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.11752 v5 pith:TGTLWJTB submitted 2023-12-18 cs.LG cs.AI

classification cs.LGcs.AI
keywords diffusionmodellearningmatchingpoliciespolicyq-scorestructure
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Diffusion models have become a popular choice for representing actor policies in behavior cloning and offline reinforcement learning. This is due to their natural ability to optimize an expressive class of distributions over a continuous space. However, previous works fail to exploit the score-based structure of diffusion models, and instead utilize a simple behavior cloning term to train the actor, limiting their ability in the actor-critic setting. In this paper, we present a theoretical framework linking the structure of diffusion model policies to a learned Q-function, by linking the structure between the score of the policy to the action gradient of the Q-function. We focus on off-policy reinforcement learning and propose a new policy update method from this theory, which we denote Q-score matching. Notably, this algorithm only needs to differentiate through the denoising model rather than the entire diffusion model evaluation, and converged policies through Q-score matching are implicitly multi-modal and explorative in continuous domains. We conduct experiments in simulated environments to demonstrate the viability of our proposed method and compare to popular baselines. Source code is available from the project website: https://michaelpsenka.io/qsm.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Flow Matching Policy Gradients

    cs.LG 2025-07 conditional novelty 7.0 of 10

    FPO trains flow-based policies with PPO by replacing the likelihood ratio with an exponentiated flow matching loss difference.

  2. GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning

    cs.LG 2026-03 conditional novelty 6.5 of 10

    GeMPO unifies diffusion RL reweighting as measure matching to a regularized target, enabling flexible and negative weights that improve exploration and performance.

  3. Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners

    cs.LG 2026-07 conditional novelty 6.0 of 10

    In a large empirical study on a water-treatment PID control task, bounded beta policies with adaptive critic updates were the most reliable actor-critic configuration, while common defaults like Gaussian policies with...

  4. FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning

    cs.LG 2025-10 conditional novelty 6.0 of 10

    An online imitation-learning method uses a flow-matching teacher's class-conditional loss as a reward and a regularizer to train a simple MLP policy, beating cloning and adversarial-imitation baselines on five of six tasks.

  5. Habitizing Diffusion Planning for Efficient and Effective Decision Making

    cs.LG 2025-02 conditional novelty 6.0 of 10

    A variational-Bayes distillation framework (Habi) turns slow diffusion planners into fast feedforward policies that match their performance at orders-of-magnitude higher decision frequency on D4RL benchmarks.

  6. Efficient Online Reinforcement Learning for Diffusion Policy

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Diffusion policies can be trained online with reweighted score matching using only Q-functions, and the resulting DPMD and SDAC algorithms beat SAC and prior diffusion-policy RL on most MuJoCo tasks.

  7. A Single Diffusion-Policy Controller for Multi-Task Block Pushing with Zero-Shot Sim-to-Real Transfer

    cs.RO 2026-07 conditional novelty 5.0 of 10

    One diffusion policy trained via energy-guided RL solves multi-shape block pushing without demos and transfers zero-shot to real robots under varied conditions.

  8. Exploratory Diffusion Model for Unsupervised Reinforcement Learning

    cs.LG 2025-02 conditional novelty 5.0 of 10

    A diffusion-model denoising loss serves as an intrinsic reward to guide unsupervised RL exploration, plus an alternating fine-tuning scheme for diffusion policies.

  9. Dual Control for Interactive Autonomous Merging with Model Predictive Diffusion

    cs.RO 2025-02 reject novelty 4.0 of 10

    An active-learning dual controller with model predictive diffusion is validated on F1-Tenth hardware, merging in 4.3 m on average versus 7.1 m for the prior dual MPPI method.

Pith tools