REVIEW 9 cited by
Learning a Diffusion Model Policy from Rewards via Q-Score Matching
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Diffusion models have become a popular choice for representing actor policies in behavior cloning and offline reinforcement learning. This is due to their natural ability to optimize an expressive class of distributions over a continuous space. However, previous works fail to exploit the score-based structure of diffusion models, and instead utilize a simple behavior cloning term to train the actor, limiting their ability in the actor-critic setting. In this paper, we present a theoretical framework linking the structure of diffusion model policies to a learned Q-function, by linking the structure between the score of the policy to the action gradient of the Q-function. We focus on off-policy reinforcement learning and propose a new policy update method from this theory, which we denote Q-score matching. Notably, this algorithm only needs to differentiate through the denoising model rather than the entire diffusion model evaluation, and converged policies through Q-score matching are implicitly multi-modal and explorative in continuous domains. We conduct experiments in simulated environments to demonstrate the viability of our proposed method and compare to popular baselines. Source code is available from the project website: https://michaelpsenka.io/qsm.
Forward citations
Cited by 9 Pith papers
-
Flow Matching Policy Gradients
FPO trains flow-based policies with PPO by replacing the likelihood ratio with an exponentiated flow matching loss difference.
-
GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning
GeMPO unifies diffusion RL reweighting as measure matching to a regularized target, enabling flexible and negative weights that improve exploration and performance.
-
Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners
In a large empirical study on a water-treatment PID control task, bounded beta policies with adaptive critic updates were the most reliable actor-critic configuration, while common defaults like Gaussian policies with...
-
FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning
An online imitation-learning method uses a flow-matching teacher's class-conditional loss as a reward and a regularizer to train a simple MLP policy, beating cloning and adversarial-imitation baselines on five of six tasks.
-
Habitizing Diffusion Planning for Efficient and Effective Decision Making
A variational-Bayes distillation framework (Habi) turns slow diffusion planners into fast feedforward policies that match their performance at orders-of-magnitude higher decision frequency on D4RL benchmarks.
-
Efficient Online Reinforcement Learning for Diffusion Policy
Diffusion policies can be trained online with reweighted score matching using only Q-functions, and the resulting DPMD and SDAC algorithms beat SAC and prior diffusion-policy RL on most MuJoCo tasks.
-
A Single Diffusion-Policy Controller for Multi-Task Block Pushing with Zero-Shot Sim-to-Real Transfer
One diffusion policy trained via energy-guided RL solves multi-shape block pushing without demos and transfers zero-shot to real robots under varied conditions.
-
Exploratory Diffusion Model for Unsupervised Reinforcement Learning
A diffusion-model denoising loss serves as an intrinsic reward to guide unsupervised RL exploration, plus an alternating fine-tuning scheme for diffusion policies.
-
Dual Control for Interactive Autonomous Merging with Model Predictive Diffusion
An active-learning dual controller with model predictive diffusion is validated on F1-Tenth hardware, merging in 4.3 m on average versus 7.1 m for the prior dual MPPI method.
Discussion (0). Continue with ORCID to comment.