REVIEW 8 cited by
Diffusion Models for Reinforcement Learning: A Survey
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Diffusion models surpass previous generative models in sample quality and training stability. Recent works have shown the advantages of diffusion models in improving reinforcement learning (RL) solutions. This survey aims to provide an overview of this emerging field and hopes to inspire new avenues of research. First, we examine several challenges encountered by RL algorithms. Then, we present a taxonomy of existing methods based on the roles of diffusion models in RL and explore how the preceding challenges are addressed. We further outline successful applications of diffusion models in various RL-related tasks. Finally, we conclude the survey and offer insights into future research directions. We are actively maintaining a GitHub repository for papers and other related resources in utilizing diffusion models in RL: https://github.com/apexrl/Diff4RLSurvey.
Forward citations
Cited by 8 Pith papers
-
JAGG: Jacobian-Aggregated Group Gradient for Efficient GRPO Training of Diffusion Models
JAGG replaces per-step gradient backpropagation in diffusion GRPO with two endpoint backward passes joined by timestep-weighted interpolation, giving ~2x backward-pass savings at modest quality cost.
-
VINE: Taming Generative Control Policies for Reinforcement Learning
Reconstructing a fresh noisy interpolation state at every denoising step stabilizes end-to-end value-gradient training of multi-step flow-matching policies and yields state-of-the-art offline and real-robot results.
-
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF
BadReward uses clean-label feature-collision images to poison CLIP-based reward models so that a text-to-image model produces target attributes (e.g., glasses, skin tone, blood) when the trigger phrase is present.
-
STITCH-OPE: Trajectory Stitching with Guided Diffusion for Off-Policy Evaluation
STITCH-OPE uses stitched diffusion-generated sub-trajectories with negative behavior-policy guidance to perform off-policy evaluation in high-dimensional, long-horizon tasks.
-
Object-centric Denoising Diffusion Models for Physical Reasoning
An object-centric diffusion model generates multi-object trajectories with conditioning at arbitrary time steps, demonstrated on the PHYRE physics benchmark.
-
Diffusion-RL Based Air Traffic Conflict Detection and Resolution Method
Diffusion-AC, a diffusion-policy RL agent with dual-Q guidance and a density curriculum, beats PPO/TD3/DQN baselines in simulated 3D conflict resolution, cutting near-collisions by about 60% in dense traffic.
-
Exploratory Diffusion Model for Unsupervised Reinforcement Learning
A diffusion-model denoising loss serves as an intrinsic reward to guide unsupervised RL exploration, plus an alternating fine-tuning scheme for diffusion policies.
-
An Optimization-Augmented Control Framework for Single and Coordinated Multi-Arm Robotic Manipulation
A multi-modal controller that switches between optimization-based planning and force control completes simulated single-arm, bimanual, and four-arm manipulation tasks.
Discussion (0). Continue with ORCID to comment.