Pith. sign in

REVIEW 8 cited by

Diffusion Models for Reinforcement Learning: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.01223 v4 pith:KAR2WV3G submitted 2023-11-02 cs.LG cs.AI

classification cs.LGcs.AI
keywords modelsdiffusionsurveychallengesgithublearningreinforcementresearch
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Diffusion models surpass previous generative models in sample quality and training stability. Recent works have shown the advantages of diffusion models in improving reinforcement learning (RL) solutions. This survey aims to provide an overview of this emerging field and hopes to inspire new avenues of research. First, we examine several challenges encountered by RL algorithms. Then, we present a taxonomy of existing methods based on the roles of diffusion models in RL and explore how the preceding challenges are addressed. We further outline successful applications of diffusion models in various RL-related tasks. Finally, we conclude the survey and offer insights into future research directions. We are actively maintaining a GitHub repository for papers and other related resources in utilizing diffusion models in RL: https://github.com/apexrl/Diff4RLSurvey.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. JAGG: Jacobian-Aggregated Group Gradient for Efficient GRPO Training of Diffusion Models

    cs.LG 2026-07 conditional novelty 7.0 of 10

    JAGG replaces per-step gradient backpropagation in diffusion GRPO with two endpoint backward passes joined by timestep-weighted interpolation, giving ~2x backward-pass savings at modest quality cost.

  2. VINE: Taming Generative Control Policies for Reinforcement Learning

    cs.RO 2026-07 conditional novelty 7.0 of 10

    Reconstructing a fresh noisy interpolation state at every denoising step stabilizes end-to-end value-gradient training of multi-step flow-matching policies and yields state-of-the-art offline and real-robot results.

  3. BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF

    cs.LG 2025-06 conditional novelty 7.0 of 10

    BadReward uses clean-label feature-collision images to poison CLIP-based reward models so that a text-to-image model produces target attributes (e.g., glasses, skin tone, blood) when the trigger phrase is present.

  4. STITCH-OPE: Trajectory Stitching with Guided Diffusion for Off-Policy Evaluation

    cs.RO 2025-05 conditional novelty 7.0 of 10

    STITCH-OPE uses stitched diffusion-generated sub-trajectories with negative behavior-policy guidance to perform off-policy evaluation in high-dimensional, long-horizon tasks.

  5. Object-centric Denoising Diffusion Models for Physical Reasoning

    cs.LG 2025-07 conditional novelty 6.0 of 10

    An object-centric diffusion model generates multi-object trajectories with conditioning at arbitrary time steps, demonstrated on the PHYRE physics benchmark.

  6. Diffusion-RL Based Air Traffic Conflict Detection and Resolution Method

    cs.AI 2025-09 conditional novelty 5.0 of 10

    Diffusion-AC, a diffusion-policy RL agent with dual-Q guidance and a density curriculum, beats PPO/TD3/DQN baselines in simulated 3D conflict resolution, cutting near-collisions by about 60% in dense traffic.

  7. Exploratory Diffusion Model for Unsupervised Reinforcement Learning

    cs.LG 2025-02 conditional novelty 5.0 of 10

    A diffusion-model denoising loss serves as an intrinsic reward to guide unsupervised RL exploration, plus an alternating fine-tuning scheme for diffusion policies.

  8. An Optimization-Augmented Control Framework for Single and Coordinated Multi-Arm Robotic Manipulation

    cs.RO 2025-06 conditional novelty 3.0 of 10

    A multi-modal controller that switches between optimization-based planning and force control completes simulated single-arm, bimanual, and four-arm manipulation tasks.

Pith tools