Pith. sign in

REVIEW 12 cited by

Diffusion Guidance Is a Controllable Policy Improvement Operator

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.23458 v1 pith:E7YY3HPC submitted 2025-05-29 cs.LG

classification cs.LG
keywords learningguidanceperformancecfgrldatadiffusionfurtherimprovement
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

At the core of reinforcement learning is the idea of learning beyond the performance in the data. However, scaling such systems has proven notoriously tricky. In contrast, techniques from generative modeling have proven remarkably scalable and are simple to train. In this work, we combine these strengths, by deriving a direct relation between policy improvement and guidance of diffusion models. The resulting framework, CFGRL, is trained with the simplicity of supervised learning, yet can further improve on the policies in the data. On offline RL tasks, we observe a reliable trend -- increased guidance weighting leads to increased performance. Of particular importance, CFGRL can operate without explicitly learning a value function, allowing us to generalize simple supervised methods (e.g., goal-conditioned behavioral cloning) to further prioritize optimality, gaining performance for "free" across the board.

Discussion (0). Sign in to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VINE: Taming Generative Control Policies for Reinforcement Learning

    cs.RO 2026-07 conditional novelty 7.0 of 10

    Reconstructing a fresh noisy interpolation state at every denoising step stabilizes end-to-end value-gradient training of multi-step flow-matching policies and yields state-of-the-art offline and real-robot results.

  2. GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning

    cs.LG 2026-03 conditional novelty 6.5 of 10

    GeMPO unifies diffusion RL reweighting as measure matching to a regularized target, enabling flexible and negative weights that improve exploration and performance.

  3. CLIFT: Turning Gemini Robotics On-Device into Humanoid Specialists via Non-Invasive Closed-Loop Iterative Fine-Tuning

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Closed-loop fine-tuning through a supervised-data API alone — no weights, gradients, or losses — lifts a closed-weight humanoid VLA to near-perfect success on three contact-rich tasks after two self-improvement cycles.

  4. RedFlow: Redirect Failure into Action-Level Corrections for Flow-matching VLA Policy

    cs.RO 2026-07 conditional novelty 6.0 of 10

    RedFlow improves flow-matching VLA policies offline by labeling individual actions from failed rollouts with progress signals and pushing those actions toward successful alternatives found in similar contexts.

  5. TACO: TActile World Model as a Self-COrrector forScalable VLA Post-Training

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A tactile-aware world model recognizes failure-adjacent contact states, imagines local visuo-tactile corrections, and post-trains VLAs with knowledge insulation, raising average success by 44% over the base policy.

  6. Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence

    cs.AI 2026-04 unverdicted novelty 6.0 of 10

    Projection-aware activation steering using logistic regression recovers honesty and compassion under malicious prompts while preserving coherence and benchmark performance better than uniform steering.

  7. Update-Free On-Policy Steering via Verifiers

    cs.RO 2026-03 conditional novelty 6.0 of 10

    Lightweight verifiers trained on a diffusion policy’s own evaluation rollouts raise real-robot success rates ~49% on average via Best-of-N or classifier guidance, without changing base parameters.

  8. From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning

    cs.RO 2026-03 accept novelty 6.0 of 10

    Residual off-policy RL with selective BC regularization and value-guided sampling contracts a pretrained generative robot policy around successful actions, reaching high success on hard long-horizon tasks from pixels ...

  9. Latent Reasoning in TRMs is Secretly a Policy Improvement Operator

    cs.CL 2025-11 reject novelty 6.0 of 10

    Recursive reasoning in TRMs is reinterpreted as policy improvement, and a stepwise denoising supervision scheme (DIS) cuts forward passes 18× while improving small-model ARC scores.

  10. ALOE: Action-Level Off-Policy Evaluation for Vision-Language-Action Model Post-Training

    cs.RO 2026-02 conditional novelty 5.0 of 10

    ALOE uses chunked TD bootstrapping with a pessimistic Q-ensemble to enable action-level off-policy value estimation for advantage-weighted post-training of flow-based VLA policies, reporting consistent success-rate ga...

  11. Dichotomous Diffusion Policy Optimization

    cs.LG 2025-12 conditional novelty 5.0 of 10

    DIPOLE decomposes a KL-regularized RL objective into a pair of sigmoid-weighted diffusion policies whose score combination (CFG-like) yields stable and controllable policy improvement.

  12. Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance

    cs.LG 2025-08 conditional novelty 4.0 of 10

    Blending a base diffusion model with its RL-finetuned version at sampling time lets users dial alignment strength, with the blend weight corresponding to the KL-regularization coefficient beta/w.

Pith tools