Pith. sign in

REVIEW 3 cited by

Consistency Models as a Rich and Efficient Policy Class for Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.16984 v2 pith:HHNR2WVK submitted 2023-09-29 cs.LG

classification cs.LG
keywords policyconsistencydiffusionefficientmodelmodelsdatagenerative
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Score-based generative models like the diffusion model have been testified to be effective in modeling multi-modal data from image generation to reinforcement learning (RL). However, the inference process of diffusion model can be slow, which hinders its usage in RL with iterative sampling. We propose to apply the consistency model as an efficient yet expressive policy representation, namely consistency policy, with an actor-critic style algorithm for three typical RL settings: offline, offline-to-online and online. For offline RL, we demonstrate the expressiveness of generative models as policies from multi-modal data. For offline-to-online RL, the consistency policy is shown to be more computational efficient than diffusion policy, with a comparable performance. For online RL, the consistency policy demonstrates significant speedup and even higher average performances than the diffusion policy.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ClutterDexGrasp: A Sim-to-Real System for General Dexterous Grasping in Cluttered Scenes

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A simulation-trained teacher-student policy achieves zero-shot sim-to-real closed-loop target-oriented dexterous grasping in cluttered scenes, with 83.9 percent real-world success.

  2. Predictive Planner for Autonomous Driving with Consistency Models

    cs.RO 2025-02 conditional novelty 5.0 of 10

    A consistency-model-based predictive planner generates joint ego and agent trajectories in four sampling steps, with an alternating guided-sampling scheme to satisfy planning constraints.

  3. Reinforcement Learning: From Algorithms To Foundation Models

    cs.AI 2026-07 conditional novelty 3.0 of 10

    A dissertation uniting the author's published results: non-exploitable Nash-DQN policies and the FightLadder benchmark for games, plus diffusion/consistency-model world models for RL — a compilation rather than new results.

Pith tools