REVIEW 7 cited by
One-Step Diffusion Policy: Fast Visuomotor Policies via Diffusion Distillation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Diffusion models, praised for their success in generative tasks, are increasingly being applied to robotics, demonstrating exceptional performance in behavior cloning. However, their slow generation process stemming from iterative denoising steps poses a challenge for real-time applications in resource-constrained robotics setups and dynamically changing environments. In this paper, we introduce the One-Step Diffusion Policy (OneDP), a novel approach that distills knowledge from pre-trained diffusion policies into a single-step action generator, significantly accelerating response times for robotic control tasks. We ensure the distilled generator closely aligns with the original policy distribution by minimizing the Kullback-Leibler (KL) divergence along the diffusion chain, requiring only $2\%$-$10\%$ additional pre-training cost for convergence. We evaluated OneDP on 6 challenging simulation tasks as well as 4 self-designed real-world tasks using the Franka robot. The results demonstrate that OneDP not only achieves state-of-the-art success rates but also delivers an order-of-magnitude improvement in inference speed, boosting action prediction frequency from 1.5 Hz to 62 Hz, establishing its potential for dynamic and computationally constrained robotic applications. We share the project page at https://research.nvidia.com/labs/dir/onedp/.
Forward citations
Cited by 7 Pith papers
-
Genuine pair density wave order on the kagome lattice
A genuine primary pair-density-wave phase emerges as a competing ground state in a two-orbital kagome Hubbard model over a wide parameter range, driven by sublattice- and orbital-polarized Fermi pockets.
-
Spatial Attention: Adapting Execution Horizons for Diffusion Policies via Observation Sensitivity
Under a fixed sampling budget, execution horizons that minimize disturbance-induced likelihood drop should shorten as Spatial Attention rises; forecasting it yields higher success rates than fixed horizons.
-
Optimal Transport Q-Learning for Flow Policy Steering and Acceleration
Advantage-weighted conditional optimal transport flow matching simultaneously steers flow policies toward high-value actions and straightens their integration paths, enabling 2-3 step inference while improving task success.
-
High-Fidelity One-Step Generative Visuomotor Policy via Recursive Correction, Frequency Consistency, and Contrastive Flow Matching
One-step flow-matching visuomotor policy with recursive correction, dual-timestep spectral consistency, and contrastive mode separation matches or exceeds 10-step baselines at 1 NFE.
-
NavCMPO: Critic-Guided MeanFlow Policy Optimization for Adaptive Navigation
A two-stage navigation policy using five-step MeanFlow generation, critic-guided trajectory refinement, and PPO fine-tuning reports higher success and lower latency than a matched NavDP baseline.
-
RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures
RoboTALES uses hierarchical LLM subgoals and VLM reward feedback to keep video-model futures task-aligned, then trains robot policies that beat baselines on RoboCasa and LIBERO10 long-horizon tasks.
-
Retrieve-Augmented Generation for Speeding up Diffusion Policy without Additional Training
RAGDP accelerates pretrained diffusion policies by initializing denoising from the nearest retrieved expert demonstration action, improving accuracy-versus-speed trade-offs without extra training.
Discussion (0). Continue with ORCID to comment.