REVIEW 11 cited by
ManiCM: Real-time 3D Diffusion Policy via Consistency Model for Robotic Manipulation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Diffusion models have been verified to be effective in generating complex distributions from natural images to motion trajectories. Recent diffusion-based methods show impressive performance in 3D robotic manipulation tasks, whereas they suffer from severe runtime inefficiency due to multiple denoising steps, especially with high-dimensional observations. To this end, we propose a real-time robotic manipulation model named ManiCM that imposes the consistency constraint to the diffusion process, so that the model can generate robot actions in only one-step inference. Specifically, we formulate a consistent diffusion process in the robot action space conditioned on the point cloud input, where the original action is required to be directly denoised from any point along the ODE trajectory. To model this process, we design a consistency distillation technique to predict the action sample directly instead of predicting the noise within the vision community for fast convergence in the low-dimensional action manifold. We evaluate ManiCM on 31 robotic manipulation tasks from Adroit and Metaworld, and the results demonstrate that our approach accelerates the state-of-the-art method by 10 times in average inference speed while maintaining competitive average success rate.
Forward citations
Cited by 11 Pith papers
-
FA-RDP: A Frequency-Adaptive Reactive Diffusion Policy for Contact-Rich Manipulation
A shared visual-force diffusion policy with a multimodality indicator and manifold consistency distillation raises contact-rich task success to 81.7% while keeping diverse pre-contact modes.
-
SegDiff: Segmented Trajectory Diffusion for Consistent and Adaptive Robot Manipulation
SegDiff predicts continuous trajectories anchored to the next keypose and uses DDIM inversion for dynamic temporal ensembling, outperforming continuous and keypose baselines on RLBench, RoboMimic, and five real tasks.
-
Optimal Transport Q-Learning for Flow Policy Steering and Acceleration
Advantage-weighted conditional optimal transport flow matching simultaneously steers flow policies toward high-value actions and straightens their integration paths, enabling 2-3 step inference while improving task success.
-
High-Fidelity One-Step Generative Visuomotor Policy via Recursive Correction, Frequency Consistency, and Contrastive Flow Matching
One-step flow-matching visuomotor policy with recursive correction, dual-timestep spectral consistency, and contrastive mode separation matches or exceeds 10-step baselines at 1 NFE.
-
ManiFlow: A General Robot Manipulation Policy via Consistency Flow Training
ManiFlow trains a flow-matching policy with a continuous-time consistency objective and an adaptive cross-attention transformer, enabling dexterous manipulation with 1-2 inference steps and substantially higher succes...
-
DemoSpeedup: Accelerating Visuomotor Policies via Entropy-Guided Demonstration Acceleration
DemoSpeedup accelerates visuomotor policies by downsampling high-entropy segments of demonstrations, achieving roughly 2x faster execution with maintained or improved success rates.
-
CordViP: Correspondence-based Visuomotor Policy for Dexterous Manipulation in Real-World
CordViP achieves strong real-world dexterous manipulation by feeding a diffusion policy with pose-tracked 3D object models and hand point clouds, pretrained on contact maps and arm-hand coordination.
-
A Continuous-Time Consistency Model for 3D Point Cloud Generation
ConTiCoM-3D trains a continuous-time consistency-style model directly on raw 3D point clouds using flow matching plus Chamfer distance, with one- to two-step generation.
-
Time-Unified Diffusion Policy with Action Discrimination for Robotic Manipulation
TUDP removes timestep conditioning from diffusion policies and adds an action-discrimination signal to learn a time-unified velocity field, achieving SOTA RLBench success rates (82.6% multi-view, 83.8% single-view) an...
-
Detecting Reading-Induced Confusion Using EEG and Eye Tracking
Multimodal EEG plus eye tracking classifies reading-induced confusion at 77.3% average weighted accuracy, beating unimodal models by 4-22%, in an 11-participant study.
-
Benchmarking Generalizable Bimanual Manipulation: RoboTwin Dual-Arm Collaboration Challenge at CVPR 2025 MEIS Workshop
Results and lessons from the RoboTwin Dual-Arm Collaboration Challenge at CVPR 2025, covering 64 teams and 17 bimanual manipulation tasks across simulation and real hardware.
Discussion (0). Continue with ORCID to comment.