Pith. sign in

REVIEW 3 cited by

Diffusion Models for Robotic Manipulation: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.08438 v3 pith:3WF7Q6QX submitted 2025-04-11 cs.RO stat.ML

classification cs.ROstat.ML
keywords diffusionmodelslearningaugmentationdataimagemanipulationrobotic
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Diffusion generative models have demonstrated remarkable success in visual domains such as image and video generation. They have also recently emerged as a promising approach in robotics, especially in robot manipulations. Diffusion models leverage a probabilistic framework, and they stand out with their ability to model multi-modal distributions and their robustness to high-dimensional input and output spaces. This survey provides a comprehensive review of state-of-the-art diffusion models in robotic manipulation, including grasp learning, trajectory planning, and data augmentation. Diffusion models for scene and image augmentation lie at the intersection of robotics and computer vision for vision-based tasks to enhance generalizability and data scarcity. This paper also presents the two main frameworks of diffusion models and their integration with imitation learning and reinforcement learning. In addition, it discusses the common architectures and benchmarks and points out the challenges and advantages of current state-of-the-art diffusion-based methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification

    cs.CV 2025-08 conditional novelty 6.0 of 10

    CogVLA pairs instruction-conditioned visual-token aggregation (EFA-Routing) with transformer-layer pruning (LFP-Routing) and bidirectional action decoding (CAtten), reporting LIBERO 97.4%, real-world 70.0%, 2.5x less ...

  2. SoccerDiffusion: Toward Learning End-to-End Humanoid Robot Soccer from Gameplay Recordings

    cs.RO 2025-04 conditional novelty 6.0 of 10

    An end-to-end transformer diffusion policy, distilled to one inference step, reproduces low-level humanoid soccer behaviors from real RoboCup game recordings but lacks high-level tactical behavior.

  3. Constraint-Aware Diffusion Guidance for Robotics: Real-Time Obstacle Avoidance for Autonomous Racing

    cs.RO 2025-05 conditional novelty 4.0 of 10

    A diffusion trajectory planner with a barrier-function guidance term and warm starting avoids obstacles in real time on a miniature race car, with 100% success in the reported trials.

Pith tools