Pith. sign in

REVIEW 5 cited by

Diffusion Dynamics Models with Generative State Estimation for Cloth Manipulation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.11999 v2 pith:TPTH4X36 submitted 2025-03-15 cs.RO cs.CVcs.SYeess.SY

classification cs.ROcs.CVcs.SYeess.SY
keywords dynamicsclothmodelsstategenerativeestimationmanipulationmodeling
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Cloth manipulation is challenging due to its highly complex dynamics, near-infinite degrees of freedom, and frequent self-occlusions, which complicate both state estimation and dynamics modeling. Inspired by recent advances in generative models, we hypothesize that these expressive models can effectively capture intricate cloth configurations and deformation patterns from data. Therefore, we propose a diffusion-based generative approach for both perception and dynamics modeling. Specifically, we formulate state estimation as reconstructing full cloth states from partial observations and dynamics modeling as predicting future states given the current state and robot actions. Leveraging a transformer-based diffusion model, our method achieves accurate state reconstruction and reduces long-horizon dynamics prediction errors by an order of magnitude compared to prior approaches. We integrate our dynamics models with model predictive control and show that our framework enables effective cloth folding on real robotic systems, demonstrating the potential of generative models for deformable object manipulation under partial observability and complex dynamics.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Scaling Cross-Embodiment World Models for Dexterous Manipulation

    cs.RO 2025-11 conditional novelty 6.0 of 10

    A single particle-based world model trained on many simulated robot hands and real human hands can plan dexterous manipulation on robot hands it never trained on.

  2. LaGarNet: Goal-Conditioned Recurrent State-Space Models for Pick-and-Place Garment Flattening

    cs.RO 2025-08 unverdicted novelty 6.0 of 10

    The submission's abstract describes a new garment-flattening robot model, but the body is a different paper about document retrieval, making the submission internally inconsistent.

  3. Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new 23-dimension benchmark finds that even the best VLMs score near random on motion trajectory, temporal extension, and several prediction tasks, far below humans, suggesting weak internal world models.

  4. Language-Guided Long Horizon Manipulation with LLM-based Planning and Visual Perception

    cs.RO 2025-09 conditional novelty 5.0 of 10

    A robot folds cloth from spoken language by decomposing instructions with GPT-4o and grounding each step with a SigLIP2-based pick-and-place perception module.

  5. PinchBot: Long-Horizon Deformable Manipulation with Guided Diffusion Policy

    cs.RO 2025-07 conditional novelty 5.0 of 10

    A single goal-conditioned diffusion policy, combined with pre-trained point cloud embeddings and collision-constrained action projection, can create pottery bowls of 8, 10, and 12 centimeter diameters.

Pith tools