Pith. sign in

REVIEW 12 cited by

RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins (early version)

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.02920 v3 pith:DAZPJGYX submitted 2024-09-04 cs.RO cs.AIcs.CL

classification cs.ROcs.AIcs.CL
keywords dual-armdataevaluationmodelsreal-worldrobotwintasksdigital
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In the rapidly advancing field of robotics, dual-arm coordination and complex object manipulation are essential capabilities for developing advanced autonomous systems. However, the scarcity of diverse, high-quality demonstration data and real-world-aligned evaluation benchmarks severely limits such development. To address this, we introduce RoboTwin, a generative digital twin framework that uses 3D generative foundation models and large language models to produce diverse expert datasets and provide a real-world-aligned evaluation platform for dual-arm robotic tasks. Specifically, RoboTwin creates varied digital twins of objects from single 2D images, generating realistic and interactive scenarios. It also introduces a spatial relation-aware code generation framework that combines object annotations with large language models to break down tasks, determine spatial constraints, and generate precise robotic movement code. Our framework offers a comprehensive benchmark with both simulated and real-world data, enabling standardized evaluation and better alignment between simulated training and real-world performance. We validated our approach using the open-source COBOT Magic Robot platform. Policies pre-trained on RoboTwin-generated data and fine-tuned with limited real-world samples improve the success rate of over 70% for single-arm tasks and over 40% for dual-arm tasks compared to models trained solely on real-world data. This significant improvement demonstrates RoboTwin's potential to enhance the development and evaluation of dual-arm robotic manipulation systems. Project Page: https://robotwin-benchmark.github.io/early-version/.

Discussion (0). Sign in to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter

    cs.RO 2026-02 conditional novelty 6.0 of 10

    A decoupled multimodal diffusion transformer with a LoRA tactile adapter improves real-world bimanual manipulation success by 21 percentage points over a diffusion-policy baseline; a new 50-hour tactile bimanual datas...

  2. ManiFlow: A General Robot Manipulation Policy via Consistency Flow Training

    cs.RO 2025-09 conditional novelty 6.0 of 10

    ManiFlow trains a flow-matching policy with a continuous-time consistency objective and an adaptive cross-attention transformer, enabling dexterous manipulation with 1-2 inference steps and substantially higher succes...

  3. DISCOVERSE: Efficient Robot Simulation in Complex High-Fidelity Environments

    cs.RO 2025-07 conditional novelty 6.0 of 10

    A simulation framework that combines 3D Gaussian Splatting with MuJoCo reports improved zero-shot transfer of manipulation policies from simulation to real robots.

  4. Diffusion-Based Imaginative Coordination for Bimanual Manipulation

    cs.RO 2025-07 conditional novelty 6.0 of 10

    A diffusion-based policy that jointly predicts future video latents and actions improves bimanual manipulation success, with video prediction used only during training.

  5. TypeTele: Releasing Dexterity in Teleoperation by Dexterous Manipulation Types

    cs.RO 2025-07 conditional novelty 6.0 of 10

    A type-guided teleoperation system that selects predefined dexterous hand poses with a language model outperforms retargeting-based teleoperation on nine real-world tasks and improves imitation learning success.

  6. ControlVLA: Few-shot Object-centric Adaptation for Pre-trained Vision-Language-Action Models

    cs.RO 2025-06 conditional novelty 6.0 of 10

    ControlVLA adapts a DROID-pretrained diffusion VLA policy to new manipulation tasks with 10 to 20 demos by injecting object-centric features through zero-initialized cross-attention layers, achieving 76.7% success acr...

  7. BiAssemble: Learning Collaborative Affordance for Bimanual Geometric Assembly

    cs.RO 2025-06 conditional novelty 6.0 of 10

    BiAssemble predicts bimanual grasp and assembly actions for geometric reassembly of fractured objects via point-level collaborative affordance, and reports simulation gains over baselines plus a real-world benchmark.

  8. Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction

    cs.CV 2025-05 conditional novelty 6.0 of 10

    CogRobot uses optical flow as an intermediate variable to fine-tune a text-to-video model for predicting bimanual robot trajectories, then maps those predictions to actions with a goal-conditioned diffusion policy.

  9. SimLauncher: Launching Sample-Efficient Real-world Robotic Reinforcement Learning via Simulation Pre-training

    cs.RO 2025-07 conditional novelty 5.0 of 10

    Simulation-pretrained policies, with digital-twin demos for critic bootstrapping and action proposals, cut real-world RL training time while reaching near-perfect success on three manipulation tasks.

  10. AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation

    cs.RO 2025-07 conditional novelty 5.0 of 10

    AC-DiT adds mobility-to-body conditioning and perception-aware 2D/3D weighting to a diffusion transformer, improving success rates on simulated and real-world mobile manipulation tasks.

  11. PRISM: Pointcloud Reintegrated Inference via Segmentation and Cross-attention for Manipulation

    cs.RO 2025-07 conditional novelty 4.0 of 10

    PRISM trains a diffusion policy on segmented point-cloud object tokens fused with joint states via cross-attention, reporting 82.0 percent average success across six RoboTwin tasks versus 58.4 percent for DP3 and 22.3...

  12. Data Pyramid for Embodied Manipulation

    cs.RO 2026-07 conditional novelty 3.0 of 10

    Embodied training data form a five-layer pyramid—real-robot, UMI, ego/exo, simulation, general V–L—ordered by the trade-off between scale and robot alignment, and model capabilities track how those layers are mixed.

Pith tools