Pith. sign in

REVIEW 19 cited by

ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2107.14483 v5 pith:CU5PORFU submitted 2021-07-30 cs.LG cs.AIcs.CVcs.RO

classification cs.LGcs.AIcs.CVcs.RO
keywords benchmarkmanipulationmaniskilllearningresearchersassetsbaselineschallenge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Object manipulation from 3D visual inputs poses many challenges on building generalizable perception and policy models. However, 3D assets in existing benchmarks mostly lack the diversity of 3D shapes that align with real-world intra-class complexity in topology and geometry. Here we propose SAPIEN Manipulation Skill Benchmark (ManiSkill) to benchmark manipulation skills over diverse objects in a full-physics simulator. 3D assets in ManiSkill include large intra-class topological and geometric variations. Tasks are carefully chosen to cover distinct types of manipulation challenges. Latest progress in 3D vision also makes us believe that we should customize the benchmark so that the challenge is inviting to researchers working on 3D deep learning. To this end, we simulate a moving panoramic camera that returns ego-centric point clouds or RGB-D images. In addition, we would like ManiSkill to serve a broad set of researchers interested in manipulation research. Besides supporting the learning of policies from interactions, we also support learning-from-demonstrations (LfD) methods, by providing a large number of high-quality demonstrations (~36,000 successful trajectories, ~1.5M point cloud/RGB-D frames in total). We provide baselines using 3D deep learning and LfD algorithms. All code of our benchmark (simulator, environment, SDK, and baselines) is open-sourced, and a challenge facing interdisciplinary researchers will be held based on the benchmark.

Discussion (0). Sign in to comment.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Dynamic Execution Commitment of Vision-Language-Action Models

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    A3 reframes dynamic action chunk commitment in VLA models as self-speculative prefix verification, accepting the longest continuous sequence of actions that satisfies consensus-ordered conditional invariance and prefi...

  2. GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert

    cs.CV 2025-10 conditional novelty 7.0 of 10

    A frozen, large-scale pretrained diffusion policy converts sparse 3D waypoints from a VLM into dense robot actions, enabling zero-shot reuse of the action expert on new tasks and environments.

  3. PartInstruct: Part-level Instruction Following for Fine-grained Robot Manipulation

    cs.RO 2025-05 conditional novelty 7.0 of 10

    PartInstruct is a new large-scale simulated benchmark with part-level language instructions and training demonstrations; current robot policies achieve at most 31.72% average success on it.

  4. SSC: A Verifiable Structured Representation for Bimanual Manipulation Labelling

    cs.RO 2026-08 conditional novelty 6.0 of 10

    The Structured Subtask Chain (SSC) formats each manipulation subtask as a state-transition template and checks the whole chain for consistency, with vision-language models resolving ambiguous fields.

  5. Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A taxonomy of robot learning on a weights-versus-skills axis, with a five-rung self-improvement ladder whose top cell (feedback plus memory plus search) holds only a few recent systems.

  6. WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A critic that jointly predicts future latent states and values improves RL fine-tuning and out-of-distribution generalization for vision-language-action robot policies.

  7. Scale Up Strategically: Learning Compositional Generalization via Bias-Aware Evaluation and Data Collection for Robotic Manipulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Robot instruction-following policies consistently over-rely on color and under-ground verbs and size, and reallocating training demonstrations to under-grounded factors improves compositional generalization with fewer demos.

  8. PhyAgentOS: A Self-Evolving Operating System for Embodied Agents with Decoupled Cognitive Planning and Physical Execution

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A file-based operating-system layer with a session verifier and persistent memory improves embodied-agent task completion on game, simulated, and real-robot platforms without retraining policies.

  9. Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A unified 38B autoregressive world model for multi-view embodied scene generation, controllable transfer, and video synthesis improves real-robot OOD robustness when used as a data engine.

  10. HRIBench: Benchmarking Interaction-Centric Human-Robot Collaboration

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A new benchmark for intent-aware human-robot collaboration shows current VLA robot policies fail at coordination and safety, while simulated practice improves real-world task success from 0.10 to 0.43.

  11. Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models

    cs.RO 2026-02 conditional novelty 6.0 of 10

    Adding a real-world supervised loss to simulation reinforcement learning improves real-robot success and data efficiency for VLA co-training.

  12. VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models

    cs.RO 2025-12 conditional novelty 6.0 of 10

    An open benchmark with 170 graded manipulation tasks shows current VLA robot policies memorize their training settings, degrade sharply under visual shifts, ignore safety constraints, and fail to compose long-horizon skills.

  13. BiAssemble: Learning Collaborative Affordance for Bimanual Geometric Assembly

    cs.RO 2025-06 conditional novelty 6.0 of 10

    BiAssemble predicts bimanual grasp and assembly actions for geometric reassembly of fractured objects via point-level collaborative affordance, and reports simulation gains over baselines plus a real-world benchmark.

  14. ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making

    cs.RO 2025-05 conditional novelty 6.0 of 10

    ManiTaskGen automatically generates diverse, feasible mobile manipulation tasks from any input scene, and uses them to benchmark and improve vision-language robot agents.

  15. WorldEval: World Model as Real-World Robot Policies Evaluator

    cs.RO 2025-05 conditional novelty 6.0 of 10

    WorldEval conditions a video generation model on a policy's internal action embeddings (Policy2Vec) and shows generated-video success rates correlate with real-world robot success rates.

  16. IMBench: A Benchmark for Intuitive Robotic Manipulation

    cs.RO 2026-07 conditional novelty 5.0 of 10

    IMBench is a 35-task robosuite benchmark with a three-stage evaluation showing current VLMs and robot policies fail to convert physical reasoning into executable manipulation.

  17. DIPOLE: Fusing Vision and Geometry for Robust Visuomotor Generalization

    cs.RO 2025-11 conditional novelty 5.0 of 10

    Fusing RGB and point-cloud inputs with training-time modality dropout plus cross-attention makes a diffusion visuomotor policy markedly more robust to visual and spatial shifts than unimodal or naively fused baselines.

  18. VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

    cs.RO 2025-07 conditional novelty 5.0 of 10

    A convolutional residual VQ-VAE action tokenizer trained on over 100x more data than prior work improves OpenVLA success rates and inference speed on several manipulation tasks.

  19. PRISM: Pointcloud Reintegrated Inference via Segmentation and Cross-attention for Manipulation

    cs.RO 2025-07 conditional novelty 4.0 of 10

    PRISM trains a diffusion policy on segmented point-cloud object tokens fused with joint states via cross-attention, reporting 82.0 percent average success across six RoboTwin tasks versus 58.4 percent for DP3 and 22.3...

Pith tools