REVIEW 19 cited by
ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Object manipulation from 3D visual inputs poses many challenges on building generalizable perception and policy models. However, 3D assets in existing benchmarks mostly lack the diversity of 3D shapes that align with real-world intra-class complexity in topology and geometry. Here we propose SAPIEN Manipulation Skill Benchmark (ManiSkill) to benchmark manipulation skills over diverse objects in a full-physics simulator. 3D assets in ManiSkill include large intra-class topological and geometric variations. Tasks are carefully chosen to cover distinct types of manipulation challenges. Latest progress in 3D vision also makes us believe that we should customize the benchmark so that the challenge is inviting to researchers working on 3D deep learning. To this end, we simulate a moving panoramic camera that returns ego-centric point clouds or RGB-D images. In addition, we would like ManiSkill to serve a broad set of researchers interested in manipulation research. Besides supporting the learning of policies from interactions, we also support learning-from-demonstrations (LfD) methods, by providing a large number of high-quality demonstrations (~36,000 successful trajectories, ~1.5M point cloud/RGB-D frames in total). We provide baselines using 3D deep learning and LfD algorithms. All code of our benchmark (simulator, environment, SDK, and baselines) is open-sourced, and a challenge facing interdisciplinary researchers will be held based on the benchmark.
Forward citations
Cited by 19 Pith papers
-
Dynamic Execution Commitment of Vision-Language-Action Models
A3 reframes dynamic action chunk commitment in VLA models as self-speculative prefix verification, accepting the longest continuous sequence of actions that satisfies consensus-ordered conditional invariance and prefi...
-
GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert
A frozen, large-scale pretrained diffusion policy converts sparse 3D waypoints from a VLM into dense robot actions, enabling zero-shot reuse of the action expert on new tasks and environments.
-
PartInstruct: Part-level Instruction Following for Fine-grained Robot Manipulation
PartInstruct is a new large-scale simulated benchmark with part-level language instructions and training demonstrations; current robot policies achieve at most 31.72% average success on it.
-
SSC: A Verifiable Structured Representation for Bimanual Manipulation Labelling
The Structured Subtask Chain (SSC) formats each manipulation subtask as a state-transition template and checks the whole chain for consistency, with vision-language models resolving ambiguous fields.
-
Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills
A taxonomy of robot learning on a weights-versus-skills axis, with a five-rung self-improvement ladder whose top cell (feedback plus memory plus search) holds only a few recent systems.
-
WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning
A critic that jointly predicts future latent states and values improves RL fine-tuning and out-of-distribution generalization for vision-language-action robot policies.
-
Scale Up Strategically: Learning Compositional Generalization via Bias-Aware Evaluation and Data Collection for Robotic Manipulation
Robot instruction-following policies consistently over-rely on color and under-ground verbs and size, and reallocating training demonstrations to under-grounded factors improves compositional generalization with fewer demos.
-
PhyAgentOS: A Self-Evolving Operating System for Embodied Agents with Decoupled Cognitive Planning and Physical Execution
A file-based operating-system layer with a session verifier and persistent memory improves embodied-agent task completion on game, simulated, and real-robot platforms without retraining policies.
-
Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model
A unified 38B autoregressive world model for multi-view embodied scene generation, controllable transfer, and video synthesis improves real-robot OOD robustness when used as a data engine.
-
HRIBench: Benchmarking Interaction-Centric Human-Robot Collaboration
A new benchmark for intent-aware human-robot collaboration shows current VLA robot policies fail at coordination and safety, while simulated practice improves real-world task success from 0.10 to 0.43.
-
Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models
Adding a real-world supervised loss to simulation reinforcement learning improves real-robot success and data efficiency for VLA co-training.
-
VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models
An open benchmark with 170 graded manipulation tasks shows current VLA robot policies memorize their training settings, degrade sharply under visual shifts, ignore safety constraints, and fail to compose long-horizon skills.
-
BiAssemble: Learning Collaborative Affordance for Bimanual Geometric Assembly
BiAssemble predicts bimanual grasp and assembly actions for geometric reassembly of fractured objects via point-level collaborative affordance, and reports simulation gains over baselines plus a real-world benchmark.
-
ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making
ManiTaskGen automatically generates diverse, feasible mobile manipulation tasks from any input scene, and uses them to benchmark and improve vision-language robot agents.
-
WorldEval: World Model as Real-World Robot Policies Evaluator
WorldEval conditions a video generation model on a policy's internal action embeddings (Policy2Vec) and shows generated-video success rates correlate with real-world robot success rates.
-
IMBench: A Benchmark for Intuitive Robotic Manipulation
IMBench is a 35-task robosuite benchmark with a three-stage evaluation showing current VLMs and robot policies fail to convert physical reasoning into executable manipulation.
-
DIPOLE: Fusing Vision and Geometry for Robust Visuomotor Generalization
Fusing RGB and point-cloud inputs with training-time modality dropout plus cross-attention makes a diffusion visuomotor policy markedly more robust to visual and spatial shifts than unimodal or naively fused baselines.
-
VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers
A convolutional residual VQ-VAE action tokenizer trained on over 100x more data than prior work improves OpenVLA success rates and inference speed on several manipulation tasks.
-
PRISM: Pointcloud Reintegrated Inference via Segmentation and Cross-attention for Manipulation
PRISM trains a diffusion policy on segmented point-cloud object tokens fused with joint states via cross-attention, reporting 82.0 percent average success across six RoboTwin tasks versus 58.4 percent for DP3 and 22.3...
Discussion (0). Sign in to comment.