REVIEW 15 cited by
THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
To realize effective large-scale, real-world robotic applications, we must evaluate how well our robot policies adapt to changes in environmental conditions. Unfortunately, a majority of studies evaluate robot performance in environments closely resembling or even identical to the training setup. We present THE COLOSSEUM, a novel simulation benchmark, with 20 diverse manipulation tasks, that enables systematical evaluation of models across 14 axes of environmental perturbations. These perturbations include changes in color, texture, and size of objects, table-tops, and backgrounds; we also vary lighting, distractors, physical properties perturbations and camera pose. Using THE COLOSSEUM, we compare 5 state-of-the-art manipulation models to reveal that their success rate degrades between 30-50% across these perturbation factors. When multiple perturbations are applied in unison, the success rate degrades $\geq$75%. We identify that changing the number of distractor objects, target object color, or lighting conditions are the perturbations that reduce model performance the most. To verify the ecological validity of our results, we show that our results in simulation are correlated ($\bar{R}^2 = 0.614$) to similar perturbations in real-world experiments. We open source code for others to use THE COLOSSEUM, and also release code to 3D print the objects used to replicate the real-world perturbations. Ultimately, we hope that THE COLOSSEUM will serve as a benchmark to identify modeling decisions that systematically improve generalization for manipulation. See https://robot-colosseum.github.io/ for more details.
Forward citations
Cited by 15 Pith papers
-
Towards Generalizable Robotic Manipulation in Dynamic Environments
DOMINO supplies 110K+ dynamic dual-arm trajectories across 35 tasks, and PUMA’s optical-flow history plus object-centric future queries raise dynamic success rate by 6.3 points over strong VLA baselines.
-
ARGUS: Aligning Robot Scene Geometry Under Shifting Views with Large 3D Vision Models
A preprocessing pipeline that reconstructs a 3D point cloud from RGB images and re-renders it from a fixed viewpoint improves viewpoint robustness and data efficiency for vision-based robot policies.
-
BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation
Adding stage-wise temporal and spatial memory to a heatmap-prediction 3D VLA policy yields strong results on memory-dependent manipulation benchmarks while keeping data efficiency.
-
Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data
Pretraining a VLA model on 18,561 hours of robot-synthesized egocentric human video mixed with robot data improves out-of-distribution manipulation success in simulation and on a real dual-arm robot.
-
It's Not Just More Demos: Counterfactual Action Sensitivity Coverage for Data-Efficient Robust Robot Imitation
Pairing clean and nuisance observations to measure action drift lets CFNBC select 20–30 counterfactual repair examples that outperform matched random selection for robust imitation.
-
RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation
Dense per-frame intermediate representations (traces, masks, grasp poses, subtasks) improve embodied VQA, VLA action generation, and world-model video prediction in the new 230k-episode RoboInter-Data suite.
-
Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control
Contrastive activation directions and a reduced-order LQR (WA-LQR) steer world-action models to recover robustness under camera, gripper, and noise shifts whenever the models' activations are linearly separable for th...
-
Beyond Visual Grasping: Benchmarking Complex Grasping from Detection to Execution
A new benchmark, GCA-Bench, evaluates robotic grasping from detection to execution across 102 complex tasks and finds current VLA and detection-based methods score below 70% success.
-
SoftVTBench: A Safety-Aware Visuo-Tactile Benchmark for Physically Constrained Robotic Manipulation of Deformable Objects
Success-only metrics overstate deformable-manipulation performance; tactile sensing raises Safety Success (e.g. 21.4%→35.6% on Object-Soft) while Goal Success stays comparable.
-
VLM-TDP: VLM-guided Trajectory-conditioned Diffusion Policy for Robust Long-Horizon Manipulation
VLM-TDP guides a diffusion-based robot policy with VLM-generated voxel trajectories, improving success rates by roughly 30-44% and adding robustness to noise and scene changes.
-
Gondola: Grounded Vision Language Planning for Generalizable Robotic Manipulation
Gondola generates multi-view segmentation-mask-grounded next-step plans for robotic manipulation and reports improved generalization on the GemBench benchmark over a prior LLM-based planner.
-
GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulation
GenManip is a benchmark and simulation platform with LLM-generated scene graphs for testing how robot policies generalize to new instructions, layouts, and objects.
-
Active Real-World Factor-Based Evaluation for Generalist Robot Policies
An active evaluation framework selects the most informative task configurations for real-robot tests, matching random testing's accuracy in 20-40% fewer trials.
-
SimLauncher: Launching Sample-Efficient Real-world Robotic Reinforcement Learning via Simulation Pre-training
Simulation-pretrained policies, with digital-twin demos for critic bootstrapping and action proposals, cut real-world RL training time while reaching near-perfect success on three manipulation tasks.
-
RoboPearls: Editable Video Simulation for Robot Manipulation
RoboPearls is a 3D Gaussian Splatting based framework that edits demonstration videos into varied photorealistic simulations, and training on them improves robot manipulation success rates on RLBench and COLOSSEUM.
Discussion (0). Continue with ORCID to comment.