REVIEW 8 cited by
FMB: a Functional Manipulation Benchmark for Generalizable Robotic Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this paper, we propose a real-world benchmark for studying robotic learning in the context of functional manipulation: a robot needs to accomplish complex long-horizon behaviors by composing individual manipulation skills in functionally relevant ways. The core design principles of our Functional Manipulation Benchmark (FMB) emphasize a harmonious balance between complexity and accessibility. Tasks are deliberately scoped to be narrow, ensuring that models and datasets of manageable scale can be utilized effectively to track progress. Simultaneously, they are diverse enough to pose a significant generalization challenge. Furthermore, the benchmark is designed to be easily replicable, encompassing all essential hardware and software components. To achieve this goal, FMB consists of a variety of 3D-printed objects designed for easy and accurate replication by other researchers. The objects are procedurally generated, providing a principled framework to study generalization in a controlled fashion. We focus on fundamental manipulation skills, including grasping, repositioning, and a range of assembly behaviors. The FMB can be used to evaluate methods for acquiring individual skills, as well as methods for combining and ordering such skills to solve complex, multi-stage manipulation tasks. We also offer an imitation learning framework that includes a suite of policies trained to solve the proposed tasks. This enables researchers to utilize our tasks as a versatile toolkit for examining various parts of the pipeline. For example, researchers could propose a better design for a grasping controller and evaluate it in combination with our baseline reorientation and assembly policies as part of a pipeline for solving multi-stage tasks. Our dataset, object CAD files, code, and evaluation videos can be found on our project website: https://functional-manipulation-benchmark.github.io
Forward citations
Cited by 8 Pith papers
-
Prediction with Action: Visual Policy Learning via Joint Denoising Process
PAD jointly denoises future images and robot actions in a single diffusion transformer, using video co-training to improve multi-task imitation learning.
-
Play2Perfect: What Matters in Dexterous Play Pretraining for Precise Assembly?
Task-agnostic RL play pretraining on diverse objects yields a reusable dexterous prior that makes sparse-reward assembly learning ~33× more sample-efficient and enables zero-shot sim-to-real transfer on tight insertio...
-
Fabrica: Dual-Arm Assembly of General Multi-Part Objects via Integrated Planning and Learning
A dual-arm robotic system combining hierarchical planning with equivariant residual RL policies demonstrates multi-part assembly of five-to-nine-part objects, with strong step-level but weaker end-to-end real-world success.
-
RLDG: Robotic Generalist Policy Distillation via Reinforcement Learning
Fine-tuning generalist robot policies on data generated by specialist RL agents beats fine-tuning on human demonstrations on precise manipulation tasks.
-
ClevrSkills: Compositional Language and Visual Reasoning in Robotics
ClevrSkills provides a 33-task, 330k-trajectory benchmark showing that vision-language robot policies struggle to compose base manipulation skills into novel long-horizon tasks.
-
Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning
FiS-VLA embeds a diffusion-based action module into the final transformer blocks of a vision-language model, achieving 69% mean success on RLBench and a claimed 117.7 Hz control frequency.
-
Improving Vision-Language-Action Model with Online Reinforcement Learning
Alternating online RL on a frozen vision-language backbone with supervised fine-tuning on collected successes improves a VLA policy's task success and generalization.
-
Data Pyramid for Embodied Manipulation: A Survey
Embodied training data form a five-layer pyramid—real-robot, UMI, ego/exo, simulation, general V–L—ordered by the trade-off between scale and robot alignment, and model capabilities track how those layers are mixed.
Discussion (0). Continue with ORCID to comment.