REVIEW 6 cited by
Vision-based Manipulation from Single Human Video with Open-World Object Graphs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This work presents an object-centric approach to learning vision-based manipulation skills from human videos. We investigate the problem of robot manipulation via imitation in the open-world setting, where a robot learns to manipulate novel objects from a single video demonstration. We introduce ORION, an algorithm that tackles the problem by extracting an object-centric manipulation plan from a single RGB or RGB-D video and deriving a policy that conditions on the extracted plan. Our method enables the robot to learn from videos captured by daily mobile devices and to generalize the policies to deployment environments with varying visual backgrounds, camera angles, spatial layouts, and novel object instances. We systematically evaluate our method on both short-horizon and long-horizon tasks, using RGB-D and RGB-only demonstration videos. Across varied tasks and demonstration types (RGB-D / RGB), we observe an average success rate of 74.4%, demonstrating the efficacy of ORION in learning from a single human video in the open world. Additional materials can be found on our project website: https://ut-austin-rpl.github.io/ORION-release.
Forward citations
Cited by 6 Pith papers
-
MimicDroid: In-Context Learning for Humanoid Robot Manipulation from Human Play Videos
Trained only on unlabeled human play videos, MimicDroid lets a GR1 humanoid perform new manipulation tasks from one to three demonstration videos, with roughly twice the real-world success of prior video-conditioned methods.
-
Weakly-Supervised Learning of Dense Functional Correspondences
A weakly-supervised pipeline that distills VLM functional part knowledge and multi-view spatial structure into a model for dense cross-category functional correspondence, outperforming baselines on new synthetic and r...
-
Act, Sense, Act: Learning Active Perception from Large-Scale Egocentric Human Data
CoMe-VLA combines cognitive subtask labels and dual-track memory with human egocentric pretraining, reaching 83% mean success on five active-perception manipulation tasks.
-
MimicFunc: Imitating Tool Manipulation from a Single Human Video via Functional Correspondence
MimicFunc transfers tool-using skills from one human video to novel tools by aligning function-centric keypoint frames, achieving 79.5% success across five manipulation tasks.
-
Object-Focus Actor for Data-efficient Robot Generalization Dexterous Manipulation
Object-Focus Actor makes dexterous manipulation policies generalize to new object positions and backgrounds by focusing on the hand-object region and using relative poses and actions, needing only 10 to 30 demonstrations.
-
ObjRetarget: An Object-Aware Motion Retargeting Framework with Anthropomorphic Arm Constraints and Polyhedral Hand Modeling
Decoupled arm–hand retargeting with anthropomorphic arm-plane constraints and polyhedral contact invariants raises real-robot dexterous-task success to 75.8% versus 61.6% and 50.8% for OKAMI and ORION.
Discussion (0). Sign in to comment.