REVIEW 9 cited by
Robot See Robot Do: Imitating Articulated Object Manipulation with Monocular 4D Reconstruction
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Humans can learn to manipulate new objects by simply watching others; providing robots with the ability to learn from such demonstrations would enable a natural interface specifying new behaviors. This work develops Robot See Robot Do (RSRD), a method for imitating articulated object manipulation from a single monocular RGB human demonstration given a single static multi-view object scan. We first propose 4D Differentiable Part Models (4D-DPM), a method for recovering 3D part motion from a monocular video with differentiable rendering. This analysis-by-synthesis approach uses part-centric feature fields in an iterative optimization which enables the use of geometric regularizers to recover 3D motions from only a single video. Given this 4D reconstruction, the robot replicates object trajectories by planning bimanual arm motions that induce the demonstrated object part motion. By representing demonstrations as part-centric trajectories, RSRD focuses on replicating the demonstration's intended behavior while considering the robot's own morphological limits, rather than attempting to reproduce the hand's motion. We evaluate 4D-DPM's 3D tracking accuracy on ground truth annotated 3D part trajectories and RSRD's physical execution performance on 9 objects across 10 trials each on a bimanual YuMi robot. Each phase of RSRD achieves an average of 87% success rate, for a total end-to-end success rate of 60% across 90 trials. Notably, this is accomplished using only feature fields distilled from large pretrained vision models -- without any task-specific training, fine-tuning, dataset collection, or annotation. Project page: https://robot-see-robot-do.github.io
Forward citations
Cited by 9 Pith papers
-
StructureGS: Structure-aware Gaussian Splatting for Articulated Object Reconstruction
OBB-based part-fitting and contact losses on 3D Gaussians disentangle geometry, appearance, and motion for cleaner articulated reconstruction than photometric-only baselines.
-
Robots Acquire Manipulation Skills in Seconds from a Single Human Video
A single human video is enough for a robot to acquire a new manipulation skill at inference time, with no parameter updates, reaching 62% success on 50 novel tasks.
-
PokeNet: Learning Kinematic Models of Articulated Objects from Human Observations
PokeNet estimates joint types, axes, ranges, and operation order of articulated objects directly from a single-view point cloud video of a human demonstration.
-
One View, Many Worlds: Single-Image to 3D Object Meets Generative Domain Randomization for One-Shot 6D Pose Estimation
Given one RGB-D photo of an unseen object, an AI-generated 3D mesh, aligned jointly in metric scale and pose, yields state-of-the-art one-shot 6D pose estimation on YCBInEOAT, TOYL, and LM-O.
-
Learning in ImaginationLand: Omnidirectional Policies through 3D Generative Models (OP-Gen)
A robot policy trained on one real demonstration plus AI-generated 3D views succeeds from novel initial poses, including opposite-side starts, across six real manipulation tasks.
-
ScrewSplat: An End-to-End Method for Articulated Object Recognition
A method that recovers the 3D shape and the rotation or sliding axis of each movable part of an object from RGB video alone, by jointly optimizing randomly initialized screw axes with Gaussian Splatting.
-
Object-Focus Actor for Data-efficient Robot Generalization Dexterous Manipulation
Object-Focus Actor makes dexterous manipulation policies generalize to new object positions and backgrounds by focusing on the hand-object region and using relative poses and actions, needing only 10 to 30 demonstrations.
-
ODeform: Learning Continuous 4D Motion for Shape Deformation with Neural ODEs
ODeform combines two parallel neural ODEs, one for rigid motion and one for local deformation, to predict arbitrary-time 3D point-cloud deformation from an initial state and physical parameters, outperforming simpler ...
- CodeDiffuser: Attention-Enhanced Diffusion Policy via VLM-Generated Code for Instruction Ambiguity
Discussion (0). Sign in to comment.