Pith. sign in

REVIEW 4 cited by

R+X: Retrieval and Execution from Everyday Human Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.12957 v2 pith:4JXPGGJZ submitted 2024-07-17 cs.RO cs.LG

classification cs.ROcs.LG
keywords videoseverydayhumanskillsbehaviourexecutionin-contextlanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present R+X, a framework which enables robots to learn skills from long, unlabelled, first-person videos of humans performing everyday tasks. Given a language command from a human, R+X first retrieves short video clips containing relevant behaviour, and then executes the skill by conditioning an in-context imitation learning method (KAT) on this behaviour. By leveraging a Vision Language Model (VLM) for retrieval, R+X does not require any manual annotation of the videos, and by leveraging in-context learning for execution, robots can perform commanded skills immediately, without requiring a period of training on the retrieved videos. Experiments studying a range of everyday household tasks show that R+X succeeds at translating unlabelled human videos into robust robot skills, and that R+X outperforms several recent alternative methods. Videos and code are available at https://www.robot-learning.uk/r-plus-x.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Deep Sensorimotor Control by Imitating Predictive Models of Human Motion

    cs.RO 2025-08 conditional novelty 7.0 of 10

    A predictive model of human hand motion, trained on human interaction data, can reward a robot policy for tracking predicted future keypoints and enable learning of dexterous manipulation from sparse rewards.

  2. Knowledge-Driven Imitation Learning: Enabling Generalization Across Diverse Conditions

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A semantic keypoint graph matched to novel objects lets imitation-learned manipulation policies generalize with a quarter of the demonstrations.

  3. EgoZero: Robot Learning from Smart Glasses

    cs.RO 2025-05 conditional novelty 6.0 of 10

    Robot policies trained only on egocentric human videos from smart glasses transfer zero-shot to a Franka gripper, with 70% success across 7 manipulation tasks.

  4. ImMimic: Cross-Domain Imitation from Human Videos via Mapping and Interpolation

    cs.RO 2025-09 conditional novelty 5.0 of 10

    A co-training framework that maps retargeted human hand trajectories to robot demonstrations with dynamic time warping and MixUp interpolation improves robot manipulation success rates and smoothness across four embodiments.

Pith tools