Pith. sign in

REVIEW 4 cited by

Hand-Object Interaction Pretraining from Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.08273 v1 pith:4IODRXPS submitted 2024-09-12 cs.RO cs.AIcs.CV

classification cs.ROcs.AIcs.CV
keywords policyrobotgeneralhand-objecthumaninteractionmanipulationprior
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present an approach to learn general robot manipulation priors from 3D hand-object interaction trajectories. We build a framework to use in-the-wild videos to generate sensorimotor robot trajectories. We do so by lifting both the human hand and the manipulated object in a shared 3D space and retargeting human motions to robot actions. Generative modeling on this data gives us a task-agnostic base policy. This policy captures a general yet flexible manipulation prior. We empirically demonstrate that finetuning this policy, with both reinforcement learning (RL) and behavior cloning (BC), enables sample-efficient adaptation to downstream tasks and simultaneously improves robustness and generalizability compared to prior approaches. Qualitative experiments are available at: \url{https://hgaurav2k.github.io/hop/}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction Synthesis

    cs.CV 2026-07 conditional novelty 7.0 of 10

    HarmoHOI jointly generates synchronized multi-view hand-object interaction videos and globally aligned 3D point tracks from a single reference image and target camera poses.

  2. Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A dexterous VLA pretrained on a 2.5M-instance human hand motion dataset transfers skills to a real robot hand, outperforming baselines in manipulation tasks.

  3. SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios

    cs.CV 2025-06 conditional novelty 6.0 of 10

    SViMo jointly generates HOI videos and explicit 3D hand-object motion via synchronized diffusion with a closed-loop 3D interaction diffusion model.

  4. EgoZero: Robot Learning from Smart Glasses

    cs.RO 2025-05 conditional novelty 6.0 of 10

    Robot policies trained only on egocentric human videos from smart glasses transfer zero-shot to a Franka gripper, with 70% success across 7 manipulation tasks.

Pith tools