Pith. sign in

REVIEW 6 cited by

The Surprising Effectiveness of Representation Learning for Visual Imitation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2112.01511 v2 pith:GNUBTPFH submitted 2021-12-02 cs.RO cs.AIcs.CVcs.LG

classification cs.ROcs.AIcs.CVcs.LG
keywords learningvisualimitationrepresentationdatademonstrationsactionsdiverse
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While visual imitation learning offers one of the most effective ways of learning from visual demonstrations, generalizing from them requires either hundreds of diverse demonstrations, task specific priors, or large, hard-to-train parametric models. One reason such complexities arise is because standard visual imitation frameworks try to solve two coupled problems at once: learning a succinct but good representation from the diverse visual data, while simultaneously learning to associate the demonstrated actions with such representations. Such joint learning causes an interdependence between these two problems, which often results in needing large amounts of demonstrations for learning. To address this challenge, we instead propose to decouple representation learning from behavior learning for visual imitation. First, we learn a visual representation encoder from offline data using standard supervised and self-supervised learning methods. Once the representations are trained, we use non-parametric Locally Weighted Regression to predict the actions. We experimentally show that this simple decoupling improves the performance of visual imitation models on both offline demonstration datasets and real-robot door opening compared to prior work in visual imitation. All of our generated data, code, and robot videos are publicly available at https://jyopari.github.io/VINN/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CLASS: Contrastive Learning via Action Sequence Supervision for Robot Manipulation

    cs.RO 2025-08 unverdicted novelty 6.0 of 10

    CLASS pre-training with Diffusion Policy reaches 75% average success under visual shifts where baseline behavior cloning methods fail.

  2. UAD: Unsupervised Affordance Distillation for Generalization in Robotic Manipulation

    cs.RO 2025-06 conditional novelty 6.0 of 10

    UAD distills affordance knowledge from vision-language models and DINOv2 features into a lightweight task-conditioned model that predicts pixel-level manipulation regions and improves few-shot imitation learning gener...

  3. Object-Focus Actor for Data-efficient Robot Generalization Dexterous Manipulation

    cs.RO 2025-05 conditional novelty 6.0 of 10

    Object-Focus Actor makes dexterous manipulation policies generalize to new object positions and backgrounds by focusing on the hand-object region and using relative poses and actions, needing only 10 to 30 demonstrations.

  4. Adaptive Visuo-Tactile Fusion with Predictive Force Attention for Dexterous Manipulation

    cs.RO 2025-05 conditional novelty 6.0 of 10

    A force-guided attention module and future-force prediction auxiliary task improve visuo-tactile fusion for dexterous manipulation, reaching 93% average success in real robot trials.

  5. 3DFlowAction: Learning Cross-Embodiment Manipulation from 3D Flow World Model

    cs.RO 2025-06 conditional novelty 5.0 of 10

    A diffusion world model predicts 3D optical flow as an embodiment-agnostic action plan, and constrained optimization converts the flow into robot arm actions.

  6. Bootstrapping Imitation Learning for Long-horizon Manipulation via Hierarchical Data Collection Space

    cs.RO 2025-05 conditional novelty 5.0 of 10

    Breaking long manipulation tasks into atomic subtasks and collecting demonstrations from varied starting poses improves imitation learning success using fewer demonstration frames.

Pith tools