Pith. sign in

REVIEW 5 cited by

Learning to Generalize Across Long-Horizon Tasks from Human Demonstrations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.06085 v2 pith:UMAXXOKZ submitted 2020-03-13 cs.RO cs.AIcs.LG

classification cs.ROcs.AIcs.LG
keywords learninggeneralizeimitationtrainbehaviorsdemonstrationspoliciesreal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Imitation learning is an effective and safe technique to train robot policies in the real world because it does not depend on an expensive random exploration process. However, due to the lack of exploration, learning policies that generalize beyond the demonstrated behaviors is still an open challenge. We present a novel imitation learning framework to enable robots to 1) learn complex real world manipulation tasks efficiently from a small number of human demonstrations, and 2) synthesize new behaviors not contained in the collected demonstrations. Our key insight is that multi-task domains often present a latent structure, where demonstrated trajectories for different tasks intersect at common regions of the state space. We present Generalization Through Imitation (GTI), a two-stage offline imitation learning algorithm that exploits this intersecting structure to train goal-directed policies that generalize to unseen start and goal state combinations. In the first stage of GTI, we train a stochastic policy that leverages trajectory intersections to have the capacity to compose behaviors from different demonstration trajectories together. In the second stage of GTI, we collect a small set of rollouts from the unconditioned stochastic policy of the first stage, and train a goal-directed agent to generalize to novel start and goal configurations. We validate GTI in both simulated domains and a challenging long-horizon robotic manipulation domain in the real world. Additional results and videos are available at https://sites.google.com/view/gti2020/ .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PartInstruct: Part-level Instruction Following for Fine-grained Robot Manipulation

    cs.RO 2025-05 conditional novelty 7.0 of 10

    PartInstruct is a new large-scale simulated benchmark with part-level language instructions and training demonstrations; current robot policies achieve at most 31.72% average success on it.

  2. RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Dense per-frame intermediate representations (traces, masks, grasp poses, subtasks) improve embodied VQA, VLA action generation, and world-model video prediction in the new 230k-episode RoboInter-Data suite.

  3. Temporal Representation Alignment: Successor Features Enable Emergent Compositionality in Robot Instruction Following

    cs.RO 2025-02 conditional novelty 6.0 of 10

    A temporal alignment auxiliary loss on goal and language representations improves zero-shot compositional generalization in robot instruction following.

  4. MuST: Multi-Head Skill Transformer for Long-Horizon Dexterous Manipulation with Skill Progress

    cs.RO 2025-02 conditional novelty 6.0 of 10

    MuST adds per-skill action heads and a progress-guided skill selector to the Octo robot policy, improving long-horizon pick-and-pack success from about 32% to 90% in one simulated setting.

  5. Steering Robots with Inference-Time Interactions

    cs.RO 2025-06 conditional novelty 4.0 of 10

    Frozen imitation policies can be steered at inference time via user interactions, with a diffusion-sampling method and a constraint-enforcing framework that provides formal task guarantees.

Pith tools