Pith. sign in

REVIEW 5 cited by

Third-Person Imitation Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1703.01703 v2 pith:FQJ4SE5N submitted 2017-03-06 cs.LG

classification cs.LG
keywords learningthird-persondemonstrationsimitationagentdomainproblemprovided
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Reinforcement learning (RL) makes it possible to train agents capable of achieving sophisticated goals in complex and uncertain environments. A key difficulty in reinforcement learning is specifying a reward function for the agent to optimize. Traditionally, imitation learning in RL has been used to overcome this problem. Unfortunately, hitherto imitation learning methods tend to require that demonstrations are supplied in the first-person: the agent is provided with a sequence of states and a specification of the actions that it should have taken. While powerful, this kind of imitation learning is limited by the relatively hard problem of collecting first-person demonstrations. Humans address this problem by learning from third-person demonstrations: they observe other humans perform tasks, infer the task, and accomplish the same task themselves. In this paper, we present a method for unsupervised third-person imitation learning. Here third-person refers to training an agent to correctly achieve a simple goal in a simple environment when it is provided a demonstration of a teacher achieving the same goal but from a different viewpoint; and unsupervised refers to the fact that the agent receives only these third-person demonstrations, and is not provided a correspondence between teacher states and student states. Our methods primary insight is that recent advances from domain confusion can be utilized to yield domain agnostic features which are crucial during the training process. To validate our approach, we report successful experiments on learning from third-person demonstrations in a pointmass domain, a reacher domain, and inverted pendulum.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FOUNDER: Grounding Foundation Models in World Models for Open-Ended Embodied Decision Making

    cs.RO 2025-07 conditional novelty 6.0 of 10

    FOUNDER maps foundation-model embeddings of text or video prompts into world-model goal states and rewards policies by predicted temporal distance to those goals, improving reward-free multi-task offline control.

  2. State-Conditional Adversarial Learning: An Off-Policy Visual Domain Transfer Method for End-to-End Imitation Learning

    cs.RO 2025-12 unverdicted novelty 5.0 of 10

    SCAL derives an upper bound on target-domain imitation loss using source loss plus state-conditional latent KL divergence and aligns distributions via a discriminator-based adversarial estimator.

  3. Beyond Domain Randomization: Event-Inspired Perception for Visually Robust Adversarial Imitation from Videos

    cs.CV 2025-05 conditional novelty 5.0 of 10

    Converting RGB videos into event-like temporal gradients lets an imitation agent ignore appearance differences between expert and learner domains, improving robustness without data augmentation.

  4. MTDP: A Modulated Transformer based Diffusion Policy Model

    cs.RO 2025-02 conditional novelty 4.0 of 10

    Combining scale-shift conditioning with cross-attention in diffusion policies yields success-rate gains of 0 to 12 percentage points on simulated robot manipulation benchmarks, with the largest gain on Toolhang.

  5. Reinforcement Learning: From Algorithms To Foundation Models

    cs.AI 2026-07 conditional novelty 3.0 of 10

    A dissertation uniting the author's published results: non-exploitable Nash-DQN policies and the FightLadder benchmark for games, plus diffusion/consistency-model world models for RL — a compilation rather than new results.

Pith tools