Pith. sign in

REVIEW 4 cited by

Asymmetric self-play for automatic goal discovery in robotic manipulation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2101.04882 v1 pith:5O3KXROK submitted 2021-01-13 cs.LG cs.AIcs.CVcs.RO

classification cs.LGcs.AIcs.CVcs.RO
keywords alicegoalspolicytasksasymmetricdiscoverygoalgoal-conditioned
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We train a single, goal-conditioned policy that can solve many robotic manipulation tasks, including tasks with previously unseen goals and objects. We rely on asymmetric self-play for goal discovery, where two agents, Alice and Bob, play a game. Alice is asked to propose challenging goals and Bob aims to solve them. We show that this method can discover highly diverse and complex goals without any human priors. Bob can be trained with only sparse rewards, because the interaction between Alice and Bob results in a natural curriculum and Bob can learn from Alice's trajectory when relabeled as a goal-conditioned demonstration. Finally, our method scales, resulting in a single policy that can generalize to many unseen tasks such as setting a table, stacking blocks, and solving simple puzzles. Videos of a learned policy is available at https://robotics-self-play.github.io.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Robust Autonomy Emerges from Self-Play

    cs.LG 2025-02 conditional novelty 8.0 of 10

    Self-play at 1.6 billion simulated kilometers yields a generalist driving policy that outperforms benchmark-specific specialists zero-shot on CARLA, nuPlan, and Waymax.

  2. Consistent Zero-Shot Imitation with Contrastive Goal Inference

    cs.LG 2025-10 reject novelty 6.0 of 10

    CIRL trains an agent with no rewards or demonstrations by having it propose and reach its own goals, then imitates a single test-time demonstration by inferring and reaching the demonstrated goal.

  3. Self-Challenging Language Model Agents

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A language model agent can generate its own verifiable training tasks and improve its tool-use success rate by about 2x without human-annotated data.

  4. Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback

    cs.AI 2025-06 conditional novelty 5.0 of 10

    EXIF repeatedly has a teacher agent explore an environment, relabel the exploration as tasks, train a student agent on it, and use the student's failures to guide the next round, improving 7B-8B agents in Webshop and Crafter.

Pith tools