Pith. sign in

REVIEW 11 cited by

Open-Ended Learning Leads to Generally Capable Agents

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2107.12808 v2 pith:223MK4VH submitted 2021-07-27 cs.LG cs.AIcs.MA

classification cs.LGcs.AIcs.MA
keywords agentagentslearningbehaviourtasksspacebehavioursbeyond
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this work we create agents that can perform well beyond a single, individual task, that exhibit much wider generalisation of behaviour to a massive, rich space of challenges. We define a universe of tasks within an environment domain and demonstrate the ability to train agents that are generally capable across this vast space and beyond. The environment is natively multi-agent, spanning the continuum of competitive, cooperative, and independent games, which are situated within procedurally generated physical 3D worlds. The resulting space is exceptionally diverse in terms of the challenges posed to agents, and as such, even measuring the learning progress of an agent is an open research problem. We propose an iterative notion of improvement between successive generations of agents, rather than seeking to maximise a singular objective, allowing us to quantify progress despite tasks being incomparable in terms of achievable rewards. We show that through constructing an open-ended learning process, which dynamically changes the training task distributions and training objectives such that the agent never stops learning, we achieve consistent learning of new behaviours. The resulting agent is able to score reward in every one of our humanly solvable evaluation levels, with behaviour generalising to many held-out points in the universe of tasks. Examples of this zero-shot generalisation include good performance on Hide and Seek, Capture the Flag, and Tag. Through analysis and hand-authored probe tasks we characterise the behaviour of our agent, and find interesting emergent heuristic behaviours such as trial-and-error experimentation, simple tool use, option switching, and cooperation. Finally, we demonstrate that the general capabilities of this agent could unlock larger scale transfer of behaviour through cheap finetuning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 55 citations worldwide. Full citation record

  1. Learning Task-Sufficient World Models by Synergizing Agentic Exploration and Structured Modeling

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Closed-loop agentic probing plus minimality/sufficiency masking recovers compact task-sufficient world-model latents that improve sample-efficient policy learning and cross-task generalization.

  2. Fly, Fail, Fix: Iterative Game Repair with Reinforcement Learning and Large Multimodal Models

    cs.AI 2025-07 conditional novelty 6.0 of 10

    An LMM iteratively repairs a Flappy Bird level generator to reach a target score using RL playtrace feedback in text, image, or combined form.

  3. FOUNDER: Grounding Foundation Models in World Models for Open-Ended Embodied Decision Making

    cs.RO 2025-07 conditional novelty 6.0 of 10

    FOUNDER maps foundation-model embeddings of text or video prompts into world-model goal states and rewards policies by predicted temporal distance to those goals, improving reward-free multi-task offline control.

  4. Uncertainty Prioritized Experience Replay

    cs.LG 2025-06 conditional novelty 6.0 of 10

    UPER uses ensemble-based epistemic and aleatoric uncertainty to compute an information gain priority for experience replay, outperforming TD-error prioritization on Atari-57.

  5. SPARQ: Synthetic Problem Generation for Reasoning via Quality-Diversity Algorithms

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Filtering self-generated math problems by a model's own solve-rate improves that model's MATH accuracy from 38% to 47% and helps out-of-distribution generalization when data is diverse.

  6. Path Generation and Evaluation in Video Games: A Nonparametric Statistical Approach

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A nonparametric model-free plus copula method generates controllable synthetic game paths, while an adapted three-sample test detects whether generated paths overfit or underfit the training data.

  7. Training RL Agents for Multi-Objective Network Defense Tasks

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Diverse, dynamically ordered training tasks make network-defense RL agents generalize to unseen attacks better than single-task training.

  8. Rethinking Agent Design: From Top-Down Workflows to Bottom-Up Skill Evolution

    cs.AI 2025-05 conditional novelty 6.0 of 10

    Agents that start with no game knowledge can build a reusable skill library through trial-and-error and visual feedback, then progress further in two complex games than baseline agents given extra hints.

  9. Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments

    cs.LG 2026-03 conditional novelty 5.0 of 10

    PPO plateaus can be avoided by increasing the number of parallel environments, which reduces both the outer-loop step size and update noise; scaling to 1M environments sustained improvement to 1T transitions.

  10. Agentic Services Computing

    cs.SE 2025-09 conditional novelty 5.0 of 10

    A position and survey paper that defines Agentic Services Computing, a lifecycle-based framework for engineering LLM agents as governed, first-class services.

  11. Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback

    cs.AI 2025-06 conditional novelty 5.0 of 10

    EXIF repeatedly has a teacher agent explore an environment, relabel the exploration as tasks, train a student agent on it, and use the student's failures to guide the next round, improving 7B-8B agents in Webshop and Crafter.

Pith tools