REVIEW 14 cited by
Human-Timescale Adaptation in an Open-Ended Task Space
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Foundation models have shown impressive adaptation and scalability in supervised and self-supervised learning problems, but so far these successes have not fully translated to reinforcement learning (RL). In this work, we demonstrate that training an RL agent at scale leads to a general in-context learning algorithm that can adapt to open-ended novel embodied 3D problems as quickly as humans. In a vast space of held-out environment dynamics, our adaptive agent (AdA) displays on-the-fly hypothesis-driven exploration, efficient exploitation of acquired knowledge, and can successfully be prompted with first-person demonstrations. Adaptation emerges from three ingredients: (1) meta-reinforcement learning across a vast, smooth and diverse task distribution, (2) a policy parameterised as a large-scale attention-based memory architecture, and (3) an effective automated curriculum that prioritises tasks at the frontier of an agent's capabilities. We demonstrate characteristic scaling laws with respect to network size, memory length, and richness of the training task distribution. We believe our results lay the foundation for increasingly general and adaptive RL agents that perform well across ever-larger open-ended domains.
Forward citations
Cited by 14 Pith papers
-
RoboTTT: Context Scaling for Robot Policies
A robot policy that updates its own weights during deployment can use 8,000 steps of history, steadily improving as context grows and enabling one-shot imitation from human videos.
-
Cross-Entropy Games for Language Models: From Implicit Knowledge to General Capability Measures
Xent Games formalize a large family of LLM evaluation tasks as games whose rewards and constraints are signed cross-entropy sums, and propose using them to build general capability measures.
-
Syllabus: Portable Curricula for Reinforcement Learning Agents
Syllabus provides a portable curriculum learning library with a unified API, reproduces prior baselines, and shows that standard automatic curricula do not transfer to NetHack and Neural MMO.
-
Training AI Scientists to Replicate Research
A 27B-parameter post-trained agent, Faraday, outperforms frontier coding agents at replicating held-out research figures by directing a larger coding model as a tool.
-
HERAKLES: Hierarchical Skill Compilation for Open-ended LLM Agents
HERAKLES couples a language-model planner to a small, continually retrained skill executor and outperforms three baselines on the 17-goal Crafter benchmark, scaling better to reworded and repeated goals.
-
Evolution and The Knightian Blindspot of Machine Learning
ML's formalisms, particularly RL's, exclude Knightian uncertainty, and evolution's diversify-and-filter mechanisms point toward a direct remedy.
-
ObjVariantEnsemble: Advancing Point Cloud LLM Evaluation in Challenging Scenes with Subtly Distinguished Objects
ObjVariantEnsemble constructs challenging 3D scenes with adversarial distractors and uses LLM-VLM annotations to reveal that current 3D grounding models are weakest at pure spatial reasoning.
-
From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons
A single MLLM-based agent, finetuned with cross-domain supervision and online RL, achieves strong zero-shot generalization across manipulation, navigation, games, UI control, and planning.
-
Parseval Regularization for Continual Reinforcement Learning
Parseval regularization, a cheap orthogonality-preserving penalty, improves continual RL agents' success on new tasks across gridworld, CARL and MetaWorld benchmarks.
-
BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
A benchmark of six long-horizon game environments shows current LLMs and VLMs struggle on hard tasks and often do worse when given images.
-
AMAGO-2: Breaking the Multi-Task Barrier in Meta-Reinforcement Learning with Transformers
Using two-hot classification for value prediction and binary-filtered imitation for policy updates makes multi-task meta-RL training scale-invariant to reward magnitudes, improving performance across five benchmarks w...
-
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments
PPO plateaus can be avoided by increasing the number of parallel environments, which reduces both the outer-loop step size and update noise; scaling to 1M environments sustained improvement to 1T transitions.
-
Unraveling the Hidden Dynamical Structure in Recurrent Neural Policies
Recurrent neural policies trained on episodic tasks converge to stable cyclic attractors in hidden state, and the geometry of these cycles mirrors behavior structure.
-
Training Cross-Morphology Embodied AI Agents: From Practical Challenges to Theoretical Foundations
The paper proves that the cross-morphology robot training problem HEAT is PSPACE-complete by embedding any POMDP into a HEAT instance with a single morphology.
Discussion (0). Continue with ORCID to comment.