REVIEW 12 cited by
A Tutorial on Meta-Reinforcement Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
While deep reinforcement learning (RL) has fueled multiple high-profile successes in machine learning, it is held back from more widespread adoption by its often poor data efficiency and the limited generality of the policies it produces. A promising approach for alleviating these limitations is to cast the development of better RL algorithms as a machine learning problem itself in a process called meta-RL. Meta-RL is most commonly studied in a problem setting where, given a distribution of tasks, the goal is to learn a policy that is capable of adapting to any new task from the task distribution with as little data as possible. In this survey, we describe the meta-RL problem setting in detail as well as its major variations. We discuss how, at a high level, meta-RL research can be clustered based on the presence of a task distribution and the learning budget available for each individual task. Using these clusters, we then survey meta-RL algorithms and applications. We conclude by presenting the open problems on the path to making meta-RL part of the standard toolbox for a deep RL practitioner.
Forward citations
Cited by 12 Pith papers
-
A Unified Causal-Origin Taxonomy of Distributional Shifts in Reinforcement Learning
Distributional shift in RL is classified by which POMDP generative component changes (internal agent vs external environment) and by whether the time boundary is explicit, implicit, or hybrid.
-
Learning Adaptive Multi-Task Guidance, Navigation, and Control via Hypernetworks
A hypernetwork maps continuous physics-informed task embeddings to shared actor-critic weights, mastering four orbital GNC tasks and composing novel ones without retraining, with sim-to-real on a floating platform.
-
Action Chunking with Transformers for Image-Based Spacecraft Guidance and Control
ACT imitation learning from 100 meta-RL demonstrations beats the meta-RL baseline on simulated ISS docking, using about 6,300 interactions instead of 40 million.
-
How Should We Meta-Learn Reinforcement Learning Algorithms?
A systematic comparison of black-box evolution, neural and symbolic distillation, and LLM-based proposal for meta-learning RL algorithms yields practical recommendations: warm-started LLM proposal is sample-efficient,...
-
Behavioral Exploration: Learning to Explore via In-Context Adaptation
A coverage-conditioned behavioral cloning policy adapts in-context to its own history, making robots explore new expert-like behaviors online without online reinforcement learning.
-
Robust In-Context Reinforcement Learning Under Reward Poisoning Attacks
Adversarially training a Decision-Pretrained Transformer against learned reward-poisoning attackers makes it robust to test-time reward corruption, outperforming robust bandit baselines in experiments.
-
Unsupervised Meta-Testing with Conditional Neural Processes for Hybrid Meta-Reinforcement Learning
A hybrid meta-RL method uses offline-trained conditional neural processes to generate extra rollouts, enabling reward-free adaptation to an unseen task from a single real rollout.
-
Fully Offline Reinforcement Learning
Bayesian offline RL can select hyperparameters and estimate deployed regret from posterior predictive uncertainty, with a regret bound at the parametric minimax rate.
-
Filtering Learning Histories Enhances In-Context Reinforcement Learning
Filtering ICRL pretraining datasets by a simple improvement-and-stability score boosts downstream in-context learning performance across AD, DICP, and DPT baselines.
-
ReBRAC-v2: The Return of the King
A fixed-recipe offline RL method combining normalizing-flow actors, categorical critics, staged training, and test-time refinement beats recent flow-based baselines by 22.5 points averaged over ten OGBench categories.
-
STMA: A Spatio-Temporal Memory Agent for Long-Horizon Embodied Task Planning
A spatio-temporal memory agent combining a textual history summarizer, a spatial knowledge graph, and a planner-critic loop outperforms ReAct, Reflexion, and AdaPlanner on TextWorld cooking tasks.
-
Single-Agent Planning in a Multi-Agent System: A Unified Framework for Type-Based Planners
A layered tree-search framework unifies type-based opponent-modelling planners, and myopic safe-agents emerge as the strongest practical choice in a large multi-agent route planning benchmark.
Discussion (0). Continue with ORCID to comment.