Pith. sign in

REVIEW 12 cited by

A Tutorial on Meta-Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.08028 v4 pith:7RAXB4LP submitted 2023-01-19 cs.LG

classification cs.LG
keywords meta-rllearningtaskdistributionproblemalgorithmsdatadeep
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While deep reinforcement learning (RL) has fueled multiple high-profile successes in machine learning, it is held back from more widespread adoption by its often poor data efficiency and the limited generality of the policies it produces. A promising approach for alleviating these limitations is to cast the development of better RL algorithms as a machine learning problem itself in a process called meta-RL. Meta-RL is most commonly studied in a problem setting where, given a distribution of tasks, the goal is to learn a policy that is capable of adapting to any new task from the task distribution with as little data as possible. In this survey, we describe the meta-RL problem setting in detail as well as its major variations. We discuss how, at a high level, meta-RL research can be clustered based on the presence of a task distribution and the learning budget available for each individual task. Using these clusters, we then survey meta-RL algorithms and applications. We conclude by presenting the open problems on the path to making meta-RL part of the standard toolbox for a deep RL practitioner.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 48 citations worldwide. Full citation record

  1. A Unified Causal-Origin Taxonomy of Distributional Shifts in Reinforcement Learning

    cs.LG 2026-06 conditional novelty 6.5 of 10

    Distributional shift in RL is classified by which POMDP generative component changes (internal agent vs external environment) and by whether the time boundary is explicit, implicit, or hybrid.

  2. Learning Adaptive Multi-Task Guidance, Navigation, and Control via Hypernetworks

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A hypernetwork maps continuous physics-informed task embeddings to shared actor-critic weights, mastering four orbital GNC tasks and composing novel ones without retraining, with sim-to-real on a floating platform.

  3. Action Chunking with Transformers for Image-Based Spacecraft Guidance and Control

    cs.RO 2025-09 conditional novelty 6.0 of 10

    ACT imitation learning from 100 meta-RL demonstrations beats the meta-RL baseline on simulated ISS docking, using about 6,300 interactions instead of 40 million.

  4. How Should We Meta-Learn Reinforcement Learning Algorithms?

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A systematic comparison of black-box evolution, neural and symbolic distillation, and LLM-based proposal for meta-learning RL algorithms yields practical recommendations: warm-started LLM proposal is sample-efficient,...

  5. Behavioral Exploration: Learning to Explore via In-Context Adaptation

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A coverage-conditioned behavioral cloning policy adapts in-context to its own history, making robots explore new expert-like behaviors online without online reinforcement learning.

  6. Robust In-Context Reinforcement Learning Under Reward Poisoning Attacks

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Adversarially training a Decision-Pretrained Transformer against learned reward-poisoning attackers makes it robust to test-time reward corruption, outperforming robust bandit baselines in experiments.

  7. Unsupervised Meta-Testing with Conditional Neural Processes for Hybrid Meta-Reinforcement Learning

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A hybrid meta-RL method uses offline-trained conditional neural processes to generate extra rollouts, enabling reward-free adaptation to an unseen task from a single real rollout.

  8. Fully Offline Reinforcement Learning

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Bayesian offline RL can select hyperparameters and estimate deployed regret from posterior predictive uncertainty, with a regret bound at the parametric minimax rate.

  9. Filtering Learning Histories Enhances In-Context Reinforcement Learning

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Filtering ICRL pretraining datasets by a simple improvement-and-stability score boosts downstream in-context learning performance across AD, DICP, and DPT baselines.

  10. ReBRAC-v2: The Return of the King

    cs.LG 2026-08 conditional novelty 5.0 of 10

    A fixed-recipe offline RL method combining normalizing-flow actors, categorical critics, staged training, and test-time refinement beats recent flow-based baselines by 22.5 points averaged over ten OGBench categories.

  11. STMA: A Spatio-Temporal Memory Agent for Long-Horizon Embodied Task Planning

    cs.AI 2025-02 conditional novelty 5.0 of 10

    A spatio-temporal memory agent combining a textual history summarizer, a spatial knowledge graph, and a planner-critic loop outperforms ReAct, Reflexion, and AdaPlanner on TextWorld cooking tasks.

  12. Single-Agent Planning in a Multi-Agent System: A Unified Framework for Type-Based Planners

    cs.MA 2025-02 conditional novelty 5.0 of 10

    A layered tree-search framework unifies type-based opponent-modelling planners, and myopic safe-agents emerge as the strongest practical choice in a large multi-agent route planning benchmark.

Pith tools