Pith. sign in

REVIEW 5 cited by

VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.08348 v2 pith:AGO6S6VL submitted 2019-10-18 cs.LG stat.ML

classification cs.LGstat.ML
keywords environmentvaribaduncertaintybayes-adaptivebayes-optimaldeepduringexploration
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Trading off exploration and exploitation in an unknown environment is key to maximising expected return during learning. A Bayes-optimal policy, which does so optimally, conditions its actions not only on the environment state but on the agent's uncertainty about the environment. Computing a Bayes-optimal policy is however intractable for all but the smallest tasks. In this paper, we introduce variational Bayes-Adaptive Deep RL (variBAD), a way to meta-learn to perform approximate inference in an unknown environment, and incorporate task uncertainty directly during action selection. In a grid-world domain, we illustrate how variBAD performs structured online exploration as a function of task uncertainty. We further evaluate variBAD on MuJoCo domains widely used in meta-RL and show that it achieves higher online return than existing methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. In-Context World Modeling for Robotic Control

    cs.RO 2026-06 unverdicted novelty 7.0 of 10

    Prepending a few self-generated random interaction clips as context lets VLA policies identify novel camera viewpoints and morphologies at test time and outperform multi-view baselines without parameter updates.

  2. Behavioral Exploration: Learning to Explore via In-Context Adaptation

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A coverage-conditioned behavioral cloning policy adapts in-context to its own history, making robots explore new expert-like behaviors online without online reinforcement learning.

  3. Uncertainty Prioritized Experience Replay

    cs.LG 2025-06 conditional novelty 6.0 of 10

    UPER uses ensemble-based epistemic and aleatoric uncertainty to compute an information gain priority for experience replay, outperforming TD-error prioritization on Atari-57.

  4. Unsupervised Meta-Testing with Conditional Neural Processes for Hybrid Meta-Reinforcement Learning

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A hybrid meta-RL method uses offline-trained conditional neural processes to generate extra rollouts, enabling reward-free adaptation to an unseen task from a single real rollout.

  5. Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens

    cs.AI 2025-06 conditional novelty 5.0 of 10

    An offline RL policy can be improved at test time by inferring a latent belief over environment dynamics from past transitions and planning with model-based rollouts averaged over that belief.

Pith tools