Pith. sign in

REVIEW 4 cited by

Model-Based Reinforcement Learning via Meta-Policy Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1809.05214 v1 pith:4AW2EDTE submitted 2018-09-14 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords dynamicsmodel-basedensemblelearningmb-mpomodelmodelsapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Model-based reinforcement learning approaches carry the promise of being data efficient. However, due to challenges in learning dynamics models that sufficiently match the real-world dynamics, they struggle to achieve the same asymptotic performance as model-free methods. We propose Model-Based Meta-Policy-Optimization (MB-MPO), an approach that foregoes the strong reliance on accurate learned dynamics models. Using an ensemble of learned dynamic models, MB-MPO meta-learns a policy that can quickly adapt to any model in the ensemble with one policy gradient step. This steers the meta-policy towards internalizing consistent dynamics predictions among the ensemble while shifting the burden of behaving optimally w.r.t. the model discrepancies towards the adaptation step. Our experiments show that MB-MPO is more robust to model imperfections than previous model-based approaches. Finally, we demonstrate that our approach is able to match the asymptotic performance of model-free methods while requiring significantly less experience.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Benchmarking Model-Based Reinforcement Learning

    cs.LG 2019-07 accept novelty 7.0 of 10

    Introduces a benchmark suite of over 18 MBRL environments, evaluates multiple algorithms under consistent settings, and identifies three core challenges: dynamics bottleneck, planning horizon dilemma, and early-termin...

  2. SimuDICE: Offline Policy Optimization Through World Model Updates and DICE Estimation

    cs.LG 2024-12 conditional novelty 5.0 of 10

    SimuDICE uses DualDICE weights and model confidence to re-sample synthetic transitions in a tabular world model, improving offline Dyna-Q style policy optimization in small discrete environments.

  3. Uncertainty-aware Model-based Policy Optimization

    cs.LG 2019-06 unverdicted novelty 5.0 of 10

    Introduces a framework that learns an uncertainty-aware dynamics model and optimizes the policy via automatic differentiation through the model, reporting competitive asymptotic performance with significantly lower sa...

  4. Calibrated Model-Based Deep Reinforcement Learning

    cs.LG 2019-06 unverdicted novelty 5.0 of 10

    Augmenting model-based RL agents with calibrated predictive uncertainties improves planning, sample efficiency, and exploration on continuous control tasks.

Pith tools