Pith. sign in

REVIEW 5 cited by

When to Trust Your Model: Model-Based Policy Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1906.08253 v3 pith:F65TGKWH submitted 2019-06-19 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords model-baseddatamodelalgorithmsanalysismodel-generatedlearningmethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Designing effective model-based reinforcement learning algorithms is difficult because the ease of data generation must be weighed against the bias of model-generated data. In this paper, we study the role of model usage in policy optimization both theoretically and empirically. We first formulate and analyze a model-based reinforcement learning algorithm with a guarantee of monotonic improvement at each step. In practice, this analysis is overly pessimistic and suggests that real off-policy data is always preferable to model-generated on-policy data, but we show that an empirical estimate of model generalization can be incorporated into such analysis to justify model usage. Motivated by this analysis, we then demonstrate that a simple procedure of using short model-generated rollouts branched from real data has the benefits of more complicated model-based algorithms without the usual pitfalls. In particular, this approach surpasses the sample efficiency of prior model-based methods, matches the asymptotic performance of the best model-free algorithms, and scales to horizons that cause other model-based methods to fail entirely.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When Is a Learned Command Adapter Worth It? Closed-Loop Identification and Counterfactual Auditing of Frozen Locomotion Policies

    cs.AI 2026-07 conditional novelty 7.0 of 10

    A counterfactual audit separates same-state headroom from recoverable state-allocation gain, returning NO-GO or ABSTAIN for learned command adapters on frozen Go2 and H1 locomotion policies at 1% thresholds.

  2. Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A taxonomy of robot learning on a weights-versus-skills axis, with a five-rung self-improvement ladder whose top cell (feedback plus memory plus search) holds only a few recent systems.

  3. Quo Vadis, World Modeling?

    cs.CV 2026-08 conditional novelty 5.0 of 10

    An agent-centric reframing of world modeling, replacing physical state prediction with 'information transitions' organized into six proxy functions and three empowerment levels.

  4. A Comprehensive Review of Reinforcement Learning for Autonomous Driving in the CARLA Simulator

    cs.RO 2025-09 conditional novelty 4.0 of 10

    A survey of roughly 100 CARLA reinforcement learning papers, mapping algorithm families, representations, rewards, evaluation metrics, towns, and open challenges.

  5. Model-free Reinforcement Learning for Model-based Control: Towards Safe, Interpretable and Sample-efficient Agents

    cs.LG 2025-07 conditional novelty 3.0 of 10

    A perspective paper argues that model predictive control can be used as a learned policy in model-free reinforcement learning and reviews the methods and open problems.

Pith tools