Pith. sign in

REVIEW 1 cited by

Learning Dynamics Models for Model Predictive Agents

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.14311 v1 pith:RLZUH3FV submitted 2021-09-29 cs.LG cs.RO

classification cs.LGcs.RO
keywords modeldynamicslearningmodelschoicesdesignplanningdifferent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Model-Based Reinforcement Learning involves learning a \textit{dynamics model} from data, and then using this model to optimise behaviour, most often with an online \textit{planner}. Much of the recent research along these lines presents a particular set of design choices, involving problem definition, model learning and planning. Given the multiple contributions, it is difficult to evaluate the effects of each. This paper sets out to disambiguate the role of different design choices for learning dynamics models, by comparing their performance to planning with a ground-truth model -- the simulator. First, we collect a rich dataset from the training sequence of a model-free agent on 5 domains of the DeepMind Control Suite. Second, we train feed-forward dynamics models in a supervised fashion, and evaluate planner performance while varying and analysing different model design choices, including ensembling, stochasticity, multi-step training and timestep size. Besides the quantitative analysis, we describe a set of qualitative findings, rules of thumb, and future research directions for planning with learned dynamics models. Videos of the results are available at https://sites.google.com/view/learning-better-models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Time-Aware World Model for Adaptive Prediction and Control

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A time-conditioned world model trained on mixed time steps matches or beats a fixed-time-step baseline at its native rate and far exceeds it at slower observation rates, using the same sample count.

Pith tools