Pith. sign in

REVIEW 2 cited by

Models, Pixels, and Rewards: Evaluating Design Trade-offs in Visual Model-Based Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.04603 v1 pith:F4PKQI2F submitted 2020-12-08 cs.LG

classification cs.LG
keywords performancedesignmodelstaskexplorationmethodsmodelperform
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Model-based reinforcement learning (MBRL) methods have shown strong sample efficiency and performance across a variety of tasks, including when faced with high-dimensional visual observations. These methods learn to predict the environment dynamics and expected reward from interaction and use this predictive model to plan and perform the task. However, MBRL methods vary in their fundamental design choices, and there is no strong consensus in the literature on how these design decisions affect performance. In this paper, we study a number of design decisions for the predictive model in visual MBRL algorithms, focusing specifically on methods that use a predictive model for planning. We find that a range of design decisions that are often considered crucial, such as the use of latent spaces, have little effect on task performance. A big exception to this finding is that predicting future observations (i.e., images) leads to significant task performance improvement compared to only predicting rewards. We also empirically find that image prediction accuracy, somewhat surprisingly, correlates more strongly with downstream task performance than reward prediction accuracy. We show how this phenomenon is related to exploration and how some of the lower-scoring models on standard benchmarks (that require exploration) will perform the same as the best-performing models when trained on the same training data. Simultaneously, in the absence of exploration, models that fit the data better usually perform better on the downstream task as well, but surprisingly, these are often not the same models that perform the best when learning and exploring from scratch. These findings suggest that performance and exploration place important and potentially contradictory requirements on the model.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Surprising Ineffectiveness of Pre-Trained Visual Representations for Model-Based Reinforcement Learning

    cs.LG 2024-11 conditional novelty 6.0 of 10

    In model-based reinforcement learning, frozen pre-trained visual representations do not improve sample efficiency or out-of-distribution generalization over representations learned from scratch.

  2. Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes

    cs.LG 2025-01 reject novelty 4.0 of 10

    FLEXplore combines an L2 dynamics loss with a Wasserstein-style critic loss, FGSM reward smoothing, and a mutual-information auxiliary reward to improve sample efficiency in parameterized-action MDPs.

Pith tools