Pith. sign in

REVIEW 3 cited by

Objective Mismatch in Model-based Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2002.04523 v3 pith:SPEC4S6G submitted 2020-02-11 cs.LG cs.ROstat.ML

classification cs.LGcs.ROstat.ML
keywords issuemismatchobjectivembrlcontrolframeworkaccuratedynamics
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Model-based reinforcement learning (MBRL) has been shown to be a powerful framework for data-efficiently learning control of continuous tasks. Recent work in MBRL has mostly focused on using more advanced function approximators and planning schemes, with little development of the general framework. In this paper, we identify a fundamental issue of the standard MBRL framework -- what we call the objective mismatch issue. Objective mismatch arises when one objective is optimized in the hope that a second, often uncorrelated, metric will also be optimized. In the context of MBRL, we characterize the objective mismatch between training the forward dynamics model w.r.t.~the likelihood of the one-step ahead prediction, and the overall goal of improving performance on a downstream control task. For example, this issue can emerge with the realization that dynamics models effective for a specific task do not necessarily need to be globally accurate, and vice versa globally accurate models might not be sufficiently accurate locally to obtain good control performance on a specific task. In our experiments, we study this objective mismatch issue and demonstrate that the likelihood of one-step ahead predictions is not always correlated with control performance. This observation highlights a critical limitation in the MBRL framework which will require further research to be fully understood and addressed. We propose an initial method to mitigate the mismatch issue by re-weighting dynamics model training. Building on it, we conclude with a discussion about other potential directions of research for addressing this issue.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When a Verified World Model Still Loses: Play-Adequacy vs Prediction-Accuracy in LLM-Synthesized Code World Models

    cs.AI 2026-07 accept novelty 7.0 of 10

    A synthesized game model can pass a 100%-accurate transition gate yet be systematically outplayed when the missed dynamics are rare under random play but pivotal under competent play.

  2. Policy-shaped prediction: avoiding distractions in model-based reinforcement learning

    cs.LG 2024-12 conditional novelty 6.0 of 10

    Policy-Shaped Prediction weights a world model's reconstruction loss by policy-gradient salience aggregated through SAM segmentation, plus an adversarial action head, improving MBRL robustness to learnable distractors.

  3. SimuDICE: Offline Policy Optimization Through World Model Updates and DICE Estimation

    cs.LG 2024-12 conditional novelty 5.0 of 10

    SimuDICE uses DualDICE weights and model confidence to re-sample synthetic transitions in a tabular world model, improving offline Dyna-Q style policy optimization in small discrete environments.

Pith tools