Pith. sign in

REVIEW 2 cited by

Self-Correcting Models for Model-Based Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1612.06018 v2 pith:AVGVY7LV submitted 2016-12-19 cs.LG cs.AI

classification cs.LGcs.AI
keywords modelmbrlwhenerrorerrorsflawedlearningmodel-based
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

When an agent cannot represent a perfectly accurate model of its environment's dynamics, model-based reinforcement learning (MBRL) can fail catastrophically. Planning involves composing the predictions of the model; when flawed predictions are composed, even minor errors can compound and render the model useless for planning. Hallucinated Replay (Talvitie 2014) trains the model to "correct" itself when it produces errors, substantially improving MBRL with flawed models. This paper theoretically analyzes this approach, illuminates settings in which it is likely to be effective or ineffective, and presents a novel error bound, showing that a model's ability to self-correct is more tightly related to MBRL performance than one-step prediction error. These results inspire an MBRL algorithm for deterministic MDPs with performance guarantees that are robust to model class limitations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. World Models in Pieces: Structural Certification for General Agents

    cs.AI 2026-06 unverdicted novelty 7.0 of 10

    Structural certification maps bounded goal-conditioned performance to O(1/n) + O(δ) entry-wise error bounds on an agent's internal world model for transitions filtered by deep compositional goals.

  2. Beyond Myopic World Models: Long-Horizon End-to-End Training for Direct Future Prediction

    cs.LG 2026-08 conditional novelty 6.0 of 10

    A direct endpoint-prediction world model trained on long-horizon objectives substantially beats recursively-rolled-out baselines, and the objective, not the backbone, drives the improvement.

Pith tools