Pith. sign in

REVIEW 16 cited by

Model-Based Reinforcement Learning for Atari

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1903.00374 v5 pith:5AGTEDSD submitted 2019-03-01 cs.LG stat.ML

classification cs.LGstat.ML
keywords gamesatarilearnlearningmodel-freesimpleinteractionsmodel-based
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Model-free reinforcement learning (RL) can be used to learn effective policies for complex tasks, such as Atari games, even from image observations. However, this typically requires very large amounts of interaction -- substantially more, in fact, than a human would need to learn the same games. How can people learn so quickly? Part of the answer may be that people can learn how the game works and predict which actions will lead to desirable outcomes. In this paper, we explore how video prediction models can similarly enable agents to solve Atari games with fewer interactions than model-free methods. We describe Simulated Policy Learning (SimPLe), a complete model-based deep RL algorithm based on video prediction models and present a comparison of several model architectures, including a novel architecture that yields the best results in our setting. Our experiments evaluate SimPLe on a range of Atari games in low data regime of 100k interactions between the agent and the environment, which corresponds to two hours of real-time play. In most games SimPLe outperforms state-of-the-art model-free algorithms, in some games by over an order of magnitude.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. World Modeling with Probabilistic Structure Integration

    cs.CV 2025-09 conditional novelty 7.0 of 10

    A single probabilistic video model extracts optical flow, depth, and segments via counterfactual prompts, then integrates those structures as new token types to improve its own video predictions.

  2. Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A taxonomy of robot learning on a weights-versus-skills axis, with a five-rung self-improvement ladder whose top cell (feedback plus memory plus search) holds only a few recent systems.

  3. Endpoint Replay: Compressing the Recency Buffer in Deep Reinforcement Learning

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Chained n-step endpoint coresets plus expectile Sarsa match million-transition DQN buffers at 10–50× less storage by keeping bootstrap targets anchored.

  4. SoK: The Pitfalls of Deep Reinforcement Learning for Cybersecurity

    cs.LG 2026-02 accept novelty 6.0 of 10

    Across 66 DRL-for-cybersecurity papers, the authors identify 11 recurring methodological pitfalls—averaging 5.8 per paper—and demonstrate their impact in four environments.

  5. Non-differentiable Reward Optimization for Diffusion-based Autonomous Motion Planning

    cs.RO 2025-07 conditional novelty 6.0 of 10

    A reinforcement learning fine-tuning method with dynamic reward thresholding lets diffusion motion planners directly optimize non-differentiable safety and goal-reaching metrics, improving collision rate and success r...

  6. Chargax: A JAX Accelerated EV Charging Simulator

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Chargax is a JAX-based EV charging simulator that accelerates reinforcement learning training by 100x to 1000x compared to existing environments, with modular real-world scenarios.

  7. Time-Aware World Model for Adaptive Prediction and Control

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A time-conditioned world model trained on mixed time steps matches or beats a fixed-time-step baseline at its native rate and far exceeds it at slower observation rates, using the same sample count.

  8. Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One

    cs.LG 2025-05 reject novelty 6.0 of 10

    An LLM classifies game states into situations and selects the RL agent with the best historical average reward for each situation, outperforming static ensemble baselines on Atari.

  9. Quo Vadis, World Modeling?

    cs.CV 2026-08 conditional novelty 5.0 of 10

    An agent-centric reframing of world modeling, replacing physical state prediction with 'information transitions' organized into six proxy functions and three empowerment levels.

  10. Relative Value Learning

    cs.LG 2026-07 conditional novelty 5.0 of 10

    A critic that learns antisymmetric value differences ∆(s_i,s_j)=V(s_i)−V(s_j) has a provably contracting Bellman operator and an unbiased advantage estimator, and PPO with this critic matches standard PPO on Atari.

  11. HypEMBER: Hypernetwork-based Ensemble for Robust Policy Learning of Parametrized Dynamical Systems

    cs.LG 2026-07 conditional novelty 5.0 of 10

    HypEMBER joins hypernetwork-generated policies with an ensemble critic to improve robustness of reinforcement-learning controllers for parametrized dynamical systems.

  12. Assessing Adaptive World Models in Machines with Novel Games

    cs.AI 2025-07 conditional novelty 5.0 of 10

    The paper proposes a framework called world model induction and a novel-game benchmark paradigm for evaluating rapid adaptation in AI.

  13. Where to Intervene: Action Selection in Deep Reinforcement Learning

    stat.ML 2025-07 conditional novelty 5.0 of 10

    Knockoff sampling selects the minimal sufficient action set during online deep reinforcement learning with false discovery rate control.

  14. STMA: A Spatio-Temporal Memory Agent for Long-Horizon Embodied Task Planning

    cs.AI 2025-02 conditional novelty 5.0 of 10

    A spatio-temporal memory agent combining a textual history summarizer, a spatial knowledge graph, and a planner-critic loop outperforms ReAct, Reflexion, and AdaPlanner on TextWorld cooking tasks.

  15. Mask-based Predictive Representations for Reinforcement Learning

    cs.LG 2026-07 conditional novelty 4.0 of 10

    Mask-based predictive representations (MPR) as an auxiliary self-supervised task improve sample efficiency of vision-based RL over prior SOTA on continuous and discrete control benchmarks.

  16. Reinforcement Learning: From Algorithms To Foundation Models

    cs.AI 2026-07 conditional novelty 3.0 of 10

    A dissertation uniting the author's published results: non-exploitable Nash-DQN policies and the FightLadder benchmark for games, plus diffusion/consistency-model world models for RL — a compilation rather than new results.

Pith tools