Pith. sign in

REVIEW 5 cited by

RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated Environments

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2002.12292 v2 pith:6QSYRDWR submitted 2020-02-27 cs.LG cs.AI

classification cs.LGcs.AI
keywords agentenvironmentsexplorationintrinsicprocedurally-generatedrewardmethodsrewards
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Exploration in sparse reward environments remains one of the key challenges of model-free reinforcement learning. Instead of solely relying on extrinsic rewards provided by the environment, many state-of-the-art methods use intrinsic rewards to encourage exploration. However, we show that existing methods fall short in procedurally-generated environments where an agent is unlikely to visit a state more than once. We propose a novel type of intrinsic reward which encourages the agent to take actions that lead to significant changes in its learned state representation. We evaluate our method on multiple challenging procedurally-generated tasks in MiniGrid, as well as on tasks with high-dimensional observations used in prior work. Our experiments demonstrate that this approach is more sample efficient than existing exploration methods, particularly for procedurally-generated MiniGrid environments. Furthermore, we analyze the learned behavior as well as the intrinsic reward received by our agent. In contrast to previous approaches, our intrinsic reward does not diminish during the course of training and it rewards the agent substantially more for interacting with objects that it can control.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 34 citations worldwide. Full citation record

  1. Beyond Single-Model Optimization: Preserving Plasticity in Continual Reinforcement Learning

    cs.LG 2026-04 unverdicted novelty 7.0 of 10

    TeLAPA preserves behaviorally diverse policy neighborhoods in a shared latent space, improving MiniGrid continual RL transfer, revisit recovery, and retention over single-model preservation.

  2. Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks

    cs.LG 2026-07 conditional novelty 6.0 of 10

    ProGPO adds a first-visit observation-coverage advantage only when an entire rollout group fails, improving group-based RL for long-horizon LLM agents on ALFWorld and WebShop.

  3. ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making

    cs.RO 2025-05 conditional novelty 6.0 of 10

    ManiTaskGen automatically generates diverse, feasible mobile manipulation tasks from any input scene, and uses them to benchmark and improve vision-language robot agents.

  4. Episodic Novelty Through Temporal Distance

    cs.LG 2025-01 conditional novelty 6.0 of 10

    An episodic intrinsic reward based on a contrastively learned temporal distance quasimetric improves exploration in sparse-reward Contextual MDPs.

  5. The impact of intrinsic rewards on exploration in Reinforcement Learning

    cs.AI 2025-01 conditional novelty 5.0 of 10

    An empirical MiniGrid study shows state-counting is best for low-dimensional observations, maximum entropy is more robust with images, and DIAYN skill learning does not aid exploration.

Pith tools