Pith. sign in

REVIEW 2 cited by

RLeXplore: Accelerating Research in Intrinsically-Motivated Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.19548 v2 pith:3J4QZLLT submitted 2024-05-29 cs.LG

classification cs.LG
keywords intrinsicrewardsrlexploreagentsdetailsextrinsicimplementationintrinsically-motivated
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Extrinsic rewards can effectively guide reinforcement learning (RL) agents in specific tasks. However, extrinsic rewards frequently fall short in complex environments due to the significant human effort needed for their design and annotation. This limitation underscores the necessity for intrinsic rewards, which offer auxiliary and dense signals and can enable agents to learn in an unsupervised manner. Although various intrinsic reward formulations have been proposed, their implementation and optimization details are insufficiently explored and lack standardization, thereby hindering research progress. To address this gap, we introduce RLeXplore, a unified, highly modularized, and plug-and-play framework offering reliable implementations of eight state-of-the-art intrinsic reward methods. Furthermore, we conduct an in-depth study that identifies critical implementation details and establishes well-justified standard practices in intrinsically-motivated RL. Our documentation, examples, and source code are available at https://github.com/RLE-Foundation/RLeXplore.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Information-Based Exploration via Random Features for Reinforcement Learning

    cs.LG 2026-07 conditional novelty 4.0 of 10

    Random-feature Gaussian-process information gain is turned into a closed-form exploration bonus for PPO that matches RND/VIME/#Explo on 12 control, navigation, and sparse-locomotion tasks, with error bounds on the app...

  2. Mixture of Autoencoder Experts Guidance using Unlabeled and Incomplete Data for Exploration in Reinforcement Learning

    cs.LG 2025-07 conditional novelty 4.0 of 10

    MoE-GUIDE guides RL exploration by rewarding states that a mixture of autoencoders, trained on sparse state-only expert demonstrations, considers similar to expert data.

Pith tools