Pith. sign in

REVIEW 3 cited by

Exploration and Anti-Exploration with Distributional Random Network Distillation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.09750 v4 pith:6XKVO3SZ submitted 2024-01-18 cs.LG

classification cs.LG
keywords explorationissuebonusdrndrandomalgorithmallocationanti-exploration
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Exploration remains a critical issue in deep reinforcement learning for an agent to attain high returns in unknown environments. Although the prevailing exploration Random Network Distillation (RND) algorithm has been demonstrated to be effective in numerous environments, it often needs more discriminative power in bonus allocation. This paper highlights the "bonus inconsistency" issue within RND, pinpointing its primary limitation. To address this issue, we introduce the Distributional RND (DRND), a derivative of the RND. DRND enhances the exploration process by distilling a distribution of random networks and implicitly incorporating pseudo counts to improve the precision of bonus allocation. This refinement encourages agents to engage in more extensive exploration. Our method effectively mitigates the inconsistency issue without introducing significant computational overhead. Both theoretical analysis and experimental results demonstrate the superiority of our approach over the original RND algorithm. Our method excels in challenging online exploration scenarios and effectively serves as an anti-exploration mechanism in D4RL offline tasks. Our code is publicly available at https://github.com/yk7333/DRND.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Offline Reinforcement Learning with Penalized Action Noise Injection

    cs.LG 2025-07 conditional novelty 5.0 of 10

    Injecting noise-perturbed actions into offline Q-learning with a distance penalty improves D4RL performance over IQL and TD3 baselines, formalized as Q-learning in a Noisy Action MDP.

  2. Information-Based Exploration via Random Features for Reinforcement Learning

    cs.LG 2026-07 conditional novelty 4.0 of 10

    Random-feature Gaussian-process information gain is turned into a closed-form exploration bonus for PPO that matches RND/VIME/#Explo on 12 control, navigation, and sparse-locomotion tasks, with error bounds on the app...

  3. Mixture of Autoencoder Experts Guidance using Unlabeled and Incomplete Data for Exploration in Reinforcement Learning

    cs.LG 2025-07 conditional novelty 4.0 of 10

    MoE-GUIDE guides RL exploration by rewarding states that a mixture of autoencoders, trained on sparse state-only expert demonstrations, considers similar to expert data.

Pith tools