Pith. sign in

REVIEW 1 cited by

Curiosity-Driven Experience Prioritization via Density Estimation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1902.08039 v3 pith:BBTW6MZ2 submitted 2019-02-20 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords learningenvironmentproblemtrajectoriesachievedagentcuriosity-drivenexperience
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In Reinforcement Learning (RL), an agent explores the environment and collects trajectories into the memory buffer for later learning. However, the collected trajectories can easily be imbalanced with respect to the achieved goal states. The problem of learning from imbalanced data is a well-known problem in supervised learning, but has not yet been thoroughly researched in RL. To address this problem, we propose a novel Curiosity-Driven Prioritization (CDP) framework to encourage the agent to over-sample those trajectories that have rare achieved goal states. The CDP framework mimics the human learning process and focuses more on relatively uncommon events. We evaluate our methods using the robotic environment provided by OpenAI Gym. The environment contains six robot manipulation tasks. In our experiments, we combined CDP with Deep Deterministic Policy Gradient (DDPG) with or without Hindsight Experience Replay (HER). The experimental results show that CDP improves both performance and sample-efficiency of reinforcement learning agents, compared to state-of-the-art methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Improving Environment Novelty Quantification for Effective Unsupervised Environment Design

    cs.LG 2025-02 conditional novelty 6.0 of 10

    CENIE augments regret-based unsupervised environment design with a GMM-based novelty score derived from the student's state-action coverage, improving zero-shot transfer in Minigrid, BipedalWalker, and CarRacing.

Pith tools