REVIEW 2 cited by
k-Means Maximum Entropy Exploration
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Exploration in high-dimensional, continuous spaces with sparse rewards is an open problem in reinforcement learning. Artificial curiosity algorithms address this by creating rewards that lead to exploration. Given a reinforcement learning algorithm capable of maximizing rewards, the problem reduces to finding an optimization objective consistent with exploration. Maximum entropy exploration uses the entropy of the state visitation distribution as such an objective. However, efficiently estimating the entropy of the state visitation distribution is challenging in high-dimensional, continuous spaces. We introduce an artificial curiosity algorithm based on lower bounding an approximation to the entropy of the state visitation distribution. The bound relies on a result we prove for non-parametric density estimation in arbitrary dimensions using k-means. We show that our approach is both computationally efficient and competitive on benchmarks for exploration in high-dimensional, continuous spaces, especially on tasks where reinforcement learning algorithms are unable to find rewards.
Forward citations
Cited by 2 Pith papers
-
ELEMENT: Episodic and Lifelong Exploration via Maximum Entropy
ELEMENT combines an average episodic state entropy reward with a kNN-graph lifelong entropy reward for reward-free RL exploration.
-
Enhancing Diversity in Parallel Agents: A Maximum State Entropy Exploration Story
A centralized policy gradient for parallel state entropy maximization improves state coverage on small gridworlds, but the paper's concentration-rate proof is invalid.
Discussion (0). Continue with ORCID to comment.