REVIEW 1 cited by
Efficient decorrelation of features using Gramian in Reinforcement Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Learning good representations is a long standing problem in reinforcement learning (RL). One of the conventional ways to achieve this goal in the supervised setting is through regularization of the parameters. Extending some of these ideas to the RL setting has not yielded similar improvements in learning. In this paper, we develop an online regularization framework for decorrelating features in RL and demonstrate its utility in several test environments. We prove that the proposed algorithm converges in the linear function approximation setting and does not change the main objective of maximizing cumulative reward. We demonstrate how to scale the approach to deep RL using the Gramian of the features achieving linear computational complexity in the number of features and squared complexity in size of the batch. We conduct an extensive empirical study of the new approach on Atari 2600 games and show a significant improvement in sample efficiency in 40 out of 49 games.
Forward citations
Cited by 1 Pith paper
-
Decorrelated Soft Actor-Critic for Efficient Deep Reinforcement Learning
Decorrelated Soft Actor-Critic (DSAC) adds layerwise input decorrelation to discrete SAC and reports faster wall-clock training in 5 of 7 Atari games and better reward in 2, though the gains are partly confounded by p...
Discussion (0). Continue with ORCID to comment.