Pith. sign in

REVIEW 2 cited by

Learning Representations in Model-Free Hierarchical Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1810.10096 v3 pith:LT2FCGFP submitted 2018-10-23 cs.AI cs.LGmath.OC

classification cs.AIcs.LGmath.OC
keywords learningenvironmentsubgoalsmethodmodelreinforcementsubgoalabstraction
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Common approaches to Reinforcement Learning (RL) are seriously challenged by large-scale applications involving huge state spaces and sparse delayed reward feedback. Hierarchical Reinforcement Learning (HRL) methods attempt to address this scalability issue by learning action selection policies at multiple levels of temporal abstraction. Abstraction can be had by identifying a relatively small set of states that are likely to be useful as subgoals, in concert with the learning of corresponding skill policies to achieve those subgoals. Many approaches to subgoal discovery in HRL depend on the analysis of a model of the environment, but the need to learn such a model introduces its own problems of scale. Once subgoals are identified, skills may be learned through intrinsic motivation, introducing an internal reward signal marking subgoal attainment. In this paper, we present a novel model-free method for subgoal discovery using incremental unsupervised learning over a small memory of the most recent experiences (trajectories) of the agent. When combined with an intrinsic motivation learning mechanism, this method learns both subgoals and skills, based on experiences in the environment. Thus, we offer an original approach to HRL that does not require the acquisition of a model of the environment, suitable for large-scale applications. We demonstrate the efficiency of our method on two RL problems with sparse delayed feedback: a variant of the rooms environment and the first screen of the ATARI 2600 Montezuma's Revenge game.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Quasi-Newton Optimization Methods For Deep Learning Applications

    cs.LG 2019-09 reject novelty 4.0 of 10

    L-BFGS with line search or trust region is applied to deep learning and deep Q-learning, with convergence theorems that rely on strong convexity and mixed empirical results on MNIST and Atari.

  2. Learning sparse representations in reinforcement learning

    cs.LG 2019-09 conditional novelty 3.0 of 10

    Adding a k-winners-take-all sparsity mechanism to the hidden layer of a TD-learning network improves performance on three classic control tasks compared to standard backpropagation and linear networks.

Pith tools