Pith. sign in

REVIEW 7 cited by

Learning Multi-Level Hierarchies with Hindsight

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1712.00948 v5 pith:BUM2NKX7 submitted 2017-12-04 cs.AI cs.LGcs.NEcs.RO

classification cs.AIcs.LGcs.NEcs.RO
keywords hierarchicallevelslearninglevelpoliciesagentslearnmultiple
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Hierarchical agents have the potential to solve sequential decision making tasks with greater sample efficiency than their non-hierarchical counterparts because hierarchical agents can break down tasks into sets of subtasks that only require short sequences of decisions. In order to realize this potential of faster learning, hierarchical agents need to be able to learn their multiple levels of policies in parallel so these simpler subproblems can be solved simultaneously. Yet, learning multiple levels of policies in parallel is hard because it is inherently unstable: changes in a policy at one level of the hierarchy may cause changes in the transition and reward functions at higher levels in the hierarchy, making it difficult to jointly learn multiple levels of policies. In this paper, we introduce a new Hierarchical Reinforcement Learning (HRL) framework, Hierarchical Actor-Critic (HAC), that can overcome the instability issues that arise when agents try to jointly learn multiple levels of policies. The main idea behind HAC is to train each level of the hierarchy independently of the lower levels by training each level as if the lower level policies are already optimal. We demonstrate experimentally in both grid world and simulated robotics domains that our approach can significantly accelerate learning relative to other non-hierarchical and hierarchical methods. Indeed, our framework is the first to successfully learn 3-level hierarchies in parallel in tasks with continuous state and action spaces.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Combined Constrained Sampling and Reinforcement Learning for Robotic Manipulation

    cs.RO 2026-02 conditional novelty 7.0 of 10

    Guiding goal-conditioned reinforcement learning with samples from a constrained feasible-state manifold lets a simulated double-sphere and a Panda-arm policy succeed far more often than RL with random resets.

  2. Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A taxonomy of robot learning on a weights-versus-skills axis, with a five-rung self-improvement ladder whose top cell (feedback plus memory plus search) holds only a few recent systems.

  3. Hierarchical Residual Policy Optimization for Generative Recommendations

    cs.IR 2026-08 conditional novelty 6.0 of 10

    HRPO decomposes item-level rewards into token-level 'residual credits' along semantic identifier hierarchies and optimizes the generator with a PPO-style objective, improving session utility in KuaiSim and production ...

  4. S3: Stable Subgoal Selection by Constraining Uncertainty of Coarse Dynamics in Hierarchical Reinforcement Learning

    cs.LG 2026-07 conditional novelty 6.0 of 10

    S3 adds a high-level intrinsic reward that penalizes the predicted variance of coarse multi-step subgoal outcomes, improving HRL performance on bottleneck-heavy MuJoCo tasks.

  5. Vision-Based Obstacle Separation for Strawberry Harvesting in Clusters Using Hierarchical Reinforcement Learning

    cs.RO 2026-07 conditional novelty 4.0 of 10

    A hierarchical reinforcement-learning harvester that separates obstacle strawberries before grasping improves real-world strawberry-picking success from 59.7% to 80.0% versus direct picking, at a cost of 1.22 s extra ...

  6. Goal-conditioned Hierarchical Reinforcement Learning for Sample-efficient and Safe Autonomous Driving at Intersections

    cs.RO 2025-06 conditional novelty 4.0 of 10

    A hierarchical RL agent with a goal-conditioned collision prediction module achieves 94.7% success and 3.3% collisions in SMARTS intersection tasks, outperforming flat RL baselines.

  7. Hierarchical Soft Actor-Critic for Sparse-Reward Long-Horizon Reinforcement Learning

    cs.RO 2026-07 conditional novelty 3.5 of 10

    A two-level entropy-regularized hierarchical SAC agent beats flat SAC on a SAR-2-inspired sparse-reward continuous search task in reported success rate and coverage.

Pith tools