Pith. sign in

REVIEW 6 cited by

Why Does Hierarchy (Sometimes) Work So Well in Reinforcement Learning?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.10618 v2 pith:TCNJDQHS submitted 2019-09-23 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords hierarchicallearningbenefitshierarchyreinforcementexplorationexploringextended
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Hierarchical reinforcement learning has demonstrated significant success at solving difficult reinforcement learning (RL) tasks. Previous works have motivated the use of hierarchy by appealing to a number of intuitive benefits, including learning over temporally extended transitions, exploring over temporally extended periods, and training and exploring in a more semantically meaningful action space, among others. However, in fully observed, Markovian settings, it is not immediately clear why hierarchical RL should provide benefits over standard "shallow" RL architectures. In this work, we isolate and evaluate the claimed benefits of hierarchical RL on a suite of tasks encompassing locomotion, navigation, and manipulation. Surprisingly, we find that most of the observed benefits of hierarchy can be attributed to improved exploration, as opposed to easier policy learning or imposed hierarchical structures. Given this insight, we present exploration techniques inspired by hierarchy that achieve performance competitive with hierarchical RL while at the same time being much simpler to use and implement.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Efficient Robotic Policy Learning via Latent Space Backward Planning

    cs.RO 2025-05 conditional novelty 6.0 of 10

    Planning backward from a predicted final latent goal, instead of forward into the future, reduces error accumulation in long-horizon robot manipulation and outperforms prior planning methods on LIBERO-LONG.

  2. LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback

    cs.AI 2026-07 conditional novelty 5.0 of 10

    LEMUR jointly learns a separate reward model for each teacher's preferences and uses them to train a population of multi-objective policies, beating baselines that merge feedback into one reward.

  3. Unsupervised Hierarchical Skill Discovery

    cs.LG 2026-01 conditional novelty 5.0 of 10

    An observation-only pipeline that combines optimal-transport skill segmentation (ASOT) with Sequitur grammar induction produces reusable skill hierarchies that accelerate reinforcement learning in Craftax and Minecraft.

  4. Dynamic Legged Ball Manipulation on Rugged Terrains with Hierarchical Reinforcement Learning

    cs.RO 2025-04 reject novelty 5.0 of 10

    A hierarchical reinforcement learning controller switches between dribbling and locomotion skills to move a ball across rugged terrain.

  5. A survey on intrinsic motivation in reinforcement learning

    cs.LG 2019-08 accept novelty 4.0 of 10

    A survey that classifies intrinsic motivation methods in deep RL as knowledge acquisition or skill learning and proposes their unification through information compression.

  6. Hierarchical Reinforcement Learning in Multi-Goal Spatial Navigation with Autonomous Mobile Robots

    cs.AI 2025-04 conditional novelty 3.0 of 10

    In Webots-simulated maze navigation, Option-Critic HRL converged faster than PPO on two harder mazes, but the paper's evidence that sub-goals cause this advantage is compromised by confounded ablations.

Pith tools