Pith. sign in

REVIEW 1 cited by

Hierarchical Reinforcement Learning Based on Planning Operators

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.14237 v2 pith:LEKVGRUB submitted 2023-09-25 cs.RO

classification cs.RO
keywords planninglearninghierarchicalhigh-leveloperatorspoliciessequenceapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Long-horizon manipulation tasks such as stacking represent a longstanding challenge in the field of robotic manipulation, particularly when using reinforcement learning (RL) methods which often struggle to learn the correct sequence of actions for achieving these complex goals. To learn this sequence, symbolic planning methods offer a good solution based on high-level reasoning, however, planners often fall short in addressing the low-level control specificity needed for precise execution. This paper introduces a novel framework that integrates symbolic planning with hierarchical RL through the cooperation of high-level operators and low-level policies. Our contribution integrates planning operators (e.g. preconditions and effects) as part of the hierarchical RL algorithm based on the Scheduled Auxiliary Control (SAC-X) method. We developed a dual-purpose high-level operator, which can be used both in holistic planning and as independent, reusable policies. Our approach offers a flexible solution for long-horizon tasks, e.g., stacking a cube. The experimental results show that our proposed method obtained an average of 97.2% success rate for learning and executing the whole stack sequence, and the success rate for learning independent policies, e.g. reach (98.9%), lift (99.7%), stack (85%), etc. The training time is also reduced by 68% when using our proposed approach.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Decoupled Hierarchical Reinforcement Learning with State Abstraction for Discrete Grids

    cs.LG 2025-06 conditional novelty 4.0 of 10

    A decoupled hierarchical RL framework with a rule-based low-level policy and DeepMDP state abstraction outperforms PPO on two custom discrete grid environments, but with a single baseline and sparse experimental detail.

Pith tools