Pith. sign in

REVIEW 3 cited by

Divide-and-Conquer Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1711.09874 v2 pith:DCVKGGOD submitted 2017-11-27 cs.LG cs.RO

classification cs.LGcs.RO
keywords statedivide-and-conquerinitialmethodspoliciespolicyalgorithmchallenging
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Standard model-free deep reinforcement learning (RL) algorithms sample a new initial state for each trial, allowing them to optimize policies that can perform well even in highly stochastic environments. However, problems that exhibit considerable initial state variation typically produce high-variance gradient estimates for model-free RL, making direct policy or value function optimization challenging. In this paper, we develop a novel algorithm that instead partitions the initial state space into "slices", and optimizes an ensemble of policies, each on a different slice. The ensemble is gradually unified into a single policy that can succeed on the whole state space. This approach, which we term divide-and-conquer RL, is able to solve complex tasks where conventional deep RL methods are ineffective. Our results show that divide-and-conquer RL greatly outperforms conventional policy gradient methods on challenging grasping, manipulation, and locomotion tasks, and exceeds the performance of a variety of prior methods. Videos of policies learned by our algorithm can be viewed at http://bit.ly/dnc-rl

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Distributionally Robust Multi-Task Reinforcement Learning via Adaptive Task Sampling

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    DRATS derives a minimax objective from a feasibility formulation of MTRL to adaptively sample tasks with the largest return gaps, leading to better worst-task performance on MetaWorld benchmarks.

  2. Prescribe-then-Select: Adaptive Policy Selection for Contextual Stochastic Optimization

    cs.LG 2025-09 conditional novelty 6.0 of 10

    A decision-tree ensemble selects, per context, the best candidate policy from a library, and is shown to beat the best single policy on synthetic newsvendor and shipment problems.

  3. Mesh-RL: Coupled subgrid reinforcement learning

    cs.LG 2026-06 unverdicted novelty 5.0 of 10

    Mesh-RL applies finite-element-inspired domain decomposition with overlapping subgrids to accelerate temporal-difference learning across distant states in grid-world environments.

Pith tools