Pith. sign in

REVIEW 1 cited by

Off-Policy Actor-Critic with Shared Experience Replay

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.11583 v2 pith:3BD7ZFVR submitted 2019-09-25 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords experiencereplayactor-criticagentslearningdataoff-policypropose
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We investigate the combination of actor-critic reinforcement learning algorithms with uniform large-scale experience replay and propose solutions for two challenges: (a) efficient actor-critic learning with experience replay (b) stability of off-policy learning where agents learn from other agents behaviour. We employ those insights to accelerate hyper-parameter sweeps in which all participating agents run concurrently and share their experience via a common replay module. To this end we analyze the bias-variance tradeoffs in V-trace, a form of importance sampling for actor-critic methods. Based on our analysis, we then argue for mixing experience sampled from replay with on-policy experience, and propose a new trust region scheme that scales effectively to data distributions where V-trace becomes unstable. We provide extensive empirical validation of the proposed solution. We further show the benefits of this setup by demonstrating state-of-the-art data efficiency on Atari among agents trained up until 200M environment frames.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. An Arbitration Control for an Ensemble of Diversified DQN variants in Continual Reinforcement Learning

    cs.LG 2025-09 conditional novelty 5.0 of 10

    ACED-DQN combines heterogeneous DQN variants with loss-based reliability weighting and experience assignment, but the paper's own ablation indicates that arbitration control is not the key factor behind the performance gain.

Pith tools