Pith. sign in

REVIEW 11 cited by

What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.05990 v1 pith:ZYXQIT4P submitted 2020-06-10 cs.LG stat.ML

classification cs.LGstat.ML
keywords on-policyagentsalgorithmschoicescontinuouscontroldifferentempirical
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In recent years, on-policy reinforcement learning (RL) has been successfully applied to many different continuous control tasks. While RL algorithms are often conceptually simple, their state-of-the-art implementations take numerous low- and high-level design decisions that strongly affect the performance of the resulting agents. Those choices are usually not extensively discussed in the literature, leading to discrepancy between published descriptions of algorithms and their implementations. This makes it hard to attribute progress in RL and slows down overall progress [Engstrom'20]. As a step towards filling that gap, we implement >50 such ``choices'' in a unified on-policy RL framework, allowing us to investigate their impact in a large-scale empirical study. We train over 250'000 agents in five continuous control environments of different complexity and provide insights and practical recommendations for on-policy training of RL agents.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 105 citations worldwide. Full citation record

  1. MuJoCo Playground

    cs.RO 2025-02 conditional novelty 7.0 of 10

    An open-source, MJX-based robot learning framework with integrated batch rendering that provides fast training and demonstrates sim-to-real transfer on six robot platforms.

  2. The Importance of Encoder Choice:A Tabular-Image Study

    cs.LG 2026-07 conditional novelty 6.5 of 10

    Tabular encoder choice reorders multimodal rankings, can erase apparent fusion gains, and requires non-vanilla extraction for in-context learning models to avoid train-test representation shift.

  3. Differentiable Weightless Controllers: Learning Logic Circuits for Continuous Control

    cs.LG 2025-12 conditional novelty 6.0 of 10

    Logic-gate circuits trained with gradient descent can match neural-network policies on most MuJoCo continuous-control tasks and run on FPGAs in a few clock cycles.

  4. Online Training and Pruning of Deep Reinforcement Learning Networks

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A method that prunes OFENet-based reinforcement learning networks during training, reducing them to a fraction of their original size with minimal performance loss.

  5. Learning human-to-robot handovers through 3D scene reconstruction

    cs.RO 2025-07 conditional novelty 6.0 of 10

    A handover policy trained only on images rendered from a sparse-view Gaussian Splatting scene can deploy on a real robot without real-robot training data.

  6. On the Effect of Regularization in Policy Mirror Descent

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A large empirical sweep shows that in Policy Mirror Descent, MDP and Drift regularizers are partly substitutable, yet their precise combination determines temperature robustness.

  7. Magistral

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Pure RL alone, without distilled reasoning traces, turned Mistral's base models into strong reasoning models on math and coding benchmarks.

  8. Simulating Errors in Touchscreen Typing

    cs.HC 2025-02 conditional novelty 6.0 of 10

    Typoist is a simulation model that reproduces slips, lapses, and mistakes in touchscreen typing, with error rates matching human data in a new benchmark.

  9. Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments

    cs.LG 2026-03 conditional novelty 5.0 of 10

    PPO plateaus can be avoided by increasing the number of parallel environments, which reduces both the outer-loop step size and update noise; scaling to 1M environments sustained improvement to 1T transitions.

  10. RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming

    cs.LG 2025-06 reject novelty 5.0 of 10

    RedRFT is a new open-source benchmark with a unified PPO backbone, five reimplemented red teaming baselines, a proposed diversity metric, and ablation insights.

  11. Communicating Smartly in Molecular Communication Environments: Neural Networks in the Internet of Bio-Nano Things

    eess.SP 2025-06 conditional novelty 3.0 of 10

    A broad, code-augmented survey of neural network methods for molecular communication in the Internet of Bio-Nano Things, including a dataset accessibility audit and open challenges.

Pith tools