REVIEW 6 cited by
Open RL Benchmark: Comprehensive Tracked Experiments for Reinforcement Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In many Reinforcement Learning (RL) papers, learning curves are useful indicators to measure the effectiveness of RL algorithms. However, the complete raw data of the learning curves are rarely available. As a result, it is usually necessary to reproduce the experiments from scratch, which can be time-consuming and error-prone. We present Open RL Benchmark, a set of fully tracked RL experiments, including not only the usual data such as episodic return, but also all algorithm-specific and system metrics. Open RL Benchmark is community-driven: anyone can download, use, and contribute to the data. At the time of writing, more than 25,000 runs have been tracked, for a cumulative duration of more than 8 years. Open RL Benchmark covers a wide range of RL libraries and reference implementations. Special care is taken to ensure that each experiment is precisely reproducible by providing not only the full parameters, but also the versions of the dependencies used to generate it. In addition, Open RL Benchmark comes with a command-line interface (CLI) for easy fetching and generating figures to present the results. In this document, we include two case studies to demonstrate the usefulness of Open RL Benchmark in practice. To the best of our knowledge, Open RL Benchmark is the first RL benchmark of its kind, and the authors hope that it will improve and facilitate the work of researchers in the field.
Forward citations
Cited by 6 Pith papers
-
Flow Matching Policy Gradients
FPO trains flow-based policies with PPO by replacing the likelihood ratio with an exponentiated flow matching loss difference.
-
Safe Domain Randomization via Uncertainty-Aware Out-of-Distribution Detection and Policy Adaptation
UARL gates policy deployment on ensemble critic variance computed on a target-domain dataset, iteratively expanding domain randomization until the uncertainty threshold is met.
-
When Maximum Entropy Misleads Policy Optimization
Maximum entropy RL can be formally steered into arbitrary suboptimal policies at convergence by adding entropy trap states, while standard RL is unaffected.
-
First Order Model-Based RL through Decoupled Backpropagation
By computing gradients through a learned dynamics model while unrolling trajectories in the real simulator, DMO achieves SHAC-level sample efficiency with standard simulators and deploys on a real quadruped.
-
CleanQRL: Lightweight Single-file Implementations of Quantum Reinforcement Learning Algorithms
The paper introduces CleanQRL, a collection of single-file implementations of quantum reinforcement learning algorithms designed to make QRL research easier to replicate and compare.
-
Combining Pre-Trained Models for Enhanced Feature Representation in Reinforcement Learning
A lightweight attention module that weights embeddings from multiple pre-trained models achieves comparable Atari RL performance to end-to-end training, with improved robustness to visual changes.
Discussion (0). Continue with ORCID to comment.