REVIEW 5 cited by
Datasets and Benchmarks for Offline Safe Reinforcement Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper presents a comprehensive benchmarking suite tailored to offline safe reinforcement learning (RL) challenges, aiming to foster progress in the development and evaluation of safe learning algorithms in both the training and deployment phases. Our benchmark suite contains three packages: 1) expertly crafted safe policies, 2) D4RL-styled datasets along with environment wrappers, and 3) high-quality offline safe RL baseline implementations. We feature a methodical data collection pipeline powered by advanced safe RL algorithms, which facilitates the generation of diverse datasets across 38 popular safe RL tasks, from robot control to autonomous driving. We further introduce an array of data post-processing filters, capable of modifying each dataset's diversity, thereby simulating various data collection conditions. Additionally, we provide elegant and extensible implementations of prevalent offline safe RL algorithms to accelerate research in this area. Through extensive experiments with over 50000 CPU and 800 GPU hours of computations, we evaluate and compare the performance of these baseline algorithms on the collected datasets, offering insights into their strengths, limitations, and potential areas of improvement. Our benchmarking framework serves as a valuable resource for researchers and practitioners, facilitating the development of more robust and reliable offline safe RL solutions in safety-critical applications. The benchmark website is available at \url{www.offline-saferl.org}.
Forward citations
Cited by 5 Pith papers
-
A Provable Approach for End-to-End Safe Reinforcement Learning
PLS combines offline return-conditioned policy training with Gaussian-process safe optimization of target returns to provide high-probability safety throughout deployment.
-
CDCP: Conditional Diffusion Model with Contextual Prompts for Multi-task Offline Safe Reinforcement Learning
A conditional diffusion model with contextual prompts and classifier-free cost guidance learns a shared safe multi-task policy from offline data and meets varying cost limits without retraining.
-
Skill Expansion and Composition in Parameter Space
PSEC shows that weighting and summing LoRA skill modules inside a diffusion policy network outperforms composing the same skills in action or noise space across D4RL, DSRL, DMC, and Meta-World tasks.
-
From Uncertain to Safe: Conformal Adaptation of Diffusion Models for Safe PDE Control
SafeDiffCon adapts a diffusion model for PDE control by adding a conformal-prediction uncertainty quantile to its safety penalty, and reports zero safety violations on three control benchmarks.
-
PyTupli: A Scalable Infrastructure for Collaborative Offline Reinforcement Learning Projects
PyTupli is a client-server tool for recording, storing, filtering, and sharing offline reinforcement learning datasets tied to custom gymnasium environments.
Discussion (0). Continue with ORCID to comment.