Pith. sign in

REVIEW 5 cited by

Datasets and Benchmarks for Offline Safe Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.09303 v2 pith:FMZ6UB2U submitted 2023-06-15 cs.LG cs.AIcs.RO

classification cs.LGcs.AIcs.RO
keywords safeofflinealgorithmsdatasetsdatalearningbaselinebenchmark
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper presents a comprehensive benchmarking suite tailored to offline safe reinforcement learning (RL) challenges, aiming to foster progress in the development and evaluation of safe learning algorithms in both the training and deployment phases. Our benchmark suite contains three packages: 1) expertly crafted safe policies, 2) D4RL-styled datasets along with environment wrappers, and 3) high-quality offline safe RL baseline implementations. We feature a methodical data collection pipeline powered by advanced safe RL algorithms, which facilitates the generation of diverse datasets across 38 popular safe RL tasks, from robot control to autonomous driving. We further introduce an array of data post-processing filters, capable of modifying each dataset's diversity, thereby simulating various data collection conditions. Additionally, we provide elegant and extensible implementations of prevalent offline safe RL algorithms to accelerate research in this area. Through extensive experiments with over 50000 CPU and 800 GPU hours of computations, we evaluate and compare the performance of these baseline algorithms on the collected datasets, offering insights into their strengths, limitations, and potential areas of improvement. Our benchmarking framework serves as a valuable resource for researchers and practitioners, facilitating the development of more robust and reliable offline safe RL solutions in safety-critical applications. The benchmark website is available at \url{www.offline-saferl.org}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Provable Approach for End-to-End Safe Reinforcement Learning

    cs.LG 2025-05 conditional novelty 7.0 of 10

    PLS combines offline return-conditioned policy training with Gaussian-process safe optimization of target returns to provide high-probability safety throughout deployment.

  2. CDCP: Conditional Diffusion Model with Contextual Prompts for Multi-task Offline Safe Reinforcement Learning

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A conditional diffusion model with contextual prompts and classifier-free cost guidance learns a shared safe multi-task policy from offline data and meets varying cost limits without retraining.

  3. Skill Expansion and Composition in Parameter Space

    cs.LG 2025-02 conditional novelty 6.0 of 10

    PSEC shows that weighting and summing LoRA skill modules inside a diffusion policy network outperforms composing the same skills in action or noise space across D4RL, DSRL, DMC, and Meta-World tasks.

  4. From Uncertain to Safe: Conformal Adaptation of Diffusion Models for Safe PDE Control

    cs.LG 2025-02 conditional novelty 6.0 of 10

    SafeDiffCon adapts a diffusion model for PDE control by adding a conformal-prediction uncertainty quantile to its safety penalty, and reports zero safety violations on three control benchmarks.

  5. PyTupli: A Scalable Infrastructure for Collaborative Offline Reinforcement Learning Projects

    cs.LG 2025-05 conditional novelty 4.0 of 10

    PyTupli is a client-server tool for recording, storing, filtering, and sharing offline reinforcement learning datasets tied to custom gymnasium environments.

Pith tools