Pith. sign in

REVIEW 2 cited by

A Survey of Constraint Formulations in Safe Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.02025 v2 pith:YJBREWYP submitted 2024-02-03 cs.LG cs.AI

classification cs.LGcs.AI
keywords safesafetyconstraintformulationslearningreinforcementadditionagent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Safety is critical when applying reinforcement learning (RL) to real-world problems. As a result, safe RL has emerged as a fundamental and powerful paradigm for optimizing an agent's policy while incorporating notions of safety. A prevalent safe RL approach is based on a constrained criterion, which seeks to maximize the expected cumulative reward subject to specific safety constraints. Despite recent effort to enhance safety in RL, a systematic understanding of the field remains difficult. This challenge stems from the diversity of constraint representations and little exploration of their interrelations. To bridge this knowledge gap, we present a comprehensive review of representative constraint formulations, along with a curated selection of algorithms designed specifically for each formulation. In addition, we elucidate the theoretical underpinnings that reveal the mathematical mutual relations among common problem formulations. We conclude with a discussion of the current state and future directions of safe reinforcement learning research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Network Automation to Trustworthy Autonomous Networking in the LLM Era: A Network Control Intelligence Perspective

    cs.NI 2026-08 conditional novelty 6.0 of 10

    A new five-axis framework for profiling network control systems shows that automation progress is uneven and that LLMs should be limited to proposal generation, not commit authority.

  2. Adjustment Speed as a Safety Constraint for Nonstationary Reinforcement Learning

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Forecasting context shifts and treating adaptation demand against calibrated recovery capacity as a safety gate reduces transient violations in a nonstationary highway-driving simulator.

Pith tools