Pith. sign in

REVIEW 3 cited by

Feasible Actor-Critic: Constrained Reinforcement Learning for Ensuring Statewise Safety

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2105.10682 v3 pith:NFJT4E2F submitted 2021-05-22 cs.LG

classification cs.LG
keywords safetyfeasiblestatesstatewisestateconstrainedpolicytasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The safety constraints commonly used by existing safe reinforcement learning (RL) methods are defined only on expectation of initial states, but allow each certain state to be unsafe, which is unsatisfying for real-world safety-critical tasks. In this paper, we introduce the feasible actor-critic (FAC) algorithm, which is the first model-free constrained RL method that considers statewise safety, e.g, safety for each initial state. We claim that some states are inherently unsafe no matter what policy we choose, while for other states there exist policies ensuring safety, where we say such states and policies are feasible. By constructing a statewise Lagrange function available on RL sampling and adopting an additional neural network to approximate the statewise Lagrange multiplier, we manage to obtain the optimal feasible policy which ensures safety for each feasible state and the safest possible policy for infeasible states. Furthermore, the trained multiplier net can indicate whether a given state is feasible or not through the statewise complementary slackness condition. We provide theoretical guarantees that FAC outperforms previous expectation-based constrained RL methods in terms of both constraint satisfaction and reward optimization. Experimental results on both robot locomotive tasks and safe exploration tasks verify the safety enhancement and feasibility interpretation of the proposed method.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Continual Reinforcement Learning for Digital Twin Synchronization Optimization

    cs.NI 2025-01 conditional novelty 5.0 of 10

    A continual reinforcement learning scheduler with multi-timescale replay and a resource-constrained actor-critic reduces digital twin state estimation error by up to 55.2% in simulation.

  2. Action Mapping for Reinforcement Learning in Continuous Environments with Constraints

    cs.LG 2024-12 conditional novelty 5.0 of 10

    Decoupling feasibility from objective optimization by training the RL policy over latent actions that map to feasible actions improves sample efficiency and constraint satisfaction in continuous constrained RL.

  3. FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning

    cs.LG 2024-12 reject novelty 4.0 of 10

    FAWAC adds a cost-advantage penalty to advantage weighted regression to keep offline-trained policies within a safety budget, with variants for standard and high-reward-but-unsafe datasets.

Pith tools