Pith. sign in

REVIEW 6 cited by

Accelerated Primal-Dual Policy Optimization for Safe Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1802.06480 v1 pith:WPUHZ2WS submitted 2018-02-19 cs.AI cs.LGstat.ML

classification cs.AIcs.LGstat.ML
keywords cmdpsdualpolicyprimal-dualacceleratedapdoconstraintsconvergence
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Constrained Markov Decision Process (CMDP) is a natural framework for reinforcement learning tasks with safety constraints, where agents learn a policy that maximizes the long-term reward while satisfying the constraints on the long-term cost. A canonical approach for solving CMDPs is the primal-dual method which updates parameters in primal and dual spaces in turn. Existing methods for CMDPs only use on-policy data for dual updates, which results in sample inefficiency and slow convergence. In this paper, we propose a policy search method for CMDPs called Accelerated Primal-Dual Optimization (APDO), which incorporates an off-policy trained dual variable in the dual update procedure while updating the policy in primal space with on-policy likelihood ratio gradient. Experimental results on a simulated robot locomotion task show that APDO achieves better sample efficiency and faster convergence than state-of-the-art approaches for CMDPs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 43 citations worldwide. Full citation record

  1. Scalable Constrained Multi-Agent Reinforcement Learning via State Augmentation and Consensus for Separable Dynamics

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    A consensus-based distributed algorithm for constrained MARL with separable dynamics achieves linear scalability and bounded constraint violations through state-augmented policies and dual variable agreement.

  2. Stationary Robust Mean-Field Games under Model Mismatches

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    Develops infinite-horizon stationary robust mean-field games incorporating distributional uncertainty, proves equilibrium existence via fixed-point on contractive Bellman operator, gives convergent algorithm, and deri...

  3. PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning

    cs.LG 2025-07 conditional novelty 6.0 of 10

    PNAct trains a safe RL agent to take unsafe actions only when a hidden trigger is present, while keeping normal safe behavior and reward when the trigger is absent.

  4. Tilted Quantile Gradient Updates for Quantile-Constrained Reinforcement Learning

    cs.LG 2024-12 reject novelty 5.0 of 10

    TQPO estimates gradients of quantile safety constraints directly through sampling and adds a tilted update to the Lagrange multiplier, improving return while satisfying the constraints.

  5. Control Synthesis with Reinforcement Learning: A Modeling Perspective

    eess.SY 2025-10 conditional novelty 4.0 of 10

    A simplified linear training model yields an RL cart-pole controller that fails in physical deployment, while a high-fidelity nonlinear model yields a deployable, disturbance-robust controller.

  6. Effective Reward Specification in Deep Reinforcement Learning

    cs.LG 2024-12 conditional novelty 4.0 of 10

    A thesis presenting four methods (ASAF, TeamReg, CoachReg, constrained RL, goal-conditioned GFlowNets) that improve reward specification for deep RL through demonstrations, policy regularization, behavior constraints,...

Pith tools