Pith. sign in

REVIEW 10 cited by

Projection-Based Constrained Policy Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.03152 v1 pith:LUOEGFC2 submitted 2020-10-07 cs.LG cs.AIcs.RO

classification cs.LGcs.AIcs.RO
keywords constraintpcpopolicyrewardviolationboundconstrainedcontrol
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

We consider the problem of learning control policies that optimize a reward function while satisfying constraints due to considerations of safety, fairness, or other costs. We propose a new algorithm, Projection-Based Constrained Policy Optimization (PCPO). This is an iterative method for optimizing policies in a two-step process: the first step performs a local reward improvement update, while the second step reconciles any constraint violation by projecting the policy back onto the constraint set. We theoretically analyze PCPO and provide a lower bound on reward improvement, and an upper bound on constraint violation, for each policy update. We further characterize the convergence of PCPO based on two different metrics: $\normltwo$ norm and Kullback-Leibler divergence. Our empirical results over several control tasks demonstrate that PCPO achieves superior performance, averaging more than 3.5 times less constraint violation and around 15\% higher reward compared to state-of-the-art methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning

    cs.LG 2025-07 conditional novelty 6.0 of 10

    PNAct trains a safe RL agent to take unsafe actions only when a hidden trigger is present, while keeping normal safe behavior and reward when the trigger is absent.

  2. Operator Splitting for Convex Constrained Markov Decision Processes

    math.OC 2024-12 conditional novelty 6.0 of 10

    OS-CMDP uses Douglas-Rachford splitting to solve convex-constrained MDPs by alternating between a quadratically regularized MDP update and a projection onto the constraint set, with convergence and infeasibility-detec...

  3. Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation

    cs.LG 2024-12 conditional novelty 6.0 of 10

    The paper introduces Gradient-based Estimation (GBE), a Taylor-expansion method using analytic trajectory gradients, and the trust-region algorithm CGPO for safe RL with finite-horizon constraints.

  4. From Text to Trajectory: Exploring Complex Constraint Representation and Decomposition in Safe Reinforcement Learning

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A text-trajectory contrastive model with per-step cost assignment reduces safety violations in reinforcement learning agents under natural language constraints.

  5. Q-learning-based Model-free Safety Filter

    cs.RO 2024-11 reject novelty 6.0 of 10

    A Q-learning safety filter with a time-dependent reward blocks unsafe actions from arbitrary task policies, but its theoretical guarantee is not valid as written.

  6. Towards Adaptive External Communication in Autonomous Vehicles: A Conceptual Design Framework

    cs.HC 2025-08 unverdicted novelty 5.0 of 10

    A three-layer framework (input, processing, output) for adaptive external human-machine interfaces in autonomous vehicles is introduced to systematize design and analysis.

  7. Proactive Constrained Policy Optimization with Preemptive Penalty

    cs.LG 2025-08 reject novelty 5.0 of 10

    A constrained policy optimization method that uses a preemptive log-barrier penalty and a constraint-aware intrinsic reward to reduce safety violations in reinforcement learning.

  8. Tilted Quantile Gradient Updates for Quantile-Constrained Reinforcement Learning

    cs.LG 2024-12 reject novelty 5.0 of 10

    TQPO estimates gradients of quantile safety constraints directly through sampling and adds a tilted update to the Lagrange multiplier, improving return while satisfying the constraints.

  9. HAEPO: History-Aggregated Exploratory Policy Optimization

    cs.LG 2025-08 conditional novelty 4.0 of 10

    HAEPO weights each trajectory by its softmax-normalized cumulative log-likelihood, adds entropy and KL penalties, and matches or slightly surpasses PPO, GRPO, and DPO on small RL and summarization tasks.

  10. Situational-Constrained Sequential Resources Allocation via Reinforcement Learning

    cs.AI 2025-06 conditional novelty 4.0 of 10

    A constrained-reinforcement-learning algorithm encodes situational if-then allocation rules as disjunctive penalties and shows lower violations on simulated medical and agricultural allocation tasks.

Pith tools