Pith. sign in

REVIEW 1 cited by

Upper Confidence Primal-Dual Reinforcement Learning for CMDP with Adversarial Loss

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.00660 v3 pith:BB2HMX7B submitted 2020-03-02 cs.LG math.OCstat.ML

classification cs.LGmath.OCstat.ML
keywords learningmathcalupperanalysisconfidencelossprocessesregret
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We consider online learning for episodic stochastically constrained Markov decision processes (CMDPs), which plays a central role in ensuring the safety of reinforcement learning. Here the loss function can vary arbitrarily across the episodes, and both the loss received and the budget consumption are revealed at the end of each episode. Previous works solve this problem under the restrictive assumption that the transition model of the Markov decision processes (MDPs) is known a priori and establish regret bounds that depend polynomially on the cardinalities of the state space $\mathcal{S}$ and the action space $\mathcal{A}$. In this work, we propose a new \emph{upper confidence primal-dual} algorithm, which only requires the trajectories sampled from the transition model. In particular, we prove that the proposed algorithm achieves $\widetilde{\mathcal{O}}(L|\mathcal{S}|\sqrt{|\mathcal{A}|T})$ upper bounds of both the regret and the constraint violation, where $L$ is the length of each episode. Our analysis incorporates a new high-probability drift analysis of Lagrange multiplier processes into the celebrated regret analysis of upper confidence reinforcement learning, which demonstrates the power of "optimism in the face of uncertainty" in constrained online learning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Adjustment Speed as a Safety Constraint for Nonstationary Reinforcement Learning

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Forecasting context shifts and treating adaptation demand against calibrated recovery capacity as a safety gate reduces transient violations in a nonstationary highway-driving simulator.

Pith tools