Pith. sign in

REVIEW 1 cited by

First-order Policy Optimization for Robust Markov Decision Process

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.10579 v2 pith:QXKTLSM6 submitted 2022-09-21 cs.LG cs.AImath.OC

classification cs.LGcs.AImath.OC
keywords robustpolicyfirst-orderepsilonmethodcomplexitydecisiondescent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We consider the problem of solving robust Markov decision process (MDP), which involves a set of discounted, finite state, finite action space MDPs with uncertain transition kernels. The goal of planning is to find a robust policy that optimizes the worst-case values against the transition uncertainties, and thus encompasses the standard MDP planning as a special case. For $(\mathbf{s},\mathbf{a})$-rectangular uncertainty sets, we establish several structural observations on the robust objective, which facilitates the development of a policy-based first-order method, namely the robust policy mirror descent (RPMD). An $\mathcal{O}(\log(1/\epsilon))$ iteration complexity for finding an $\epsilon$-optimal policy is established with linearly increasing stepsizes. We further develop a stochastic variant of the robust policy mirror descent method, named SRPMD, when the first-order information is only available through online interactions with the nominal environment. We show that the optimality gap converges linearly up to the noise level, and consequently establish an $\tilde{\mathcal{O}}(1/\epsilon^2)$ sample complexity by developing a temporal difference learning method for policy evaluation. Both iteration and sample complexities are also discussed for RPMD with a constant stepsize. To the best of our knowledge, all the aforementioned results appear to be new for policy-based first-order methods applied to the robust MDP problem.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning

    cs.LG 2025-06 conditional novelty 7.0 of 10

    For robust average-reward MDPs, the paper proves model-free Q-learning and actor-critic algorithms converge with tilde O(epsilon^{-2}) sample complexity via a carefully constructed semi-norm contraction.

Pith tools