Pith. sign in

REVIEW 3 cited by

Neural Proximal/Trust Region Policy Optimization Attains Globally Optimal Policy

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1906.10306 v3 pith:WG5LEPJV submitted 2019-06-25 cs.LG math.OCstat.ML

classification cs.LGmath.OCstat.ML
keywords neuralpolicynetworksoptimizationtrpoconvergenceglobalglobally
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Proximal policy optimization and trust region policy optimization (PPO and TRPO) with actor and critic parametrized by neural networks achieve significant empirical success in deep reinforcement learning. However, due to nonconvexity, the global convergence of PPO and TRPO remains less understood, which separates theory from practice. In this paper, we prove that a variant of PPO and TRPO equipped with overparametrized neural networks converges to the globally optimal policy at a sublinear rate. The key to our analysis is the global convergence of infinite-dimensional mirror descent under a notion of one-point monotonicity, where the gradient and iterate are instantiated by neural networks. In particular, the desirable representation power and optimization geometry induced by the overparametrization of such neural networks allow them to accurately approximate the infinite-dimensional gradient and iterate.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPs

    cs.LG 2019-09 conditional novelty 7.0 of 10

    Adaptive TRPO is shown to be mirror descent with an adaptive proximity term, converging at tilde O(1/sqrt(N)) and at tilde O(1/N) for regularized MDPs.

  2. Neural Policy Gradient Methods: Global Optimality and Rates of Convergence

    cs.LG 2019-08 conditional novelty 7.0 of 10

    Under strong regularity assumptions, neural natural policy gradient converges to a global optimum at rate O(1/sqrt(T)), and neural vanilla policy gradient converges to a stationary point at the same rate.

  3. Fast Convergence of Softmax Policy Mirror Ascent

    cs.LG 2024-11 conditional novelty 6.0 of 10

    Softmax policy mirror ascent is a normalization-free mirror ascent on logits that converges linearly in tabular MDPs and linearly to a neighborhood with function approximation.

Pith tools