Pith. sign in

REVIEW 3 cited by

Neural Policy Gradient Methods: Global Optimality and Rates of Convergence

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.01150 v3 pith:QXJCMMPV submitted 2019-08-29 cs.LG math.OCstat.ML

classification cs.LGmath.OCstat.ML
keywords neuralpolicygradientglobalmethodsoptimalityactorconvergence
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Policy gradient methods with actor-critic schemes demonstrate tremendous empirical successes, especially when the actors and critics are parameterized by neural networks. However, it remains less clear whether such "neural" policy gradient methods converge to globally optimal policies and whether they even converge at all. We answer both the questions affirmatively in the overparameterized regime. In detail, we prove that neural natural policy gradient converges to a globally optimal policy at a sublinear rate. Also, we show that neural vanilla policy gradient converges sublinearly to a stationary point. Meanwhile, by relating the suboptimality of the stationary points to the representation power of neural actor and critic classes, we prove the global optimality of all stationary points under mild regularity conditions. Particularly, we show that a key to the global optimality and convergence is the "compatibility" between the actor and critic, which is ensured by sharing neural architectures and random initializations across the actor and critic. To the best of our knowledge, our analysis establishes the first global optimality and convergence guarantees for neural policy gradient methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Geometry of Neural Reinforcement Learning in Continuous State and Action Spaces

    cs.LG 2025-07 conditional novelty 7.0 of 10

    For wide two-layer linearized neural policies in deterministic continuous RL, the locally attainable states concentrate on a manifold of dimension at most 2da+1, independent of the state dimension.

  2. Adaptive Partitioning and Learning for Stochastic Control of Diffusion Processes

    cs.LG 2025-12 conditional novelty 6.0 of 10

    APL-Diffusion achieves a regret bound in K whose exponent is governed by a new zooming dimension tailored to unbounded diffusions, recovering Sinclair et al.'s rate as the initial-state moment p goes to infinity.

  3. The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Reversing PPO's update order in federated learning, actor before critic, removes the pairwise critic-difference term from the convergence bound and improves performance on heterogeneous control tasks.

Pith tools