Pith. sign in

REVIEW 1 cited by

Beyond Stationarity: Convergence Analysis of Stochastic Softmax Policy Gradient Methods

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.02671 v2 pith:3PWDTNTR submitted 2023-10-04 math.OC cs.LGstat.ML

classification math.OCcs.LGstat.ML
keywords gradientdynamicpolicyproblemsconvergenceanalysismdpsoptimal
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Markov Decision Processes (MDPs) are a formal framework for modeling and solving sequential decision-making problems. In finite-time horizons such problems are relevant for instance for optimal stopping or specific supply chain problems, but also in the training of large language models. In contrast to infinite horizon MDPs optimal policies are not stationary, policies must be learned for every single epoch. In practice all parameters are often trained simultaneously, ignoring the inherent structure suggested by dynamic programming. This paper introduces a combination of dynamic programming and policy gradient called dynamic policy gradient, where the parameters are trained backwards in time. For the tabular softmax parametrisation we carry out the convergence analysis for simultaneous and dynamic policy gradient towards global optima, both in the exact and sampled gradient settings without regularisation. It turns out that the use of dynamic policy gradient training much better exploits the structure of finite- time problems which is reflected in improved convergence bounds.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Simultaneous Best-Response Dynamics in Random Potential Games

    cs.GT 2025-05 conditional novelty 7.0 of 10

    In two-player random potential games with many actions, simultaneous best-response dynamics converge with high probability to a two-cycle whose off-diagonal profiles are Nash equilibria; simulations indicate convergen...

Pith tools