Pith. sign in

REVIEW 1 cited by

Full error analysis of policy gradient learning algorithms for exploratory linear quadratic mean-field control problem in continuous time with common noise

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.02489 v1 pith:ZTM4EVFX submitted 2024-08-05 math.OC stat.ML

classification math.OCstat.ML
keywords gradientconvergencelearninglinearsettingalgorithmsanalysiscommon
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We consider reinforcement learning (RL) methods for finding optimal policies in linear quadratic (LQ) mean field control (MFC) problems over an infinite horizon in continuous time, with common noise and entropy regularization. We study policy gradient (PG) learning and first demonstrate convergence in a model-based setting by establishing a suitable gradient domination condition.Next, our main contribution is a comprehensive error analysis, where we prove the global linear convergence and sample complexity of the PG algorithm with two-point gradient estimates in a model-free setting with unknown parameters. In this setting, the parameterized optimal policies are learned from samples of the states and population distribution.Finally, we provide numerical evidence supporting the convergence of our implemented algorithms.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Policy Optimization for Continuous-time Linear-Quadratic Graphon Mean Field Games

    math.OC 2025-06 accept novelty 7.0 of 10

    A bilevel policy optimization algorithm for continuous-time linear-quadratic graphon mean field games converges linearly to best-response policies and globally to the Nash equilibrium.

Pith tools