Pith. sign in

REVIEW 1 cited by

Independent and Decentralized Learning in Markov Potential Games

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.14590 v8 pith:SO3CLQRQ submitted 2022-05-29 cs.LG cs.AIcs.GTcs.MAcs.SYeess.SY

classification cs.LGcs.AIcs.GTcs.MAcs.SYeess.SY
keywords learningdynamicsplayersq-functionupdateasynchronousdecentralizedgames
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We study a multi-agent reinforcement learning dynamics, and analyze its asymptotic behavior in infinite-horizon discounted Markov potential games. We focus on the independent and decentralized setting, where players do not know the game parameters, and cannot communicate or coordinate. In each stage, players update their estimate of Q-function that evaluates their total contingent payoff based on the realized one-stage reward in an asynchronous manner. Then, players independently update their policies by incorporating an optimal one-stage deviation strategy based on the estimated Q-function. Inspired by the actor-critic algorithm in single-agent reinforcement learning, a key feature of our learning dynamics is that agents update their Q-function estimates at a faster timescale than the policies. Leveraging tools from two-timescale asynchronous stochastic approximation theory, we characterize the convergent set of learning dynamics.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Independent Learning in Performative Markov Potential Games

    cs.LG 2025-04 reject novelty 6.0 of 10

    In performative Markov potential games, independent policy gradient and natural policy gradient algorithms are claimed to converge to an approximate performatively stable equilibrium, but the proofs contain a false eq...

Pith tools