REVIEW 1 cited by
Multi-agent online learning in time-varying games
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We examine the long-run behavior of multi-agent online learning in games that evolve over time. Specifically, we focus on a wide class of policies based on mirror descent, and we show that the induced sequence of play (a) converges to Nash equilibrium in time-varying games that stabilize in the long run to a strictly monotone limit; and (b) it stays asymptotically close to the evolving equilibrium of the sequence of stage games (assuming they are strongly monotone). Our results apply to both gradient-based and payoff-based feedback - i.e., the "bandit feedback" case where players only get to observe the payoffs of their chosen actions.
Forward citations
Cited by 1 Pith paper
-
Prediction-Aware Learning in Multi-Agent Systems
A contextual optimistic multiplicative weights algorithm (POMWU) achieves static-game regret, equilibrium convergence, and social welfare guarantees in time-varying games when players can predict the changing state of...
Discussion (0). Continue with ORCID to comment.