REVIEW 5 cited by
A Unified Approach to Reinforcement Learning, Quantal Response Equilibria, and Two-Player Zero-Sum Games
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This work studies an algorithm, which we call magnetic mirror descent, that is inspired by mirror descent and the non-Euclidean proximal gradient algorithm. Our contribution is demonstrating the virtues of magnetic mirror descent as both an equilibrium solver and as an approach to reinforcement learning in two-player zero-sum games. These virtues include: 1) Being the first quantal response equilibria solver to achieve linear convergence for extensive-form games with first order feedback; 2) Being the first standard reinforcement learning algorithm to achieve empirically competitive results with CFR in tabular settings; 3) Achieving favorable performance in 3x3 Dark Hex and Phantom Tic-Tac-Toe as a self-play deep reinforcement learning algorithm.
Forward citations
Cited by 5 Pith papers
-
IFlowNets: Extending Generative Samplers to Learn Strategies in Incomplete Information Games
IFlowNets make generative flow networks work for imperfect-information games by adding an information-set aggregation constraint that restores valid flow matching.
-
Incentivize without Bonus: Provably Efficient Model-based Online Multi-agent RL for Markov Games
Value-incentivized exploration via best-response values gives near-optimal regret for NE/CCE in linear-model Markov games without explicit uncertainty bonuses.
-
The Curvature Shadow: An Apparent Failure of Maximum-Entropy Equilibrium Selection is a Removable Artifact
The Kuhn-poker selection gap is fully explained by gap ≈ sqrt(2δ/κ), where δ is a small removable entropy shortfall and κ is the peak curvature, and it vanishes as δ approaches zero.
-
FootsiesGym: A Fighting Game Benchmark for Two-Player Zero-Sum Imperfect-Information Games
FootsiesGym is an open-source, vectorized fighting-game benchmark for two-player zero-sum imperfect-information RL that isolates non-transitive neutral-game dynamics while remaining tractable on standard hardware.
-
Two-Player Zero-Sum Differential Games with One-Sided Information
A new algorithm, CAMS, reformulates the Bellman backup in one-sided-information differential games into small nonconvex minimax problems, achieving complexity independent of the continuous action space size and demons...
Discussion (0). Continue with ORCID to comment.