Pith. sign in

REVIEW 5 cited by

A Unified Approach to Reinforcement Learning, Quantal Response Equilibria, and Two-Player Zero-Sum Games

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.05825 v4 pith:M23IBZMZ submitted 2022-06-12 cs.LG cs.AIcs.GT

classification cs.LGcs.AIcs.GT
keywords algorithmlearningreinforcementdescentfirstgamesmirrorachieve
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This work studies an algorithm, which we call magnetic mirror descent, that is inspired by mirror descent and the non-Euclidean proximal gradient algorithm. Our contribution is demonstrating the virtues of magnetic mirror descent as both an equilibrium solver and as an approach to reinforcement learning in two-player zero-sum games. These virtues include: 1) Being the first quantal response equilibria solver to achieve linear convergence for extensive-form games with first order feedback; 2) Being the first standard reinforcement learning algorithm to achieve empirically competitive results with CFR in tabular settings; 3) Achieving favorable performance in 3x3 Dark Hex and Phantom Tic-Tac-Toe as a self-play deep reinforcement learning algorithm.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. IFlowNets: Extending Generative Samplers to Learn Strategies in Incomplete Information Games

    cs.LG 2026-08 conditional novelty 6.0 of 10

    IFlowNets make generative flow networks work for imperfect-information games by adding an information-set aggregation constraint that restores valid flow matching.

  2. Incentivize without Bonus: Provably Efficient Model-based Online Multi-agent RL for Markov Games

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Value-incentivized exploration via best-response values gives near-optimal regret for NE/CCE in linear-model Markov games without explicit uncertainty bonuses.

  3. The Curvature Shadow: An Apparent Failure of Maximum-Entropy Equilibrium Selection is a Removable Artifact

    cs.AI 2026-07 conditional novelty 5.0 of 10

    The Kuhn-poker selection gap is fully explained by gap ≈ sqrt(2δ/κ), where δ is a small removable entropy shortfall and κ is the peak curvature, and it vanishes as δ approaches zero.

  4. FootsiesGym: A Fighting Game Benchmark for Two-Player Zero-Sum Imperfect-Information Games

    cs.AI 2026-07 accept novelty 5.0 of 10

    FootsiesGym is an open-source, vectorized fighting-game benchmark for two-player zero-sum imperfect-information RL that isolates non-transitive neutral-game dynamics while remaining tractable on standard hardware.

  5. Two-Player Zero-Sum Differential Games with One-Sided Information

    cs.GT 2025-02 conditional novelty 4.0 of 10

    A new algorithm, CAMS, reformulates the Bellman backup in one-sided-information differential games into small nonconvex minimax problems, achieving complexity independent of the continuous action space size and demons...

Pith tools