Pith. sign in

REVIEW 4 major objections 6 minor 66 references

Rapid Learning in Constrained Minimax Games with Negative Momentum

T0 review · 4 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Negative momentum, borrowed from unconstrained games, carries over to constrained minimax solvers and yields the first algorithm reported to surpass CFR+ across multiple game types, with MoCFR+ reaching about $10^9$ times lower…

desk verdict Empirically promising negative-momentum acceleration of RM+/CFR+, but the main theorem has a parameter-range error and the proof does not cover the algorithms that produce the headline results. read the letter →

arxiv 2501.00533 v1 pith:ZYUK5Y47 submitted 2024-12-31 cs.LG

classification cs.LG
keywords negativemomentumminimaxgamesconstrainedoptimizationregretmatchingcounterfactualminimizationextensive-formlast-iterateconvergenceexploitability
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper sets out to show that negative momentum, an acceleration trick known from unconstrained bilinear games, can be carried over to constrained minimax games and sharply speed up the standard solvers. The authors propose a momentum buffer kept in the dual space and updated with a negative coefficient $\beta<0$, plus a "Restarting Aggregated Momentum" (RAM) pattern that periodically re-anchors the buffer; the mechanism drops into online mirror descent, FTRL, and regret matching. The theoretical part proves exponential convergence to an approximate equilibrium for the entropy-regularized momentum algorithm with an infinitely long buffer, and convergence to Nash equilibrium with a sufficiently large buffer. The experimental part reports that the regret-matching variant MoCFR+ reaches roughly $10^9$ times lower exploitability than CFR+ and outperforms PCFR+ on Kuhn Poker, Leduc Poker, Goofspiel, and Liar's Dice, which the authors describe as the first algorithm to surpass CFR+ across these game types. Because CFR+ is the standard strong solver for imperfect-information games, the result matters as a simple route to faster equilibrium finding.

What carries the argument

The load-bearing object is the momentum buffer kept in the dual space, $\mu_t = \beta\mu_{t-1} - F(z_t)$ with the negative coefficient $\beta<0$, together with its "Restarting Aggregated Momentum" (RAM) extension that sums buffer snapshots over an interval of length $k$ and periodically re-anchors an attachment point $L_{\text{att}}$ in the FTRL form $L_t = L_{t-1} + F(z_t) - \beta(L_{\text{att}} - L_{t-1})$. The negative sign turns the usual momentum extrapolation into a friction-like drag that damps oscillation in the strategy trajectory. The machinery is transferred to regret matching through the known FTRL/OMD-to-RM+ correspondence, where the attachment regret vector $R_{\text{att}}$ plays the anchor role, producing Algorithm 2 (MoRM+).

What would settle it

Run MoCFR+ and CFR+ on the same four games with identical alternation, averaging, and hyperparameter-selection protocols; if the exploitability gap does not remain roughly $10^9$ across seeds, the headline claim fails. A sharper check is to write the MoRM+ update from Algorithm 2 as an instance of online linear optimization: if the negative-momentum term $-\beta(R_{\text{att}} - R_t)$ cannot be expressed as a valid loss sequence under the FTRL/OMD-to-RM+ correspondence, then the theoretical bridge that justifies MoRM+ is broken.

Watch

Extended reading notes

Core claim

The central claim is that negative momentum, stored as a dual-space buffer $\mu_t = \beta\mu_{t-1} - F(z_t)$ with $\beta<0$, is a universal acceleration mechanism for constrained zero-sum solvers. Plugged into mirror descent or FTRL, it yields updates $z_{t+1} = \arg\min_z\{\eta\langle z, -\mu_t\rangle + D_\psi(z,z_t)\}$; with negative entropy as the regularizer and an infinitely large buffer, the algorithm converges exponentially to the regularized equilibrium $z^*$ of a modified game (Theorem 2), that point is an $O(-\beta/\eta)$-Nash equilibrium of the original game (Theorem 3), and with the RAM buffer re-anchored every $k$ steps the algorithm converges to the set of Nash equilibria (Theorem 4). The same buffer is grafted onto the regret update of RM+, giving $R_{t+1} = [R_t + r(x_t) - \beta(R_{\text{att}} - R_t)]_+$, which defines MoRM+ and, through regret decomposition, MoCFR+; the experiments claim these variants dominate their base algorithms and state-of-the-art baselines, with MoCFR+ reaching roughly nine orders of magnitude lower exploitability than CFR+.

Load-bearing premise

The load-bearing premise is that the negative-momentum trick, proven to converge in one particular regularized mirror-descent setting, still works when grafted into the different regret-matching update that the headline experiments use; no theorem in the paper covers that grafted algorithm.

Editorial extensions

If this is right

  • MoCFR+ reaches exploitability about $10^9$ times lower than CFR+ and outperforms PCFR+ on all four tested extensive-form games, a claimed first for any algorithm across those game types.
  • The negative-momentum buffer applies uniformly to OMD and FTRL instantiations such as MoMWU and DMoGDA, so existing game solvers can be upgraded by altering only the buffer/loss update.
  • With negative entropy regularizer and an infinite buffer, convergence is exponential to an $O(-\beta/\eta)$-equilibrium; with finite $k$ under RAM, the algorithm provably reaches the exact Nash set.
  • Larger $|\beta|$ accelerates convergence but biases the regularized equilibrium away from Nash, exposing an explicit speed-accuracy tradeoff controlled by $\beta$ and $k$.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper proves convergence only for the entropy-regularized mirror-descent variant; whether the FTRL/OMD-to-RM+ correspondence survives with the negative-momentum term is left open, so the impressive MoCFR+ numbers currently rest on an analogy rather than a theorem.
  • An adaptive rule for the re-anchoring interval $k$ — for example, triggering a reset when exploitability stalls — could make the method free of per-game tuning, since the paper's own ablation shows a stair-step pattern when $k$ is too large.
  • The same dual-space buffer could be applied to Monte Carlo variants of CFR or to optimistic/predictive updates, where stochastic gradients would test whether the friction-like damping survives noise; that extension is not explored in the paper.
  • A protocol-stricter comparison — simultaneous updates, untuned $\beta$, and identical averaging schemes across algorithms — would clarify how much of the reported $10^9$ advantage comes from the negative momentum itself rather than from alternation and averaging choices.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper proposes a negative-momentum mechanism for constrained minimax games, instantiated as MoMD/MoFTRL (with entropy regularization), MoRM+, and their extensive-form counterparts MoCFR+ and DMoGDA/DMoMWU. The theoretical claims are exponential convergence of the entropy-regularized momentum algorithm to an approximate equilibrium (Theorems 2 and 3) and convergence to Nash equilibria with a sufficiently large restart interval (Theorem 4). The experimental sections report large exploitability reductions on normal-form and extensive-form benchmarks, most notably that MoCFR+ attains roughly 10^9 lower exploitability than CFR+ and outperforms PCFR+ across Kuhn Poker, Leduc Poker, Goofspiel, and Liar's Dice.

Significance. If the theoretical results were correct and the empirical claims reproducible, this would be a meaningful contribution: it extends negative-momentum ideas from unconstrained bilinear games to constrained decision sets, proposes a simple restarting-buffer mechanism, and demonstrates strong empirical performance across several standard game benchmarks. The paper ships code, uses standard benchmarks, and includes ablation studies for the restart interval. However, the central theorem as stated contains a parameter-range error, the proof of Theorem 4 is only a sketch, and the headline empirical algorithm MoCFR+ has no convergence theorem. The theory therefore does not currently support the paper's central claims as cleanly as the text suggests.

major comments (4)
  1. [Appendix D, proof of Theorem 2] The final step of the proof is algebraically incorrect. The proof establishes Dψ(z*,z_{t+1}) ≤ (1+β/2)DKL(z*,z_t) + C·DKL(z_{t+1},z_t) with C = -1 - (3/2)β - 4η²/β, and contraction requires C ≤ 0. The stated condition η ≤ sqrt(-(1+3β/2)β/2) does not imply this. For example, β = -0.5 and η = 0.2 are allowed by the stated bound (which equals 0.25), yet C = -1 + 0.75 + 0.32 = 0.07 > 0. A correct condition is η ≤ sqrt(-(1+3β/2)β/4), a factor √2 smaller. Since Theorem 3 inherits this condition from Theorem 2, the theoretical guarantee as stated is invalid and must be corrected.
  2. [Theorems 2-3 vs. Table 1] Even with the corrected bound, none of the MoMWU hyperparameters in Table 1 satisfy the theorem's conditions: for the size-25 random game, β = -0.02 and η = 7, while the corrected bound is about 0.0696; for the 3×3 game, β = -0.06 and η = 1, while the corrected bound is about 0.117. In addition, Theorems 2 and 3 assume k = ∞, whereas all Table 1 experiments use finite k. Thus the sentence 'The following theoretical analysis supports the experimental results above' is not justified by the stated theorems, and the paper should either relegate the theory to a qualitative role or select experimental hyperparameters that fall inside the proven regime.
  3. [Appendix D, proof of Theorem 4] Theorem 4 asserts convergence to the set of Nash equilibria with a finite restart interval, but its proof is a sketch. It invokes Lemma 7 and Lemma 8 as 'adapted from Abe et al. (2023)' without stating or proving them for the present algorithm, and it does not show that the periodic attachment update in Algorithm 1 produces the sequence z_att^{(n)} for which Lemma 7's strict decrease holds. The compactness and continuity argument is plausible but incomplete. The paper needs a full proof of these lemmas in the current setting, or a clear statement that Theorem 4 is conjectural.
  4. [Algorithm 2 and Eq. (21)] The headline algorithm MoCFR+ (and its normal-form counterpart MoRM+) has no convergence guarantee in the paper. The RM+/OMD correspondence in Eq. (16) and Lemma 5 is imported from Farina et al. (2023) for the unmodified regret-matching update; the paper does not prove that this correspondence remains valid after the negative-momentum term -β(R_att - R_t) is inserted in Algorithm 2 and Eq. (21). Since the paper's main empirical claim concerns MoCFR+, the connection between the proven entropy-regularized MoMD analysis and the actual headline algorithm is an analogy rather than a theorem. The empirical results can stand on their own, but the paper should not describe them as supported by the theoretical analysis.
minor comments (6)
  1. [Introduction, after Eq. (8)] The text refers to 'Figure 5(a)' when discussing the divergence of GDAm with non-negative momentum; this should be Figure 1(a), since Figure 5 is the ablation figure in Appendix C.
  2. [Appendix D] There are typos: 'Yong's inequality' should be 'Young's inequality', and in the proof of Theorem 4 there is a stray parenthesis in 'each z(n)_att )'.
  3. [Appendix C] The text contains 'learing rate' instead of 'learning rate' and 'The experiment results' instead of 'The experimental results.'
  4. [Preliminaries] The phrase 'norm-form' should be 'normal-form games'.
  5. [Table 1] Use consistent notation '3×3' instead of '3*3' for the matrix game.
  6. [Experiments, EFG section] The strong claim that this is 'the first instance where an algorithm surpasses CFR+ performance across various types of games' should be tempered or substantiated with a more precise comparison to existing strong CFR variants, since the claim goes beyond the experiments reported here.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: the convergence theorems and the MoCFR+ extension are independent of the experimental claims, though Theorem 2 has a separate algebraic gap.

full rationale

The paper's theoretical results (Theorems 2-4) are genuine Lyapunov-style convergence proofs for the entropy-regularized MoMWU/RAM algorithm to a regularized equilibrium defined by an anchor point; this is self-referential in the sense that the anchor is algorithm-dependent, but that is a standard regularized-saddle-point fixed-point analysis, not a fit or a parameter renamed as a prediction. The proof of Theorem 2 derives the contraction inequality from first-order optimality, Pinsker's inequality, and Young's inequality; it does not assume its own conclusion. The anchor z_att is either fixed (k=∞) or updated deterministically, and the modified game in Eq. (15) is a definition of the target equilibrium, not an input fitted to the experimental output. The extension to MoRM+/MoCFR+ is presented as an algorithmic analogy built on the external FTRL/OMD-to-RM+ correspondence of Farina et al. 2023 (Eq. 16, Lemma 5, Algorithm 2, Eq. 21); no theorem in the paper covers MoRM+/MoCFR+, so the empirical speedup is explicitly an experimental observation rather than a derived prediction of the entropy-regularizer theory. That is a correctness or support gap, not circularity. There is no load-bearing self-citation chain, no fitted parameter relabeled as a prediction, and no imported uniqueness theorem used to force the choice of algorithm. The algebraic error in Theorem 2's stated η bound, and the fact that the Table 1 and Table 2 hyperparameters violate even the corrected bound, are serious correctness defects but do not make the derivation circular. Therefore no significant circularity is present.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central theoretical results depend on standard monotonicity and entropy arguments, plus two ad-hoc assumptions: the transfer of the momentum framework to RM+/CFR+ without proof, and the unproven adaptation of magnet-descent lemmas for the periodic restart. The experimental claims depend on per-game tuned beta, k, and eta.

free parameters (3)
  • negative momentum coefficient beta = per game, e.g., -0.04 (3x3), -0.02 (random NFG), -0.2 (Kuhn)
    Controls the trade-off between convergence rate and duality gap of the modified equilibrium (Theorems 2-3); chosen by grid search per game (Tables 1-2).
  • restart interval k = per game, e.g., 50 (3x3), 100 (random NFG), 5-100 across EFGs
    Controls how often the attachment point L_att resets; ablation in Appendix C shows performance depends on k; no theoretical rule for choosing it.
  • learning rate eta = per game, e.g., 1.0 (3x3), 7.0 (random 25), 2.0 (Kuhn)
    Learning rate for MoMWU and DMoGDA; tuned per game; the convergence theorems impose a range but experiments do not report whether chosen eta satisfies it.
assumptions (4)
  • standard math F is monotone in the bilinear game, so the inner product bound used after Eq. (32) holds.
    Monotonicity of the gradient of a bilinear saddle-point problem is standard and used in the proof of Theorem 2.
  • domain assumption The regularized equilibrium z* of the modified game Eq. (15) satisfies the variational inequality invoked below Eq. (33).
    The proof assumes z* is the unique regularized equilibrium and uses its first-order condition; uniqueness is not proven for general G and anchors.
  • ad hoc to paper Lemmas 7 and 8, adapted from Abe et al. (2023), hold for the periodic attachment updates in Algorithm 1.
    Theorem 4's proof relies on these lemmas but does not prove the adaptation for the RAM restart mechanism or state the required k condition.
  • ad hoc to paper The RM+/OMD correspondence (Eq. 16 and Lemma 5 from Farina et al. 2023) remains valid when the negative-momentum term is added in Algorithm 2.
    MoRM+ is derived by analogy with this correspondence; no theorem covers the modified regret-matching update.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Rapid Learning in Constrained Minimax Games with Negative Momentum." pith.science (2026). https://pith.science/paper/ZYUK5Y47

@misc{pith2026250100533,
  author       = {Pith},
  title        = {Pith review of: Rapid Learning in Constrained Minimax Games with Negative Momentum},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ZYUK5Y47}},
  note         = {Machine review of arXiv:2501.00533}
}
read the original abstract

In this paper, we delve into the utilization of the negative momentum technique in constrained minimax games. From an intuitive mechanical standpoint, we introduce a novel framework for momentum buffer updating, which extends the findings of negative momentum from the unconstrained setting to the constrained setting and provides a universal enhancement to the classic game-solver algorithms. Additionally, we provide theoretical guarantee of convergence for our momentum-augmented algorithms with entropy regularizer. We then extend these algorithms to their extensive-form counterparts. Experimental results on both Normal Form Games (NFGs) and Extensive Form Games (EFGs) demonstrate that our momentum techniques can significantly improve algorithm performance, surpassing both their original versions and the SOTA baselines by a large margin.

Figures

Figures reproduced from arXiv: 2501.00533 by the authors.

Figure 1
Figure 1. The mechanic dynamic of a particle with different [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. The trajectory plots depict the MoMWU algorithm under varying negative momentum coefficients [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. The performance evaluation of the momentum [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Evaluating the momentum variants and baselines [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: The last-iterate convergence results of DMoGDA with different parameters k, in Kuhn Poker (left) and Leduc Poker [PITH_FULL_IMAGE:figures/full_fig_p014_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

66 extracted references · 56 canonical work pages

  1. [1]

    , " * write output.state after.block = add.period write newline

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Abe, K.; Ariu, K.; Sakamoto, M.; and Iwasaki, A. 2024. Adaptively Perturbed Mirror Descent for Learning in Games. arXiv:2305.16610

  4. [4]

    Abe, K.; Ariu, K.; Sakamoto, M.; Toyoshima, K.; and Iwasaki, A. 2023. Last-Iterate Convergence with Full and Noisy Feedback in Two-Player Zero-Sum Games. arXiv:2208.09855

  5. [5]

    D.; Hazan, E.; and Rakhlin, A

    Abernethy, J. D.; Hazan, E.; and Rakhlin, A. 2008. Competing in the Dark: An Efficient Algorithm for Bandit Linear Optimization. In Annual Conference on Learning Theory, 263--274. Omnipress

  6. [6]

    Balduzzi, D.; Racaniere, S.; Martens, J.; Foerster, J.; Tuyls, K.; and Graepel, T. 2018. The mechanics of n-player differentiable games. In International Conference on Machine Learning, 354--363. PMLR

  7. [7]

    Berry, M.; and Shukla, P. 2016. Curl force dynamics: symmetries, chaos and constants of motion. New Journal of Physics, 18(6): 063018

  8. [8]

    Brown, N.; and Sandholm, T. 2018. Superhuman AI for heads-up no-limit poker: Libratus beats top professionals. Science, 359(6374): 418--424

Show all 66 references
  1. [9]

    Burch, N.; Moravcik, M.; and Schmid, M. 2019. Revisiting CFR+ and alternating updates. Journal of Artificial Intelligence Research, 64: 429--443

  2. [10]

    Cai, Y.; Farina, G.; Grand-Clément, J.; Kroer, C.; Lee, C.-W.; Luo, H.; and Zheng, W. 2023. Last-Iterate Convergence Properties of Regret-Matching Algorithms in Games. arXiv:2311.00676

  3. [11]

    U.; Fleuret, F.; and Jaggi, M

    Chavdarova, T.; Pagliardini, M.; Stich, S. U.; Fleuret, F.; and Jaggi, M. 2021. Taming GANs with Lookahead-Minmax. arXiv:2006.14567

  4. [12]

    M.; Gidel, G.; Tracey, B.; Tuyls, K.; Omidshafiei, S.; Balduzzi, D.; and Jaderberg, M

    Czarnecki, W. M.; Gidel, G.; Tracey, B.; Tuyls, K.; Omidshafiei, S.; Balduzzi, D.; and Jaderberg, M. 2020. Real world games look like spinning tops. Advances in Neural Information Processing Systems, 33: 17443--17454

  5. [13]

    J.; and Golowich, N

    Daskalakis, C.; Foster, D. J.; and Golowich, N. 2020. Independent Policy Gradient Methods for Competitive Reinforcement Learning. In Advances in Neural Information Processing Systems

  6. [14]

    Daskalakis, C.; and Panageas, I. 2019. Last-Iterate Convergence: Zero-Sum Games and Constrained Min-Max Optimization. In Innovations in Theoretical Computer Science Conference, volume 124 of LIPIcs, 27:1--27:18. Schloss Dagstuhl - Leibniz-Zentrum f \" u r Informatik

  7. [15]

    S.; Chen, J.; Li, L.; Xiao, L.; and Zhou, D

    Du, S. S.; Chen, J.; Li, L.; Xiao, L.; and Zhou, D. 2017. Stochastic variance reduction methods for policy evaluation. In International Conference on Machine Learning, 1049--1058. PMLR

  8. [16]

    Farina, G.; Grand-Clément, J.; Kroer, C.; Lee, C.-W.; and Luo, H. 2023. Regret Matching+: (In)Stability and Fast Convergence in Games. arXiv:2305.14709

  9. [17]

    Farina, G.; Kroer, C.; Brown, N.; and Sandholm, T. 2019. Stable-Predictive Optimistic Counterfactual Regret Minimization. In International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, 1853--1862. PMLR

  10. [18]

    Farina, G.; Kroer, C.; and Sandholm, T. 2019 a . Online convex optimization for sequential decision processes and extensive-form games. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, 1917--1925

  11. [19]

    Farina, G.; Kroer, C.; and Sandholm, T. 2019 b . Optimistic Regret Minimization for Extensive-Form Games via Dilated Distance-Generating Functions. In Advances in Neural Information Processing Systems, 5222--5232

  12. [20]

    Farina, G.; Kroer, C.; and Sandholm, T. 2021 a . Better Regularization for Sequential Decision Spaces: Fast Convergence Rates for Nash, Correlated, and Team Equilibria. In EC '21: The 22nd ACM Conference on Economics and Computation , 432. ACM

  13. [21]

    Farina, G.; Kroer, C.; and Sandholm, T. 2021 b . Faster game solving via predictive blackwell approachability: Connecting regret matching and mirror descent. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, 5363--5371

  14. [22]

    Fiez, T.; and Ratliff, L. J. 2021. Local convergence analysis of gradient descent ascent with finite timescale separation. In International Conference on Learning Representation

  15. [23]

    N.; Chen, R

    Foerster, J. N.; Chen, R. Y.; Al-Shedivat, M.; Whiteson, S.; Abbeel, P.; and Mordatch, I. 2017. Learning with Opponent-Learning Awareness. In Adaptive Agents and Multi-Agent Systems

  16. [24]

    A.; Pezeshki, M.; Le Priol, R.; Huang, G.; Lacoste-Julien, S.; and Mitliagkas, I

    Gidel, G.; Hemmat, R. A.; Pezeshki, M.; Le Priol, R.; Huang, G.; Lacoste-Julien, S.; and Mitliagkas, I. 2019. Negative momentum for improved game dynamics. In The 22nd International Conference on Artificial Intelligence and Statistics, 1802--1811. PMLR

  17. [25]

    Golowich, N.; Pattathil, S.; and Daskalakis, C. 2020. Tight last-iterate convergence rates for no-regret learning in multi-player games. Advances in Neural Information Processing Systems, 33: 20766--20778

  18. [26]

    J.; Pouget - Abadie, J.; Mirza, M.; Xu, B.; Warde - Farley, D.; Ozair, S.; Courville, A

    Goodfellow, I. J.; Pouget - Abadie, J.; Mirza, M.; Xu, B.; Warde - Farley, D.; Ozair, S.; Courville, A. C.; and Bengio, Y. 2020. Generative adversarial networks. Communications of the ACM, 63(11): 139--144

  19. [27]

    Grand-Cl \'e ment, J.; and Kroer, C. 2024. Solving optimization problems with Blackwell approachability. Mathematics of Operations Research, 49(2): 697--728

  20. [28]

    Gulrajani, I.; Ahmed, F.; Arjovsky, M.; Dumoulin, V.; and Courville, A. C. 2017. Improved Training of Wasserstein GANs. In Advances in Neural Information Processing Systems, 5767--5777

  21. [29]

    Hart, S.; and Mas-Colell, A. 2000. A simple adaptive procedure leading to correlated equilibrium. Econometrica, 68(5): 1127--1150

  22. [30]

    Heusel, M.; Ramsauer, H.; Unterthiner, T.; Nessler, B.; and Hochreiter, S. 2017. GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium. In Advances in Neural Information Processing Systems, 6626--6637

  23. [31]

    Hoda, S.; Gilpin, A.; Pena, J.; and Sandholm, T. 2010. Smoothing techniques for computing Nash equilibria of sequential games. Mathematics of Operations Research, 35(2): 494--512

  24. [32]

    Huang, K.; and Zhang, S. 2022. New first-order algorithms for stochastic variational inequalities. SIAM Journal on Optimization, 32(4): 2745--2772

  25. [33]

    Korpelevich, G. M. 1976. The extragradient method for finding saddle points and other problems. Matecon, 12: 747--756

  26. [34]

    Kroer, C.; Peysakhovich, A.; Sodomka, E.; and Stier-Moses, N. E. 2019. Computing large market equilibria using abstractions. In ACM Conference on Economics and Computation, 745--746

  27. [35]

    Kroer, C.; Waugh, K.; K l n c -Karzan, F.; and Sandholm, T. 2020. Faster algorithms for extensive-form game solving via improved smoothing functions. Mathematical Programming, 179(1-2): 385--417

  28. [36]

    Kuhn, H. W. 1950. A simplified two-person poker. Contributions to the Theory of Games, 1: 97--103

  29. [37]

    D.; Saeta, B.; Bradbury, J.; Ding, D.; Borgeaud, S.; Lai, M.; Schrittwieser, J.; Anthony, T.; Hughes, E.; Danihelka, I.; and Ryan-Davis, J

    Lanctot, M.; Lockhart, E.; Lespiau, J.-B.; Zambaldi, V.; Upadhyay, S.; Pérolat, J.; Srinivasan, S.; Timbers, F.; Tuyls, K.; Omidshafiei, S.; Hennes, D.; Morrill, D.; Muller, P.; Ewalds, T.; Faulkner, R.; Kramár, J.; Vylder, B. D.; Saeta, B.; Bradbury, J.; Ding, D.; Borgeaud, S...

  30. [38]

    Lanctot, M.; Waugh, K.; Zinkevich, M.; and Bowling, M. H. 2009. Monte Carlo Sampling for Regret Minimization in Extensive Games. In Advances in Neural Information Processing Systems, 1078--1086. Curran Associates, Inc

  31. [39]

    Lee, C.; Kroer, C.; and Luo, H. 2021. Last-iterate Convergence in Extensive-Form Games. In Advances in Neural Information Processing Systems, 14293--14305

  32. [40]

    Liang, T.; and Stokes, J. 2019. Interaction matters: A note on non-asymptotic local convergence of generative adversarial networks. In International Conference on Artificial Intelligence and Statistics, 907--915. PMLR

  33. [41]

    Lis \' y , V.; Lanctot, M.; and Bowling, M. H. 2015. Online Monte Carlo Counterfactual Regret Minimization for Search in Imperfect Information Games. In International Conference on Autonomous Agents and Multiagent Systems, 27--36. ACM

  34. [42]

    Liu, M.; Ozdaglar, A.; Yu, T.; and Zhang, K. 2023. The Power of Regularization in Solving Extensive-Form Games. arXiv:2206.09495

  35. [43]

    Liu, W.; Jiang, H.; Li, B.; and Li, H. 2022. Equivalence Analysis between Counterfactual Regret Minimization and Online Mirror Descent. In International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, 13717--13745. PMLR

  36. [44]

    Lockhart, E.; Lanctot, M.; P \' e rolat, J.; Lespiau, J.; Morrill, D.; Timbers, F.; and Tuyls, K. 2019. Computing Approximate Equilibria in Sequential Adversarial Games by Exploitability Descent. In International Joint Conference on Artificial Intelligence, 464--470

  37. [45]

    P.; Acuna, D.; Vicol, P.; and Duvenaud, D

    Lorraine, J. P.; Acuna, D.; Vicol, P.; and Duvenaud, D. 2022. Complex momentum for optimization in games. In International Conference on Artificial Intelligence and Statistics, 7742--7765. PMLR

  38. [46]

    Madras, D.; Creager, E.; Pitassi, T.; and Zemel, R. S. 2018. Learning Adversarially Fair and Transferable Representations. In International Conference on Machine Learning, 3381--3390. PMLR

  39. [47]

    Mertikopoulos, P.; Lecouat, B.; Zenati, H.; Foo, C.; Chandrasekhar, V.; and Piliouras, G. 2019. Optimistic mirror descent in saddle-point problems: Going the extra (gradient) mile. In International Conference on Learning Representations

  40. [48]

    M.; Nowozin, S.; and Geiger, A

    Mescheder, L. M.; Nowozin, S.; and Geiger, A. 2017. The Numerics of GANs. In Advances in Neural Information Processing Systems, 1825--1835

  41. [49]

    Morav c \' k, M.; Schmid, M.; Burch, N.; Lis \`y , V.; Morrill, D.; Bard, N.; Davis, T.; Waugh, K.; Johanson, M.; and Bowling, M. 2017. Deepstack: Expert-level artificial intelligence in heads-up no-limit poker. Science, 356(6337): 508--513

  42. [50]

    Orabona, F. 2023. A Modern Introduction to Online Learning. arXiv:1912.13213

  43. [51]

    Peng, W.; Dai, Y.-H.; Zhang, H.; and Cheng, L. 2020. Training GANs with centripetal acceleration. Optimization Methods and Software, 35(5): 955--973

  44. [52]

    P \'e rolat, J.; Munos, R.; Lespiau, J.-B.; Omidshafiei, S.; Rowland, M.; Ortega, P.; Burch, N.; Anthony, T.; Balduzzi, D.; De Vylder, B.; et al. 2021. From poincar \'e recurrence to convergence in imperfect information games: Finding equilibrium via regularization. In Interna...

  45. [53]

    Polyak, B. T. 1964. Some methods of speeding up the convergence of iteration methods. Ussr computational mathematics and mathematical physics, 4(5): 1--17

  46. [54]

    Ross, S. M. 1971. Goofspiel—the game of pure strategy. Journal of Applied Probability, 8(3): 621--625

  47. [55]

    Sch \" a fer, F.; and Anandkumar, A. 2019. Competitive Gradient Descent. In Advances in Neural Information Processing Systems, 7623--7633

  48. [56]

    S.; Su, W

    Shi, B.; Du, S. S.; Su, W. J.; and Jordan, M. I. 2019. Acceleration via Symplectic Discretization of High-Resolution Differential Equations. In Advances in Neural Information Processing Systems

  49. [57]

    Sinha, A.; Namkoong, H.; and Duchi, J. C. 2018. Certifying Some Distributional Robustness with Principled Adversarial Training. In International Conference on Learning Representations

  50. [58]

    Z.; Loizou, N.; Lanctot, M.; Mitliagkas, I.; Brown, N.; and Kroer, C

    Sokota, S.; D'Orazio, R.; Kolter, J. Z.; Loizou, N.; Lanctot, M.; Mitliagkas, I.; Brown, N.; and Kroer, C. 2023. A Unified Approach to Reinforcement Learning, Quantal Response Equilibria, and Two-Player Zero-Sum Games. In International Conference on Learning Representations,

  51. [59]

    P.; Larson, B.; Piccione, C.; Burch, N.; Billings, D.; and Rayner, C

    Southey, F.; Bowling, M. P.; Larson, B.; Piccione, C.; Burch, N.; Billings, D.; and Rayner, C. 2012. Bayes' Bluff: Opponent Modelling in Poker. arXiv:1207.1411

  52. [60]

    Tammelin, O. 2014. Solving Large Imperfect Information Games Using CFR+. arXiv:1407.5042

  53. [61]

    Neumann, J

    v. Neumann, J. 1928. Zur theorie der gesellschaftsspiele. Mathematische annalen, 100(1): 295--320

  54. [62]

    Vlatakis - Gkaragkounis, E.; Flokas, L.; and Piliouras, G. 2019. Poincar \' e Recurrence, Cycles and Spurious Equilibria in Gradient-Descent-Ascent for Non-Convex Non-Concave Zero-Sum Games. In Advances in Neural Information Processing Systems, 10450--10461

  55. [63]

    K.; Jagota, A

    Warmuth, M. K.; Jagota, A. K.; et al. 1997. Continuous and discrete-time nonlinear gradient descent: Relative loss bounds and convergence. In International Symposium on Artificial Intelligence and Mathematics, volume 326. Citeseer

  56. [64]

    Wei, C.; Lee, C.; Zhang, M.; and Luo, H. 2021. Linear Last-iterate Convergence in Constrained Saddle-point Optimization. In International Conference on Learning Representations

  57. [65]

    Zhang, G.; and Wang, Y. 2021. On the suboptimality of negative momentum for minimax optimization. In International Conference on Artificial Intelligence and Statistics, 2098--2106. PMLR

  58. [66]

    Zinkevich, M.; Johanson, M.; Bowling, M.; and Piccione, C. 2007. Regret minimization in games with incomplete information. Advances in Neural Information Processing Systems, 20

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.