Pith. sign in

REVIEW 13 cited by

OpenSpiel: A Framework for Reinforcement Learning in Games

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1908.09453 v6 pith:IM2TDZG3 submitted 2019-08-26 cs.LG cs.AIcs.GTcs.MA

classification cs.LGcs.AIcs.GTcs.MA
keywords learningopenspielgamesreinforcementalgorithmsenvironmentssearchacross
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

OpenSpiel is a collection of environments and algorithms for research in general reinforcement learning and search/planning in games. OpenSpiel supports n-player (single- and multi- agent) zero-sum, cooperative and general-sum, one-shot and sequential, strictly turn-taking and simultaneous-move, perfect and imperfect information games, as well as traditional multiagent environments such as (partially- and fully- observable) grid worlds and social dilemmas. OpenSpiel also includes tools to analyze learning dynamics and other common evaluation metrics. This document serves both as an overview of the code base and an introduction to the terminology, core concepts, and algorithms across the fields of reinforcement learning, computational game theory, and search.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DungeonBench: A Benchmark for Rules-Rich Tactical Reasoning in Dungeons & Dragons Combat

    cs.AI 2026-07 conditional novelty 7.0 of 10

    DungeonBench scores LLM tactical play on D&D combat, finding frontier policies clear ~80% of single encounters but only 40% of linked multi-encounter days.

  2. Solving Zero-Sum Convex Markov Games

    cs.GT 2025-06 conditional novelty 7.0 of 10

    Independent policy-gradient algorithms provably compute approximate Nash equilibria in two-player zero-sum convex Markov games.

  3. Asymmetric Perturbation in Solving Bilinear Saddle-Point Optimization

    math.OC 2025-06 conditional novelty 7.0 of 10

    In bilinear saddle-point games, adding a small penalty to only one player's payoff makes that player's equilibrium exact and gives gradient ascent-descent a linear last-iterate rate.

  4. GUARD: Constructing Realistic Two-Player Matrix and Security Games for Benchmarking Game-Theoretic Algorithms

    cs.GT 2025-05 conditional novelty 7.0 of 10

    GUARD generates realistic security-game benchmarks from open data and shows that random games yield degenerate, defender-friendly equilibria.

  5. IFlowNets: Extending Generative Samplers to Learn Strategies in Incomplete Information Games

    cs.LG 2026-08 conditional novelty 6.0 of 10

    IFlowNets make generative flow networks work for imperfect-information games by adding an information-set aggregation constraint that restores valid flow matching.

  6. LeAct: Learning to Reason from Expert Actions

    cs.LG 2026-07 conditional novelty 6.0 of 10

    An AI can learn to reason by sampling explanations for an expert's actions and keeping only the ones that help it predict those actions.

  7. TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning

    cs.MA 2026-02 unverdicted novelty 6.0 of 10

    TABX is a JAX-based, GPU-accelerated, configurable multi-agent battle simulator that lets researchers vary units, terrain, and physics to benchmark cooperative MARL algorithms.

  8. NashPG: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria

    cs.LG 2025-10 unverdicted novelty 6.0 of 10

    Periodically re-centering the KL-regularizer on the current policy in self-play yields a policy-gradient algorithm that, in its exact form, provably converges to a Nash equilibrium.

  9. Benchmarking Partial Observability in Reinforcement Learning with a Suite of Memory-Improvable Domains

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A JAX benchmark suite of nine memory-improvable partially observable RL environments, with evidence that recurrent and transformer agents beat memoryless baselines but fall short of state-augmented ceilings.

  10. A Gold-Standard Study of What Makes a Lightweight Game-Playing Agent Strong

    cs.LG 2026-07 accept novelty 5.0 of 10

    A controlled study using a fixed Gin Rummy expert as a yardstick isolates which lightweight RL training choices help (TRPO, knock-first reward, curriculum, warm-start, best-checkpoint) and which fail (imitation, dense...

  11. FootsiesGym: A Fighting Game Benchmark for Two-Player Zero-Sum Imperfect-Information Games

    cs.AI 2026-07 accept novelty 5.0 of 10

    FootsiesGym is an open-source, vectorized fighting-game benchmark for two-player zero-sum imperfect-information RL that isolates non-transitive neutral-game dynamics while remaining tractable on standard hardware.

  12. Conservative Equilibrium Discovery in Offline Game-Theoretic Multiagent Reinforcement Learning

    cs.AI 2026-02 conditional novelty 5.0 of 10

    COffeE-PSRO combines conservative uncertainty penalties with robust replicator dynamics to extract lower-regret equilibrium profiles from offline multi-agent datasets.

  13. Multi-Agent Reinforcement Learning in Cybersecurity: From Fundamentals to Applications

    cs.MA 2025-05 conditional novelty 3.0 of 10

    A narrative survey of multi-agent reinforcement learning for cyber defense, reviewing game-theoretic models, cyber gyms, and applications, concluding MARL is promising but faces scalability and simulation-to-real tran...

Pith tools