Pith. sign in

REVIEW 12 cited by

GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.12348 v2 pith:DGPBQVK3 submitted 2024-02-19 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords llmsreasoningstrategicgame-theoreticgamescompetitivescenariosversus
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As Large Language Models (LLMs) are integrated into critical real-world applications, their strategic and logical reasoning abilities are increasingly crucial. This paper evaluates LLMs' reasoning abilities in competitive environments through game-theoretic tasks, e.g., board and card games that require pure logic and strategic reasoning to compete with opponents. We first propose GTBench, a language-driven environment composing 10 widely recognized tasks, across a comprehensive game taxonomy: complete versus incomplete information, dynamic versus static, and probabilistic versus deterministic scenarios. Then, we (1) Characterize the game-theoretic reasoning of LLMs; and (2) Perform LLM-vs.-LLM competitions as reasoning evaluation. We observe that (1) LLMs have distinct behaviors regarding various gaming scenarios; for example, LLMs fail in complete and deterministic games yet they are competitive in probabilistic gaming scenarios; (2) Most open-source LLMs, e.g., CodeLlama-34b-Instruct and Llama-2-70b-chat, are less competitive than commercial LLMs, e.g., GPT-4, in complex games, yet the recently released Llama-3-70b-Instruct makes up for this shortcoming. In addition, code-pretraining greatly benefits strategic reasoning, while advanced reasoning methods such as Chain-of-Thought (CoT) and Tree-of-Thought (ToT) do not always help. We further characterize the game-theoretic properties of LLMs, such as equilibrium and Pareto Efficiency in repeated games. Detailed error profiles are provided for a better understanding of LLMs' behavior. We hope our research provides standardized protocols and serves as a foundation to spur further explorations in the strategic reasoning of LLMs.

Discussion (0). Sign in to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Scaffolding the Strategist: Architecture-Dependent Reasoning Interventions in Hotelling Spatial Markets

    cs.AI 2026-07 conditional novelty 7.0 of 10

    In Hotelling spatial markets, commitment scaffolding improves a standard LLM but degrades a reasoning-optimized one, while principled separation shows the opposite crossover; adversarial self-critique harms both, more...

  2. Strategy, Not Payoffs: A Behavioural Embedding of Normal-Form Games

    cs.GT 2026-07 conditional novelty 6.0 of 10

    A two-feature game embedding (Nash entropy and best-response switching) predicts cross-game transfer of fine-tuned LLMs on held-out games, outperforming game identity and published structural embeddings.

  3. Democratizing Diplomacy: A Harness for Evaluating Any Large Language Model on Full-Press Diplomacy

    cs.AI 2025-08 conditional novelty 6.0 of 10

    An evaluation harness lets off-the-shelf local LLMs, including a 24B model, play full-press Diplomacy without fine-tuning.

  4. How Large Language Models play humans in online conversations: a simulated study of the 2016 US politics on Reddit

    cs.CL 2025-06 conditional novelty 6.0 of 10

    GPT-4 impersonating Reddit users in 2016 election threads produces comments that lean toward consensus and are semantically separable from real human comments.

  5. The Effect of State Representation on LLM Agent Behavior in Dynamic Routing Games

    cs.AI 2025-06 conditional novelty 6.0 of 10

    In a repeated Braess routing game, LLM agents given summarized, regret-based, and own-action-only state representations converge closer to Nash equilibrium and behave more stably than agents given full chat transcript...

  6. WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A new wargame-based benchmark finds large language models score far below human experts on strategic reasoning across situation awareness, opponent modeling, and policy generation.

  7. DipSVD: Dual-importance Protected SVD for Efficient LLM Compression

    cs.LG 2025-06 reject novelty 5.0 of 10

    DipSVD combines channel-weighted whitening with layer-wise compression ratios and reports better perplexity and accuracy than existing SVD-based LLM compression methods.

  8. Shapley-Coop: Credit Assignment for Emergent Cooperation in Self-Interested LLM Agents

    cs.MA 2025-06 reject novelty 5.0 of 10

    Shapley-Coop asks LLM agents to negotiate prices for contributions based on Shapley value reasoning, improving cooperation and reward fairness in three multi-agent tasks.

  9. Grammars of Formal Uncertainty: When to Trust LLMs in Automated Reasoning Tasks

    cs.CL 2025-05 reject novelty 5.0 of 10

    A grammar-based model of LLM-generated SMT-LIB code produces uncertainty signals that predict formalization errors on some reasoning tasks, with fused signals giving large error reductions only in an in-sample evaluation.

  10. Tracing LLM Reasoning Processes with Strategic Games: A Framework for Planning, Revision, and Resource-Constrained Decision Making

    cs.AI 2025-06 conditional novelty 4.0 of 10

    In a new three-game benchmark tracking planning, revision, and budget use across 12 LLMs, ChatGPT-o3-mini ranked highest, while overcorrecting models such as Qwen-Plus won few matches.

  11. Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A survey that classifies AI agent evaluation benchmarks along environment and capability axes, and proposes five traits that distinguish agents from chatbots.

  12. GenAI-based Multi-Agent Reinforcement Learning towards Distributed Agent Intelligence: A Generative-RL Agent Perspective

    cs.AI 2025-07 unverdicted novelty 3.0 of 10

    A position paper claiming that generative-AI agents that model and predict multi-agent dynamics will replace today's reactive MARL approaches.

Pith tools