REVIEW 11 cited by
GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large language models have demonstrated remarkable few-shot performance on many natural language understanding tasks. Despite several demonstrations of using large language models in complex, strategic scenarios, there lacks a comprehensive framework for evaluating agents' performance across various types of reasoning found in games. To address this gap, we introduce GameBench, a cross-domain benchmark for evaluating strategic reasoning abilities of LLM agents. We focus on 9 different game environments, where each covers at least one axis of key reasoning skill identified in strategy games, and select games for which strategy explanations are unlikely to form a significant portion of models' pretraining corpuses. Our evaluations use GPT-3 and GPT-4 in their base form along with two scaffolding frameworks designed to enhance strategic reasoning ability: Chain-of-Thought (CoT) prompting and Reasoning Via Planning (RAP). Our results show that none of the tested models match human performance, and at worst GPT-4 performs worse than random action. CoT and RAP both improve scores but not comparable to human levels.
Forward citations
Cited by 11 Pith papers
-
CHBench: A Cognitive Hierarchy Benchmark for Evaluating Strategic Reasoning Capability of LLMs
CHBench fits Level-K and Poisson cognitive hierarchy models to LLM game play and uses the fitted reasoning level as a benchmark score.
-
Democratizing Diplomacy: A Harness for Evaluating Any Large Language Model on Full-Press Diplomacy
An evaluation harness lets off-the-shelf local LLMs, including a 24B model, play full-press Diplomacy without fine-tuning.
-
A Practical Approach for Building Production-Grade Conversational Agents with Workflow Graphs
A production e-commerce chatbot using a workflow graph with node-specific prompts and response-masked fine-tuning reports large gains in accuracy and format compliance, beating a GPT-4o-based agent in human preference tests.
-
lmgame-Bench: How Good are LLMs at Playing Games?
lmgame-Bench turns six classic games into a scaffolded LLM evaluation suite, ranks 13 models, detects contamination, and reports RL transfer from Sokoban or Tetris to unseen games and planning tasks.
-
Spatial Reasoning in LLM Game Agents: Impact of Causal Context and Multi-Step Planning
On a new three-game spatial benchmark, larger Qwen3 models with thinking mode and multi-step planning achieve higher win rates, while small models struggle to localize and causal prompt hints give only marginal, model...
-
Talking-to-Build: How LLM-Assisted Interface Shapes Player Performance and Experience in Minecraft
A 30-participant Minecraft study found that an LLM chat interface improved self-reported game experience compared with typed commands, while objective task performance was not measured.
-
Agents Require Metacognitive and Strategic Reasoning to Succeed in the Coming Labor Markets
AI agents in future labor markets will need metacognitive and strategic reasoning because incomplete information creates adverse selection, moral hazard, and reputation effects.
-
Game Reasoning Arena: A Framework and Benchmark for Assessing Reasoning Capabilities of Large Language Models via Game Play
Game Reasoning Arena is a modular OpenSpiel-based framework for benchmarking LLM decision making in games, with exploratory analyses suggesting models adapt their verbalized reasoning to game structure and model size.
-
Evaluation and Benchmarking of LLM Agents: A Survey
A review that proposes a two-dimensional taxonomy for evaluating LLM agents and highlights enterprise-specific evaluation gaps.
-
Tracing LLM Reasoning Processes with Strategic Games: A Framework for Planning, Revision, and Resource-Constrained Decision Making
In a new three-game benchmark tracking planning, revision, and budget use across 12 LLMs, ChatGPT-o3-mini ranked highest, while overcorrecting models such as Qwen-Plus won few matches.
-
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey
A survey that classifies AI agent evaluation benchmarks along environment and capability axes, and proposes five traits that distinguish agents from chatbots.
Discussion (0). Continue with ORCID to comment.