REVIEW 11 cited by
AvalonBench: Evaluating LLMs Playing the Game of Avalon
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this paper, we explore the potential of Large Language Models (LLMs) Agents in playing the strategic social deduction game, Resistance Avalon. Players in Avalon are challenged not only to make informed decisions based on dynamically evolving game phases, but also to engage in discussions where they must deceive, deduce, and negotiate with other players. These characteristics make Avalon a compelling test-bed to study the decision-making and language-processing capabilities of LLM Agents. To facilitate research in this line, we introduce AvalonBench - a comprehensive game environment tailored for evaluating multi-agent LLM Agents. This benchmark incorporates: (1) a game environment for Avalon, (2) rule-based bots as baseline opponents, and (3) ReAct-style LLM agents with tailored prompts for each role. Notably, our evaluations based on AvalonBench highlight a clear capability gap. For instance, models like ChatGPT playing good-role got a win rate of 22.2% against rule-based bots playing evil, while good-role bot achieves 38.2% win rate in the same setting. We envision AvalonBench could be a good test-bed for developing more advanced LLMs (with self-playing) and agent frameworks that can effectively model the layered complexities of such game environments.
Forward citations
Cited by 11 Pith papers
-
MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games
Non-invasive per-utterance belief probes in Mafia, auto-scored against engine truth, expose poorly calibrated LLM confidence and 1.5× over-prediction of being suspected.
-
Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game
Frontier LLMs win Secret Hitler matches and can deceive, but most fail to keep a consistent false persona as evidence accumulates, with DRR often falling below 50%.
-
Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems
Changing one LLM agent's secret objective in Werewolf lowers its team's win rate and changes its reasoning, while its public chat stays deceptively normal.
-
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks
The authors propose SC2Arena, a full-coverage StarCraft II benchmark for LLMs, and StarEvolve, a planner-executor-verifier self-improvement framework, claiming superior strategic planning.
-
Toward Adaptable Multi-Agent Reinforcement Learning: An Assumption-Aware Review
A survey proposing adaptability as a three-part taxonomy (learning, policy, scenario-driven) for organizing and evaluating MARL under changing conditions.
-
StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production-Living Simulations with Stardew Valley
StarDojo is a 1,000-task benchmark in Stardew Valley combining production and social activities, and the best tested MLLM (GPT-4.1) achieves only 12.7% success on its 100-task subset.
-
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models
A new wargame-based benchmark finds large language models score far below human experts on strategic reasoning across situation awareness, opponent modeling, and policy generation.
-
ARIA: Training Language Agents with Intention-Driven Reward Aggregation
Clustering language-agent actions into shared intentions and averaging their rewards reduces reward variance and improves policy performance in open-ended dialogue tasks.
-
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models
SocialMaze is a six-task benchmark that claims to evaluate LLM social reasoning along deep reasoning, dynamic interaction, and information uncertainty dimensions.
-
KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation
KORGym introduces a 51-game, text and visual, multi-turn benchmark with a normalized scoring scheme, and uses it to compare 19 LLMs and 8 VLMs on six reasoning dimensions.
-
Tracing LLM Reasoning Processes with Strategic Games: A Framework for Planning, Revision, and Resource-Constrained Decision Making
In a new three-game benchmark tracking planning, revision, and budget use across 12 LLMs, ChatGPT-o3-mini ranked highest, while overcorrecting models such as Qwen-Plus won few matches.
Discussion (0). Sign in to comment.