Pith. sign in

REVIEW 8 cited by

AvalonBench: Evaluating LLMs Playing the Game of Avalon

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.05036 v3 pith:3433H2UV submitted 2023-10-08 cs.AI cs.CL

classification cs.AIcs.CL
keywords gameavalonagentsavalonbenchplayingllmsbotsenvironment
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we explore the potential of Large Language Models (LLMs) Agents in playing the strategic social deduction game, Resistance Avalon. Players in Avalon are challenged not only to make informed decisions based on dynamically evolving game phases, but also to engage in discussions where they must deceive, deduce, and negotiate with other players. These characteristics make Avalon a compelling test-bed to study the decision-making and language-processing capabilities of LLM Agents. To facilitate research in this line, we introduce AvalonBench - a comprehensive game environment tailored for evaluating multi-agent LLM Agents. This benchmark incorporates: (1) a game environment for Avalon, (2) rule-based bots as baseline opponents, and (3) ReAct-style LLM agents with tailored prompts for each role. Notably, our evaluations based on AvalonBench highlight a clear capability gap. For instance, models like ChatGPT playing good-role got a win rate of 22.2% against rule-based bots playing evil, while good-role bot achieves 38.2% win rate in the same setting. We envision AvalonBench could be a good test-bed for developing more advanced LLMs (with self-playing) and agent frameworks that can effectively model the layered complexities of such game environments.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games

    cs.CL 2026-07 conditional novelty 7.0 of 10

    Non-invasive per-utterance belief probes in Mafia, auto-scored against engine truth, expose poorly calibrated LLM confidence and 1.5× over-prediction of being suspected.

  2. Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Frontier LLMs win Secret Hitler matches and can deceive, but most fail to keep a consistent false persona as evidence accumulates, with DRR often falling below 50%.

  3. Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Changing one LLM agent's secret objective in Werewolf lowers its team's win rate and changes its reasoning, while its public chat stays deceptively normal.

  4. SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks

    cs.LG 2025-08 conditional novelty 6.0 of 10

    The authors propose SC2Arena, a full-coverage StarCraft II benchmark for LLMs, and StarEvolve, a planner-executor-verifier self-improvement framework, claiming superior strategic planning.

  5. Toward Adaptable Multi-Agent Reinforcement Learning: An Assumption-Aware Review

    cs.AI 2025-07 conditional novelty 6.0 of 10

    A survey proposing adaptability as a three-part taxonomy (learning, policy, scenario-driven) for organizing and evaluating MARL under changing conditions.

  6. StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production-Living Simulations with Stardew Valley

    cs.AI 2025-07 conditional novelty 6.0 of 10

    StarDojo is a 1,000-task benchmark in Stardew Valley combining production and social activities, and the best tested MLLM (GPT-4.1) achieves only 12.7% success on its 100-task subset.

  7. WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A new wargame-based benchmark finds large language models score far below human experts on strategic reasoning across situation awareness, opponent modeling, and policy generation.

  8. Tracing LLM Reasoning Processes with Strategic Games: A Framework for Planning, Revision, and Resource-Constrained Decision Making

    cs.AI 2025-06 conditional novelty 4.0 of 10

    In a new three-game benchmark tracking planning, revision, and budget use across 12 LLMs, ChatGPT-o3-mini ranked highest, while overcorrecting models such as Qwen-Plus won few matches.

Pith tools