Pith. sign in

REVIEW 19 cited by

Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.04658 v2 pith:7X65DZZ6 submitted 2023-09-09 cs.CL

classification cs.CL
keywords communicationgamesllmslanguagewerewolfempiricalengageframework
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Communication games, which we refer to as incomplete information games that heavily depend on natural language communication, hold significant research value in fields such as economics, social science, and artificial intelligence. In this work, we explore the problem of how to engage large language models (LLMs) in communication games, and in response, propose a tuning-free framework. Our approach keeps LLMs frozen, and relies on the retrieval and reflection on past communications and experiences for improvement. An empirical study on the representative and widely-studied communication game, ``Werewolf'', demonstrates that our framework can effectively play Werewolf game without tuning the parameters of the LLMs. More importantly, strategic behaviors begin to emerge in our experiments, suggesting that it will be a fruitful journey to engage LLMs in communication games and associated domains.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 23 citations worldwide. Full citation record

  1. Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Training single-layer attention with squared regret loss has stationary points that implement smoothed fictitious play (external regret) and, via a new swap-regret loss, the Blum–Mansour no-swap-regret algorithm.

  2. CaM-Wolf: Causal-Aware Multimodal Agents for Social Deduction Games

    cs.AI 2026-07 conditional novelty 6.0 of 10

    CaM-Wolf is a multimodal Werewolf agent that perceives player video, reasons about hidden roles with a counterfactual-intervention-trained RL reasoner, and responds through an animated avatar.

  3. Reproducing human biases in route choice using large language models: Toward scalable behavioral modeling

    cs.AI 2026-07 conditional novelty 6.0 of 10

    LLM agents with demographic profiles reproduce CPT-style risk attitudes in route choice and yield fitted parameters (α=0.4, β=0.64, λ=1.43) that predict human data competitively.

  4. SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks

    cs.LG 2025-08 conditional novelty 6.0 of 10

    The authors propose SC2Arena, a full-coverage StarCraft II benchmark for LLMs, and StarEvolve, a planner-executor-verifier self-improvement framework, claiming superior strategic planning.

  5. Strategy Adaptation in Large Language Model Werewolf Agents

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Dynamically switching between Support and Attack strategies based on role estimates raises win rates for Werewolf LLM agents, with mixed effects for Villagers.

  6. The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A new interactive language-game benchmark shows LLMs lag behind simple word-embedding baselines and that newer reasoning models regress on theory-of-mind tasks.

  7. IndoorWorld: Integrating Physical Task Solving and Social Simulation in A Heterogeneous Multi-Agent Environment

    cs.MA 2025-06 conditional novelty 6.0 of 10

    IndoorWorld is a new multi-agent environment that combines physical task solving with social interaction, and its experiments show effects of collaboration, resource competition, and layout on agent behavior.

  8. Sword and Shield: Uses and Strategies of LLMs in Navigating Disinformation

    cs.HC 2025-06 conditional novelty 6.0 of 10

    In a 25-participant Werewolf-style game, all roles used an LLM chatbot strategically, as a sword for disinformation and a shield against it.

  9. Empowering Economic Simulation for Massively Multiplayer Online Games through Generative Agent-Based Modeling

    cs.AI 2025-06 conditional novelty 6.0 of 10

    LLM-driven agents in a simulated MMO economy reproduce role specialization and price responses to supply and demand, though the price result is partly shaped by what the AI is told.

  10. SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    SocialMaze is a six-task benchmark that claims to evaluate LLM social reasoning along deep reasoning, dynamic interaction, and information uncertainty dimensions.

  11. Cracking Aegis: An Adversarial LLM-based Game for Raising Awareness of Vulnerabilities in Privacy Protection

    cs.HC 2025-05 conditional novelty 6.0 of 10

    Cracking Aegis, an adversarial LLM-driven dialogue game, led players to use manipulative language strategies and to self-report stronger awareness of privacy vulnerabilities after a single session.

  12. Cumulative suspicion and absorption dynamics in an agent-based Mafia game

    physics.soc-ph 2026-07 conditional novelty 5.5 of 10

    History-dependent suspicion scores in an agent Mafia model yield F(τ)∼(τ/N)^{N_m} early extinction without detectives and an empirical N_c collapse of win probabilities that detectives break.

  13. Ethical Considerations of Large Language Models in Game Playing

    cs.CL 2025-08 conditional novelty 5.0 of 10

    In Werewolf games, LLM agents change their kills, votes, and trust scores based on explicit gender labels and even based on gender-implied first names, behaving differently for male and female players.

  14. SimuPanel: A Novel Immersive Multi-Agent System to Simulate Interactive Expert Panel Discussion

    cs.HC 2025-06 conditional novelty 5.0 of 10

    A multi-agent LLM system called SimuPanel simulates expert panel discussions with personas grounded in public academic sources, and a small evaluation suggests the full reasoning pipeline produces higher LLM-judged di...

  15. KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation

    cs.CL 2025-05 conditional novelty 5.0 of 10

    KORGym introduces a 51-game, text and visual, multi-turn benchmark with a normalized scoring scheme, and uses it to compare 19 LLMs and 8 VLMs on six reasoning dimensions.

  16. Tracing LLM Reasoning Processes with Strategic Games: A Framework for Planning, Revision, and Resource-Constrained Decision Making

    cs.AI 2025-06 conditional novelty 4.0 of 10

    In a new three-game benchmark tracking planning, revision, and budget use across 12 LLMs, ChatGPT-o3-mini ranked highest, while overcorrecting models such as Qwen-Plus won few matches.

  17. Artificial Intelligence and Civil Discourse: How LLMs Moderate Climate Change Conversations

    cs.CY 2025-06 reject novelty 4.0 of 10

    LLM replies to climate change posts are more emotionally neutral and lower in intensity than the human posts they respond to.

  18. AI Agent Behavioral Science

    q-bio.NC 2025-06 conditional novelty 4.0 of 10

    AI agents should be studied as behavioral entities shaped by context and interaction, not only as trained models.

  19. Verbal Werewolf: Engage Users with Verbalized Agentic Werewolf Game Framework

    cs.CL 2025-05 reject novelty 4.0 of 10

    A system that voices LLM-driven Werewolf agents in near real time using parallel gameplay and TTS pipelines, with only anecdotal evaluation.

Pith tools