Pith. sign in

REVIEW 5 cited by

Avalon's Game of Thoughts: Battle Against Deception through Recursive Contemplation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.01320 v3 pith:RTCKXFNQ submitted 2023-10-02 cs.AI cs.CLcs.CYcs.LGcs.MA

classification cs.AIcs.CLcs.CYcs.LGcs.MA
keywords llmsavaloncontemplationdeceptivegameinformationreconefficacy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent breakthroughs in large language models (LLMs) have brought remarkable success in the field of LLM-as-Agent. Nevertheless, a prevalent assumption is that the information processed by LLMs is consistently honest, neglecting the pervasive deceptive or misleading information in human society and AI-generated content. This oversight makes LLMs susceptible to malicious manipulations, potentially resulting in detrimental outcomes. This study utilizes the intricate Avalon game as a testbed to explore LLMs' potential in deceptive environments. Avalon, full of misinformation and requiring sophisticated logic, manifests as a "Game-of-Thoughts". Inspired by the efficacy of humans' recursive thinking and perspective-taking in the Avalon game, we introduce a novel framework, Recursive Contemplation (ReCon), to enhance LLMs' ability to identify and counteract deceptive information. ReCon combines formulation and refinement contemplation processes; formulation contemplation produces initial thoughts and speech, while refinement contemplation further polishes them. Additionally, we incorporate first-order and second-order perspective transitions into these processes respectively. Specifically, the first-order allows an LLM agent to infer others' mental states, and the second-order involves understanding how others perceive the agent's mental state. After integrating ReCon with different LLMs, extensive experiment results from the Avalon game indicate its efficacy in aiding LLMs to discern and maneuver around deceptive information without extra fine-tuning and data. Finally, we offer a possible explanation for the efficacy of ReCon and explore the current limitations of LLMs in terms of safety, reasoning, speaking style, and format, potentially furnishing insights for subsequent research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CaM-Wolf: Causal-Aware Multimodal Agents for Social Deduction Games

    cs.AI 2026-07 conditional novelty 6.0 of 10

    CaM-Wolf is a multimodal Werewolf agent that perceives player video, reasons about hidden roles with a counterfactual-intervention-trained RL reasoner, and responds through an animated avatar.

  2. SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks

    cs.LG 2025-08 conditional novelty 6.0 of 10

    The authors propose SC2Arena, a full-coverage StarCraft II benchmark for LLMs, and StarEvolve, a planner-executor-verifier self-improvement framework, claiming superior strategic planning.

  3. Tactical Decision for Multi-UGV Confrontation with a Vision-Language Model-Based Commander

    cs.AI 2025-07 conditional novelty 5.0 of 10

    A VLM-plus-LLM commander trained with an expert rule system reaches 80-83% win rates in simulated multi-UGV confrontations, beating rule, RL, and vision-only baselines.

  4. Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers

    cs.AI 2025-02 conditional novelty 5.0 of 10

    A taxonomy-based survey of bidirectional game theory and LLM research, spanning evaluation, alignment, economic competition, and LLM-driven game solving.

  5. Verbal Werewolf: Engage Users with Verbalized Agentic Werewolf Game Framework

    cs.CL 2025-05 reject novelty 4.0 of 10

    A system that voices LLM-driven Werewolf agents in near real time using parallel gameplay and TTS pipelines, with only anecdotal evaluation.

Pith tools