Pith. sign in

Paper Citation Record · LEDGER

AvalonBench: Evaluating LLMs Playing the Game of Avalon

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 34 inbound Pith citation observations for arXiv:2310.05036.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.05036 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 34 of 34 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:36:25.898399Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T17:29:59.553004Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a1bdc827-7f92-4d8e-9400-42b3aba6e56d · inbound

Large Language Model based Multi-Agents: A Survey of Progress and Challenges cites this paper.

Large Language Model based Multi-Agents: A Survey of Progress and Challenges AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:58:55.946367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T06:58:54.921355Z digest=sha256:5cf3f9e4c0b279e2eac3edd55abb6357069f950062e00911a28e9d21a9e121b9

Observation 01b7f31b-91b7-4efb-aed8-cf060d0da3cb · inbound

A Survey on the Memory Mechanism of Large Language Model based Agents cites this paper.

A Survey on the Memory Mechanism of Large Language Model based Agents AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 172

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:21:39.679132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T07:21:39.440092Z digest=sha256:db2596b0618d6c633a6218c47fbcc4fcb1ebc742aa6d8d38d1957a19f63d343e

Observation 34943d93-bde3-4456-a1e2-d32ed692a665 · inbound

KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation cites this paper.

KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:25.898399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:25.898399Z digest=sha256:9765bdb5ff9f3d1e448f2c95b49be9317641da3add18be75a5aecd2df26b34ff

Observation 160bc2c0-759b-4f4d-9147-24b6198c4e06 · inbound

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models cites this paper.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:15.723111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:15.723111Z digest=sha256:69514229d5757b2cb4765091e94697c06194ecf0b0df28cd950347905e54543d

Observation f7e1c9e0-5ab3-441d-a0f4-7b0c98d53ea7 · inbound

ARIA: Training Language Agents with Intention-Driven Reward Aggregation cites this paper.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:02.061227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:02.061227Z digest=sha256:b80550b5f4493a3f4cc5581c51e7fcd7134128665e0158f9414ad9e4fe1ac552

Observation 9c92b787-ddad-496b-83e6-a9517e2cd855 · inbound

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models cites this paper.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:37.906915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:37.906915Z digest=sha256:f20321bc943a8d68fdede0d0de02e661745cffcad3a1f5558d849b354cafffe9

Observation ba95d4ab-71dc-4bd3-9418-a07a022a8bfc · inbound

Tracing LLM Reasoning Processes with Strategic Games: A Framework for Planning, Revision, and Resource-Constrained Decision Making cites this paper.

Tracing LLM Reasoning Processes with Strategic Games: A Framework for Planning, Revision, and Resource-Constrained Decision Making AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:46.441264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:46.441264Z digest=sha256:85a59ea432470ebd147bc8b591d43293f3772f7ffc2218cfa8e959742ffc40cc

Observation 290b5c75-e176-4b71-a860-a03d95f8ead9 · inbound

Bayesian Social Deduction with Graph-Informed Language Models cites this paper.

Bayesian Social Deduction with Graph-Informed Language Models AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:32:09.388738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T07:27:20.027179Z digest=sha256:d8433c24b6ac7e0bafb76b16b8252acfb8cb47a0dc864dc1262b60c438a2a5c1

Observation 6be7ce0b-2629-4ef1-9986-e986f01f6c38 · inbound

StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production-Living Simulations with Stardew Valley cites this paper.

StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production-Living Simulations with Stardew Valley AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:21.994778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:43:21.994778Z digest=sha256:b1d33d5959442a3ae6347859af2775b033fc1724319ffae27a6d2d863882d517

Observation a715f859-e9e0-480d-ab7f-5e80e2d89da0 · inbound

Toward Adaptable Multi-Agent Reinforcement Learning: An Assumption-Aware Review cites this paper.

Toward Adaptable Multi-Agent Reinforcement Learning: An Assumption-Aware Review AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 158

Resolution
unresolved
no resolver link, observed 2026-08-06T17:42:59.009288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:42:59.009288Z digest=sha256:4780d99aaf193b80653ba0c2626e359558a946a2ae39a719ee7f47485371a583

Observation 505180f6-239f-4f55-a4c2-653ea13f0b53 · inbound

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks cites this paper.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:29.614571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:29.614571Z digest=sha256:6e93c8e3600cf03c43b9c0b4406ffa78417bc2de1e35a6cc09ffd8ad493868e2

Observation ada10e2a-8fe7-4ff1-b0df-b473b714c77e · inbound

EmoMAS: Emotion-Aware Multi-Agent System for High-Stakes Edge-Deployable Negotiation with Bayesian Orchestration cites this paper.

EmoMAS: Emotion-Aware Multi-Agent System for High-Stakes Edge-Deployable Negotiation with Bayesian Orchestration AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:25:58.378836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:39:11.659294Z digest=sha256:3b82a6de0b4de12dfecf681119ecc8570efb7819a0dc9d4503eae7bb8cc58129

Observation 4f451fe1-7b98-41a2-a175-1233aff50a66 · inbound

Foresight Optimization for Strategic Reasoning in Large Language Models cites this paper.

Foresight Optimization for Strategic Reasoning in Large Language Models AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:45:27.784043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T13:43:43.769341Z digest=sha256:5dd4c3f9c74eec283c066a556460e02d0442c5a60523d0050ecf0f276937797f

Observation 15e07cb4-a7dd-4665-a187-63299dd729a4 · inbound

SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems cites this paper.

SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:48:01.645334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T08:45:54.303143Z digest=sha256:1f990349651f1f849f91a82f9718d599ed21865ee96b190bb57139b96af56238

Observation 03a35bec-b864-4d81-8aee-37c2234c7c85 · inbound

Trust, Lies, and Long Memories: Emergent Social Dynamics and Reputation in Multi-Round Avalon with LLM Agents cites this paper.

Trust, Lies, and Long Memories: Emergent Social Dynamics and Reputation in Multi-Round Avalon with LLM Agents AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:11:05.462688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T23:17:21.835439Z digest=sha256:e6da49ecaa57ee874ab033ef10376d04afde346a31fcab461d8b2e582a0a7f7f

Observation f5cdfb77-2a9a-482c-aca0-3a8724774c82 · inbound

Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks cites this paper.

Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:14:46.439004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T00:14:07.017420Z digest=sha256:cfc841630a07cca96a153c1477d5cd5bab011ea80513ccd0d611580a734da3bc

Observation 7cd5bd5e-3e33-434f-a1bb-d483448230ac · inbound

Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond cites this paper.

Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 236

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:07.961902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T12:02:07.027775Z digest=sha256:6b860f109710af0bc22e4336fafdf72ee24b2373ec68e068e8bfe621930dbee0

Observation 159e147e-aa2e-4a70-ab3e-c5401ab93f10 · inbound

Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond cites this paper.

Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 236

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:29:59.554653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-04T17:29:43.764085Z digest=sha256:f257a1d7c87cd6c57413b6e8a29040d230ef9aab06f19a3e6bbcdea4c22a6786

Observation 06ce7a50-4e3e-4e57-b69e-629ab8feccdd · inbound

Cooperate to Compete: Strategic Coordination in Multi-Agent Conquest cites this paper.

Cooperate to Compete: Strategic Coordination in Multi-Agent Conquest AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:36:36.092379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-07T16:37:58.860183Z digest=sha256:23583dfd9dd7f76f350e92922dba0537406a8b4204243363ec6ca31088a5911f

Observation 5c10b9a0-6e5e-466e-ab34-de7982eb352a · inbound

Common-agency Games for Multi-Objective Test-Time Alignment cites this paper.

Common-agency Games for Multi-Objective Test-Time Alignment AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:15:06.426047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-15T06:14:53.685486Z digest=sha256:67f10a33d7bce11e2ffa5697dcad7fd0f380ecade62fd5d9b108a01f72be8f37

Observation cb9cf47e-5ef0-4ca1-81e1-fd3c80f9aef2 · inbound

Scale-Dependent Collective Adaptation in Self-Amending LLM Societies: A Cross-Family Study of Emergent Governance cites this paper.

Scale-Dependent Collective Adaptation in Self-Amending LLM Societies: A Cross-Family Study of Emergent Governance AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-19T22:32:49.787160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T22:29:17.058647Z digest=sha256:faccd72168064549b71d2917c643ee6006d6bc2924dfc484b64ad3005de37536

Observation f56577ed-154c-4d46-b975-29068346c485 · inbound

GENSTRAT: Toward a Science of Strategic Reasoning in Large Language Models cites this paper.

GENSTRAT: Toward a Science of Strategic Reasoning in Large Language Models AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:45:20.667414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T04:41:35.532363Z digest=sha256:81f70aaf07602c4d2298461c6d3da15755635790efb93249e39df797a5aa4e02

Observation 71342bf4-7e38-447d-a111-91474e50b450 · inbound

QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction Agents cites this paper.

QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction Agents AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:03:47.993143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:59:51.871639Z digest=sha256:374d2d0c9986176e119209c9f1497093a3a4b16abfbb8e864229bfaf67e3e7e0

Observation 4ecfe833-ade7-4e2d-ab47-4aff26c4a9e3 · inbound

MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs cites this paper.

MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:23:13.235012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T07:15:27.939886Z digest=sha256:2703ec4bdb76ec6926358d243977cdacd7d5e1c5aa1334915efb7789fd8c8eb1

Observation 4ac8eaba-a506-4471-8c2b-27a32a88dbc7 · inbound

SAGE: A Quantitative Evaluation of Socialized Evolution in Agent Ecosystems cites this paper.

SAGE: A Quantitative Evaluation of Socialized Evolution in Agent Ecosystems AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:26:29.552145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T10:02:25.870386Z digest=sha256:e3fe06bea007126ed8a714fc00eeeff193cfd41d1975c24d892940481dbf6d8e

Observation 9065fdc1-89b4-41b1-9c2e-eab7f132aa78 · inbound

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application cites this paper.

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:58:02.906651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T09:46:30.702256Z digest=sha256:283a79237d29d01a980a6a104c5279a03628826f92a9808d0758715cca8e41eb

Observation a91d1000-7fc6-44a7-ac9c-6a0fa701b1f9 · inbound

RTSGameBench: An RTS Benchmark for Strategic Reasoning by Vision-Language Models cites this paper.

RTSGameBench: An RTS Benchmark for Strategic Reasoning by Vision-Language Models AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:39:17.307387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T21:02:26.081166Z digest=sha256:e037b7626d3b368b04db0b43141b5887a630f3313ca65f703284f7f822c2dd7b

Observation a92852ab-6d5e-4c1d-95c4-6f22d72fc59a · inbound

Enhancing Decision-Making with Large Language Models through Multi-Agent Fictitious Play cites this paper.

Enhancing Decision-Making with Large Language Models through Multi-Agent Fictitious Play AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:18.518428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T20:57:49.840546Z digest=sha256:f5fe5d7f4a40bee96704ea5271905edbf54ed2791e9190acab7c8094ff886555

Observation f8a15a83-5e08-4b2b-86cc-0b4598343d03 · inbound

Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games cites this paper.

Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:19:13.686463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T21:17:02.332687Z digest=sha256:ec7959758f129a0d642ff316c7ca47423830e216eec7a629f8bd73a70f7c02ef

Observation 53f0b17e-7126-4545-ab2f-a1bb757a927a · inbound

SidConArena: An Environment Evaluating Agents in Open-Ended,Positive-Sum Bargaining Game cites this paper.

SidConArena: An Environment Evaluating Agents in Open-Ended,Positive-Sum Bargaining Game AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T18:35:58.036755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T01:56:22.856352Z digest=sha256:e1fd90260878d001be13e8d3513a88325f66a9663985ea2501563c8c64735a79

Observation 6d241ef8-b7ed-47a5-8d3a-9c8ee46715cb · inbound

MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games cites this paper.

MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T10:15:59.479435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T10:15:59.479435Z digest=sha256:faabe7e4e60c61ea46b4df0c351f476719c91450968b1f325a4bbb366d597b36

Observation ab7c3c69-23ed-407f-b4ad-8852f0751c99 · inbound

MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games cites this paper.

MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T07:15:00.001267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:15:00.001267Z digest=sha256:9a6f9160af91b0f6fa523f303ef34fba391196f13614cf026594a658a50ac81b

Observation 18585a22-5ed8-4c90-bf85-0990a2baedef · inbound

Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems cites this paper.

Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T00:51:15.570493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:51:15.570493Z digest=sha256:2e209ca6afe78add3c7c42bfd36090d778cb916db5d8fbff05ba407d53f8aab1

Observation 086d3904-79f4-4f93-b6f9-b7db7a39c1c0 · inbound

Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game cites this paper.

Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-31T16:28:48.692047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:28:48.692047Z digest=sha256:4457a05871a09daa1f1160dbca041fee3e81614667ca9ec833087fe1c8b3ac4d