Pith. sign in

Paper Citation Record · LEDGER

BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2408.15971.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.15971 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:14:47.229043Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T22:57:26.233385Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bdeb7504-dd83-401b-9027-98444f02d6a4 · inbound

Why Do Multi-Agent LLM Systems Fail? cites this paper.

Why Do Multi-Agent LLM Systems Fail? BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:57.912351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:74f763639d4c866fb0dfdccbef552712a10e6a143656086c5e867fcd5a19bd76

Observation ea404cf2-4c3f-4944-91f5-ceac3ea54560 · inbound

MAEBE: Multi-Agent Emergent Behavior Framework cites this paper.

MAEBE: Multi-Agent Emergent Behavior Framework BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:47.229043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:14:47.229043Z digest=sha256:31c76f96db89650c2a8e7f031dbc86eab891d2765587ed9d4e0afd1829025438

Observation 9583adb3-51b2-41e2-bde5-666e154bd097 · inbound

TextAtari: 100K Frames Game Playing with Language Agents cites this paper.

TextAtari: 100K Frames Game Playing with Language Agents BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T10:51:58.050301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:51:58.050301Z digest=sha256:84f343ae7dc7a28b4759311d909c10662b29099b8d77dc22c27e41afe1b2c052

Observation bdc7057b-65d2-454b-a014-bacf7035e2d2 · inbound

G-Memory: Tracing Hierarchical Memory for Multi-Agent Systems cites this paper.

G-Memory: Tracing Hierarchical Memory for Multi-Agent Systems BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:39:58.873471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:39:58.873471Z digest=sha256:cadff4e867fa9138f58bd7ded671d7ae97fca0e2874b2614beb05d8ca4b1d93c

Observation b26edf21-e77e-4e9b-ba80-c3c40b23c3bc · inbound

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey cites this paper.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 152

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.709367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.709367Z digest=sha256:8d5e6377e372f120ed76055fb03385a8b96ad5c5ec760cd5f99431ec3dc7eeba

Observation 5ddcaa60-af84-408e-804c-f494c3de74ea · inbound

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind cites this paper.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.475102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.475102Z digest=sha256:8a2b926ee46dbdd7a81b6fdda38409d468b49ec7f352d005f093b3a5a549438f

Observation 5818bfc3-d269-4ac5-a028-690f96bb1ba7 · inbound

Toward Adaptable Multi-Agent Reinforcement Learning: An Assumption-Aware Review cites this paper.

Toward Adaptable Multi-Agent Reinforcement Learning: An Assumption-Aware Review BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 150

Resolution
unresolved
no resolver link, observed 2026-08-06T17:42:58.980429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:42:58.980429Z digest=sha256:d25c6eae38440beb10cb94438a5ce7ffb8b57435fa6026f35aa37493e2d14c9b

Observation fea5e3e3-1c7c-4e94-8f15-652267ca4ce0 · inbound

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory cites this paper.

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 93

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T23:13:15.347355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-14T23:13:15.016486Z digest=sha256:f1a5113846314352207f9708e9b1551925a6050a5b365eaf679eaeb71b4b8129

Observation 8810323f-30ad-4717-974d-a0aca85c4436 · inbound

Agentic Reasoning for Large Language Models cites this paper.

Agentic Reasoning for Large Language Models BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:14:25.817356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T15:14:25.558878Z digest=sha256:a5c6edf003916c67606b81ae490a501901670625301a9d9c6a6e454201b858b3

Observation 6c817dfd-babd-4460-a9de-0eb0b64aecfe · inbound

CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V cites this paper.

CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:41:01.876999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T17:59:21.887765Z digest=sha256:a983ea67be41181c4238b7c754642d8acdb1874ac619240d96cea1c2044fd22f

Observation 1939d673-476e-47e9-9aff-740c74e61895 · inbound

Cooperate to Compete: Strategic Coordination in Multi-Agent Conquest cites this paper.

Cooperate to Compete: Strategic Coordination in Multi-Agent Conquest BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:36:36.074202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-07T16:37:58.860183Z digest=sha256:555ff031320b5a19f5cb0915da3026b2ec6bae360cbd3ea864543069eb78293e

Observation 878cc54b-7d8c-4338-a0f8-f425dfa61bde · inbound

When Does Hierarchy Help? Benchmarking Agent Coordination in Event-Driven Industrial Scheduling cites this paper.

When Does Hierarchy Help? Benchmarking Agent Coordination in Event-Driven Industrial Scheduling BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:23:37.908608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T02:19:10.912487Z digest=sha256:a30d443bd361da8ad70d93404dcf9019626ba9f90a271bbd8a38fa6064538734

Observation 6f036a5a-a79e-4408-ba1e-50cc7b7eb24f · inbound

Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems cites this paper.

Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 295

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T03:08:58.183231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T03:07:38.232966Z digest=sha256:7123039f73ddac9a22f75af05dceee9de3b6084c368d54c540841f29e93835c6

Observation a0f51e4a-ddc6-42f5-bf98-e83312e239df · inbound

Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems cites this paper.

Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 296

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T16:52:39.862066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T16:51:13.491389Z digest=sha256:d41c6b43abd9097f47e176a3cc88bb0504c6c073a9ffdaa9a2cb8bcde95f4d60

Observation e417114f-0aa8-484c-b6db-7c7270a90df1 · inbound

Interactive Evaluation Requires a Design Science cites this paper.

Interactive Evaluation Requires a Design Science BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:58:14.125624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T10:55:08.135630Z digest=sha256:e5665a301be2013c30da034a9a4955a76d5777021b251e1cfbd128f768c03471

Observation f55197ef-28be-4ac5-9d2f-ff2cc0a035a1 · inbound

CollabSim: A CSCW-Grounded Methodology for Investigating Collaborative Competence of LLM Agents through Controlled Multi-Agent Experiments cites this paper.

CollabSim: A CSCW-Grounded Methodology for Investigating Collaborative Competence of LLM Agents through Controlled Multi-Agent Experiments BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 102

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:16:58.884053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T01:25:51.788544Z digest=sha256:675e4d79baa715fdce478438cdd39a5d869c91e38ca6f75f9ef3850db2eef39e

Observation 54db5e46-b1ba-41bb-82a2-6991ed0d43bc · inbound

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents cites this paper.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:57:26.483468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:0363029ce7a686c075d9dd072364a84da86fe30e5399b9d0accbb7fad8245d1e

Observation 034bce39-cc41-46be-96ed-0315da7095d4 · inbound

RAILS: Verification-Native Clearing For Agentic Commerce cites this paper.

RAILS: Verification-Native Clearing For Agentic Commerce BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:26.234866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T18:35:13.781966Z digest=sha256:104f23cf029d6399d5ef9e9afc67e00905724edd467d66150a99bf1d106b2d3d