Pith. sign in

Paper Citation Record · LEDGER

GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2406.06613.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.06613 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T13:48:54.886214Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7441e011-4bcb-4002-89b4-389f383a3360 · inbound

Cardiverse: Harnessing LLMs for Novel Card Game Prototyping cites this paper.

Cardiverse: Harnessing LLMs for Novel Card Game Prototyping GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T13:48:54.886214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:48:54.886214Z digest=sha256:eaab39b326f456c4ff1f9e355354dcbef8ceb1d8e8bdda5a01da10a4f666ab5d

Observation 7e38bd61-6f22-4a25-8709-7c5227b6d909 · inbound

lmgame-Bench: How Good are LLMs at Playing Games? cites this paper.

lmgame-Bench: How Good are LLMs at Playing Games? GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:56.894546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:56.894546Z digest=sha256:88d5619c9a116c3b490c318352453725640577753d802f9cd52a2dbac5bf9f9e

Observation ed3a4d43-6e56-4e31-a840-ac9f9a8909b8 · inbound

Agents Require Metacognitive and Strategic Reasoning to Succeed in the Coming Labor Markets cites this paper.

Agents Require Metacognitive and Strategic Reasoning to Succeed in the Coming Labor Markets GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:23.124422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:23.124422Z digest=sha256:90071bbfda7b9a49d3b31baee80239fa0c435b99827c1a0283d731263ebe5601

Observation 4aeadccc-02a9-4415-961d-6356d4293d6b · inbound

A Practical Approach for Building Production-Grade Conversational Agents with Workflow Graphs cites this paper.

A Practical Approach for Building Production-Grade Conversational Agents with Workflow Graphs GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:11.853620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:59:11.853620Z digest=sha256:8b877181601ccdeca76054140c5f6848913f9bed185188b259e0b3ed51d961ac

Observation 0fb2d63d-bee2-4b95-827f-7fc8b15ebf6b · inbound

Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games cites this paper.

Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:02:16.594562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T12:01:42.681135Z digest=sha256:cb15d5c93b62490acc1a4dee1b0fc2544d35fc9a262f001f0415d1cf2a534cb3

Observation 7ce69c54-e46c-48e7-b285-c9b16b59878d · inbound

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey cites this paper.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 109

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.529936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.529936Z digest=sha256:1ec6057eb7e59904c76a7576dafca8424b4eb7c952909c611bd63010cd27b750

Observation dd1758e4-3553-4f38-9474-fc54f9e18953 · inbound

Tracing LLM Reasoning Processes with Strategic Games: A Framework for Planning, Revision, and Resource-Constrained Decision Making cites this paper.

Tracing LLM Reasoning Processes with Strategic Games: A Framework for Planning, Revision, and Resource-Constrained Decision Making GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:45.645947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:45.645947Z digest=sha256:88b0efeb4f79f1d2571561bebc60af7e6cb1064958d3c7492effc56d087bb36a

Observation 987d5ba4-0bac-4117-82cc-d29d004a9ab9 · inbound

Talking-to-Build: How LLM-Assisted Interface Shapes Player Performance and Experience in Minecraft cites this paper.

Talking-to-Build: How LLM-Assisted Interface Shapes Player Performance and Experience in Minecraft GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T13:46:22.108679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:46:22.108679Z digest=sha256:f41c6cf25c1413efc38ebc24ce24f4a1a932df5e2ffd344b7853c18c29bb55ed

Observation fd5e71eb-3fd5-4fc8-bee5-f08a2d330abe · inbound

Evaluation and Benchmarking of LLM Agents: A Survey cites this paper.

Evaluation and Benchmarking of LLM Agents: A Survey GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.535787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.535787Z digest=sha256:50f084af9f4b1eecaeeddd19994b4b55d7a12f1668432646ee251d7a796d4b5e

Observation bcaf3dd5-840a-4fd0-8c7c-6e80e316d631 · inbound

Game Reasoning Arena: A Framework and Benchmark for Assessing Reasoning Capabilities of Large Language Models via Game Play cites this paper.

Game Reasoning Arena: A Framework and Benchmark for Assessing Reasoning Capabilities of Large Language Models via Game Play GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T04:33:42.122192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:33:42.122192Z digest=sha256:12a7f061275d80701d77a4ef4f13daa5194b9a3aacc284e72e4329a94362a93e

Observation 884612bd-9a67-48d0-bb05-fbe7a0c7a256 · inbound

Democratizing Diplomacy: A Harness for Evaluating Any Large Language Model on Full-Press Diplomacy cites this paper.

Democratizing Diplomacy: A Harness for Evaluating Any Large Language Model on Full-Press Diplomacy GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T22:07:46.942904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:07:46.942904Z digest=sha256:07dcef5b6c8327c44f5d320f4b312b11087c88b198d2eb22553aa24ba88f226a

Observation df416bba-30db-46f5-8714-5221c907d163 · inbound

CHBench: A Cognitive Hierarchy Benchmark for Evaluating Strategic Reasoning Capability of LLMs cites this paper.

CHBench: A Cognitive Hierarchy Benchmark for Evaluating Strategic Reasoning Capability of LLMs GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T19:48:54.647303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:48:54.647303Z digest=sha256:9f88da854be550accce6e0c64971e408cdb736f0965a664219cdde04f9151661

Observation cb706f97-d4a7-4e9b-ac99-8cb0ac403234 · inbound

CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V cites this paper.

CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:41:01.825660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T17:59:21.887765Z digest=sha256:48bc60afa7d0c0511b79dbf47748fe1b3b1031fd7d92c53b94ee80bee03c2038

Observation c1dfd01b-706b-4a48-9568-3b76e9ba769c · inbound

DORA Explorer: Improving the Exploration Ability of LLMs Without Training cites this paper.

DORA Explorer: Improving the Exploration Ability of LLMs Without Training GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:31:30.739567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T06:27:36.922485Z digest=sha256:cba550a38bb7539cbf76f8382fcc0ad99ede7755c6497a2fa839555fd0c5cecc

Observation 27bd0da9-c555-463f-ac6a-4c3da179efe4 · inbound

Mechanism Plausibility in Generative Agent-Based Modeling cites this paper.

Mechanism Plausibility in Generative Agent-Based Modeling GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:27:50.727667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T19:24:53.816091Z digest=sha256:8f291ddf9ad527214259dd26c9c6c08b6e1f883e264bf630ec1f08c3d436457a

Observation d55229ce-b5e4-4106-acec-2f013a55ae0f · inbound

Mechanism Plausibility in Generative Agent-Based Modeling cites this paper.

Mechanism Plausibility in Generative Agent-Based Modeling GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:53:43.331988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T20:51:54.829821Z digest=sha256:d24ef19ed355418f6a4be49f521820a4e1fb83df13a06c05e40fc8093c44dc30

Observation 909b3e46-a1c0-4de6-ab6c-9dc215409fdb · inbound

Scale-Dependent Collective Adaptation in Self-Amending LLM Societies: A Cross-Family Study of Emergent Governance cites this paper.

Scale-Dependent Collective Adaptation in Self-Amending LLM Societies: A Cross-Family Study of Emergent Governance GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-19T22:32:49.757376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T22:29:17.058647Z digest=sha256:73b72b5d7dd28c3356ba6d575143944bf6e517a0d61972ba7aba6b0e6f4396c7

Observation 829a4d9d-dfbc-4e51-afff-29120182443e · inbound

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games cites this paper.

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:28:17.223104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T12:24:06.062957Z digest=sha256:a73a069b3918eabb69c9b2ff98ea6b29499045a6ed2c62fbdaf1c45ed15fe251

Observation c77d4480-722c-49e4-a929-3313c6fbfeb0 · inbound

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games cites this paper.

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:45:23.209799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T05:45:04.573722Z digest=sha256:2e2a657b79c311cfd467de186aade87e32cbe1d611ceaecc59da01ef7a69db9e

Observation aaa34980-a9de-4766-81ec-46c655554ffe · inbound

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play cites this paper.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:41:08.490213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:82fc8264e4bf5747d62ae59ea56d89a920e79a07c9b2a060815208efa31cfa6b

Observation 58c20efe-e896-42ff-bdf7-2572582b4d01 · inbound

GENSTRAT: Toward a Science of Strategic Reasoning in Large Language Models cites this paper.

GENSTRAT: Toward a Science of Strategic Reasoning in Large Language Models GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:45:20.661900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T04:41:35.532363Z digest=sha256:9654ccbe16f113f829d8d381370b16a12d25fd0c7e2125498160adbcfe52a26d

Observation d3c50d8d-6641-45cc-ac87-441fa5e8ab56 · inbound

MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs cites this paper.

MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:23:13.225042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T07:15:27.939886Z digest=sha256:19727640973d3d82e34253c1d8bb7cc73d15f47fdb9d6213bbf35cb0b51cda7c

Observation 1e0264e1-e215-4528-bb99-535cef897dc5 · inbound

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application cites this paper.

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:58:02.959131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T09:46:30.702256Z digest=sha256:b3f2e4bc9cb1ecb529b21cab2bc8b32af20c0d4908c960caa3677ad0864fb8c4

Observation 3beae700-2d4a-4407-b34c-757e5a6b5f3b · inbound

Enhancing Decision-Making with Large Language Models through Multi-Agent Fictitious Play cites this paper.

Enhancing Decision-Making with Large Language Models through Multi-Agent Fictitious Play GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:18.400250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T20:57:49.840546Z digest=sha256:05496e01b0f793e2c239c2fd009473154aea7d542250208f15c71781ec173e5e

Observation 91441aa4-6d18-4095-86d0-037d058e4b31 · inbound

Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games cites this paper.

Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:19:13.720602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T21:17:02.332687Z digest=sha256:36379397c02c38fe5087a4edfc09da36009653da3aa188e9ea396155574b7aec

Observation d5988a91-cb98-4ff4-9431-648c00f011ac · inbound

Age of LLM: A Strategic 1v1 Benchmark for Reasoning, Diplomacy and Reliability of Large Language Models under Fog of War cites this paper.

Age of LLM: A Strategic 1v1 Benchmark for Reasoning, Diplomacy and Reliability of Large Language Models under Fog of War GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:20:00.168465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T23:49:25.799611Z digest=sha256:6e71f735f581a5e359b0cbad25bfd1ed7094e269366f73e8044e86594659d2fb

Observation 41216d83-c78f-4341-8a30-c316306f1c4b · inbound

SidConArena: An Environment Evaluating Agents in Open-Ended,Positive-Sum Bargaining Game cites this paper.

SidConArena: An Environment Evaluating Agents in Open-Ended,Positive-Sum Bargaining Game GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T18:35:58.017825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-29T01:56:22.856352Z digest=sha256:a2a054b9b05b2279f1d559221d942b685b681215fc055be2a82d417a7ec2e9c6

Observation 8b5f053d-9714-417f-90e5-f34f088869e6 · inbound

Spatial Reasoning in LLM Game Agents: Impact of Causal Context and Multi-Step Planning cites this paper.

Spatial Reasoning in LLM Game Agents: Impact of Causal Context and Multi-Step Planning GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T10:58:16.650048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:58:16.650048Z digest=sha256:61f21fc5ed6672a0025fa615a77bfe0bae863dce048e2232ba1d048b1b7bbea8