Pith. sign in

Paper Citation Record · LEDGER

GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2402.12348.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.12348 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T19:37:15.933653Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 45209f9c-2cd0-447c-90b2-50cebb81d5f7 · inbound

Estimating LLM Uncertainty with Evidence cites this paper.

Estimating LLM Uncertainty with Evidence GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T19:37:15.933653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:37:15.933653Z digest=sha256:9d9d8288a78e3392c6d77b82ca0e41dc370da49bb1e4f9b716230ea72e9f61a6

Observation d822c53b-83bc-4bc7-897f-ac9d279a87f7 · inbound

Verbalized Bayesian Persuasion cites this paper.

Verbalized Bayesian Persuasion GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T14:57:41.362935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:57:41.362935Z digest=sha256:21ce5c0c94c46d0eae627f0b0eec62b332b71590beb6803695416fe9c9317d4e

Observation 117a72f0-ac6a-4a25-9b3a-786bc51021ee · inbound

Grammars of Formal Uncertainty: When to Trust LLMs in Automated Reasoning Tasks cites this paper.

Grammars of Formal Uncertainty: When to Trust LLMs in Automated Reasoning Tasks GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:06:31.207570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:06:31.207570Z digest=sha256:4e28de9b0f1a67b1864e5dba415cedd51a97d7cbe7239c1295192852d7aa2789

Observation ccdb28e0-7a0d-4687-90bb-27dd9651a7fe · inbound

VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments cites this paper.

VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:57:16.225848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T11:57:08.314088Z digest=sha256:b7628dfe2d2da6ae1dbc1e99f36e26678e4a4c7856fa5cd5a2747feaf8aa7c8a

Observation 10149eab-6e4a-41bf-af76-6a916325d0b8 · inbound

Shapley-Coop: Credit Assignment for Emergent Cooperation in Self-Interested LLM Agents cites this paper.

Shapley-Coop: Credit Assignment for Emergent Cooperation in Self-Interested LLM Agents GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:48.118071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:48.118071Z digest=sha256:d56462348130ab17e0918111cf428f2f553f23d304c3cd3cda21b05b19f11375

Observation d37a19b2-dfde-4166-84ee-8fc49b616ab4 · inbound

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models cites this paper.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:38.244643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:38.244643Z digest=sha256:72e12c990f450d217c2b438ccbe103c676879513d7c13fea61c271dade4ab81b

Observation 3ecc2b31-fb22-4aec-984e-ffa1694e5e8a · inbound

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey cites this paper.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations

Reference 114

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.537892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.537892Z digest=sha256:f6b75bb58b90484febcd5f2fbe585adadebea2b7f98eb5667dff041b14449ffe

Observation 2794b4f1-eeb1-45e2-bcbb-2ac90dc8244e · inbound

Tracing LLM Reasoning Processes with Strategic Games: A Framework for Planning, Revision, and Resource-Constrained Decision Making cites this paper.

Tracing LLM Reasoning Processes with Strategic Games: A Framework for Planning, Revision, and Resource-Constrained Decision Making GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:45.860130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:45.860130Z digest=sha256:69b605a3d7419ef02c6845fd8da50154a817a26098aba56023a9531c25ca0f94

Observation 88deb9cc-21ec-4eea-a3d7-eeea63417592 · inbound

The Effect of State Representation on LLM Agent Behavior in Dynamic Routing Games cites this paper.

The Effect of State Representation on LLM Agent Behavior in Dynamic Routing Games GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:00:11.342661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:00:11.342661Z digest=sha256:5586c1d3e8bd1ca86f906c9d49f0d37cb32d15a3e7e00363e300a4dd558f22d4

Observation ade39c0f-f51e-4c82-8f16-222cdfb2ae79 · inbound

DipSVD: Dual-importance Protected SVD for Efficient LLM Compression cites this paper.

DipSVD: Dual-importance Protected SVD for Efficient LLM Compression GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:58:09.308919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:58:09.308919Z digest=sha256:5d3f61ab19302fb5dd6563af0afc3ebc61c2e8bd9a35f8adb19befddc7f4be89

Observation 025f81de-cd78-4010-a445-bc1cccddaa54 · inbound

How Large Language Models play humans in online conversations: a simulated study of the 2016 US politics on Reddit cites this paper.

How Large Language Models play humans in online conversations: a simulated study of the 2016 US politics on Reddit GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:18.435764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:18.435764Z digest=sha256:4c8c92b5a0053627f3d7a23950bc5760d65089a0ce578cb86e660e0115e8562f

Observation 611fddd2-c4d0-4182-8f8c-7c17f856217b · inbound

GenAI-based Multi-Agent Reinforcement Learning towards Distributed Agent Intelligence: A Generative-RL Agent Perspective cites this paper.

GenAI-based Multi-Agent Reinforcement Learning towards Distributed Agent Intelligence: A Generative-RL Agent Perspective GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:57:08.266513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:57:08.266513Z digest=sha256:bc4a9bea2041e5a10cee22dbf930b2b9780a551c8e2d25f2c95a659522d17319

Observation 30f59d22-4284-4e1c-a3ca-b9322be083f2 · inbound

Democratizing Diplomacy: A Harness for Evaluating Any Large Language Model on Full-Press Diplomacy cites this paper.

Democratizing Diplomacy: A Harness for Evaluating Any Large Language Model on Full-Press Diplomacy GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T22:07:47.740830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:07:47.740830Z digest=sha256:c97b23d428a8fded5a522dca40870e2500664b7b29047525a8ebfe3a9972d81c

Observation 12bcbc80-d732-47ad-be20-fc098b3ac09b · inbound

Strategic Heterogeneous Multi-Agent Architecture for Cost-Effective Code Vulnerability Detection cites this paper.

Strategic Heterogeneous Multi-Agent Architecture for Cost-Effective Code Vulnerability Detection GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:28.941171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T22:06:47.552043Z digest=sha256:618b01c2ba5bcaf7ea64d5dabffd8eb9d36c8fcc3af8da0730e7b921a7543f68

Observation 9f2990fc-a0f6-4512-8b83-4b798bd08e2d · inbound

Common-agency Games for Multi-Objective Test-Time Alignment cites this paper.

Common-agency Games for Multi-Objective Test-Time Alignment GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:15:06.376948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-15T06:14:53.685486Z digest=sha256:e494c01b3a9b0098ade438798eb020fac70f8c859aedb7ef45ddeb621613b516

Observation 2f7e3d0e-74a3-4145-8dd5-c89289876e4b · inbound

DemoEvolve: Overcoming Sparse Feedback in Agentic Harness Evolution with Demonstrations cites this paper.

DemoEvolve: Overcoming Sparse Feedback in Agentic Harness Evolution with Demonstrations GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:14:40.656457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T13:12:46.927103Z digest=sha256:cfd663e10996363189aa67295743855d29c38688763e7830ac96ede7d4735042

Observation aa4058ae-4165-46e4-a750-1b83aee8f181 · inbound

Robots Need More than VLA and World Models cites this paper.

Robots Need More than VLA and World Models GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations

Reference 143

Resolution
verified exact
arxiv_id, observed 2026-06-28T01:11:28.888729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T01:01:33.530167Z digest=sha256:65fb430564ed6d9e750aac672a51da71514aadc414ce45d2cda94e84c4d67ecc

Observation d50edb7b-404d-45ca-b0c9-ac88b352d22b · inbound

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application cites this paper.

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations

Reference 120

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:58:02.988053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T09:46:30.702256Z digest=sha256:0e87b393a9a6b70647e2617411aa7238e179c83eda1eff99853d1db120bc4db7

Observation 1c59a98e-472a-4f27-a2aa-27ca3b39248e · inbound

Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games cites this paper.

Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:19:13.715084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T21:17:02.332687Z digest=sha256:68b4c096ab50ff4fd8de0d694fd1a0260b58f7608e3340b1d89b732de5d9ed04

Observation b292097a-1f13-4384-a22d-90153f7ca041 · inbound

BehaviorBench: Benchmarking Foundation Models for Behavioral Science Tasks cites this paper.

BehaviorBench: Benchmarking Foundation Models for Behavioral Science Tasks GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-26T00:38:42.745298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T00:38:14.922388Z digest=sha256:fe54f688c0fe9f1c6a37da66cd5792c7ec569f009c0dd786c3dae5b259ceb8dc

Observation 45aeb99a-2a1f-4e2d-8263-e1c954e94bb1 · inbound

Age of LLM: A Strategic 1v1 Benchmark for Reasoning, Diplomacy and Reliability of Large Language Models under Fog of War cites this paper.

Age of LLM: A Strategic 1v1 Benchmark for Reasoning, Diplomacy and Reliability of Large Language Models under Fog of War GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:20:00.182045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T23:49:25.799611Z digest=sha256:4a116503fd9baab4e55d9abbac11aea5e6fb43af4689fb9788e80f00291bba78

Observation 71dc7915-1b89-4660-a538-38354c769283 · inbound

Towards Reliable and Robust LLM Planning: Symbolic Feedback-Driven Iterative Self-Refinement Framework cites this paper.

Towards Reliable and Robust LLM Planning: Symbolic Feedback-Driven Iterative Self-Refinement Framework GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:13:52.970237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T04:53:17.854375Z digest=sha256:05fd007955c674123d27c0774dda72159e9aaf0d896d7f3e83bef901e7ba06b0

Observation d655b8b7-d299-4544-a174-9373c020718b · inbound

Scaffolding the Strategist: Architecture-Dependent Reasoning Interventions in Hotelling Spatial Markets cites this paper.

Scaffolding the Strategist: Architecture-Dependent Reasoning Interventions in Hotelling Spatial Markets GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T16:35:42.294582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:35:42.294582Z digest=sha256:2bdcf8d2b842f95f5e1d49c0f0571865086e095d56cd0eef9d296020d7890adc

Observation 99ee0253-bde2-48e4-acba-bb511d2da326 · inbound

Strategy, Not Payoffs: A Behavioural Embedding of Normal-Form Games cites this paper.

Strategy, Not Payoffs: A Behavioural Embedding of Normal-Form Games GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T06:09:13.779994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:09:13.779994Z digest=sha256:ed64dc92628f0fca298f548d673e4235d1cea53ad1f0cf7648cda4cdedfcb8a1