Pith. sign in

Paper Citation Record · LEDGER

Evaluating Large Language Models with Grid-Based Game Competitions: An Extensible LLM Benchmark and Leaderboard

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2407.07796.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.07796 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T22:50:32.995433Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T13:41:36.684597Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation dcfa133e-3d45-45ae-a553-635e64b4425c · inbound

Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers cites this paper.

Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers Evaluating Large Language Models with Grid-Based Game Competitions: An Extensible LLM Benchmark and Leaderboard

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T22:50:32.995433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:50:32.995433Z digest=sha256:9f780daa223f96bb78a6e5b93b0839e797e01725514a7910bef7083b84d40ef0

Observation bd94c50a-1a6f-407c-878b-f44e1c7a8789 · inbound

MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation cites this paper.

MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation Evaluating Large Language Models with Grid-Based Game Competitions: An Extensible LLM Benchmark and Leaderboard

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:41:36.687125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T13:37:49.475416Z digest=sha256:6be5de32142aead0193c69fcb7b63526f69a29047c44e12d21c737a806f5371b

Observation e773708e-4e7f-4d08-95e7-99412259d6b0 · inbound

MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning cites this paper.

MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning Evaluating Large Language Models with Grid-Based Game Competitions: An Extensible LLM Benchmark and Leaderboard

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:58:36.058002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:58:36.058002Z digest=sha256:7a8e50d38f9d79bd8216fb2c9dfc54f1d577c3903686e65cae63b2337c3d9b37

Observation 526200cd-e14d-497e-bfc1-7226018d7aa2 · inbound

Game Reasoning Arena: A Framework and Benchmark for Assessing Reasoning Capabilities of Large Language Models via Game Play cites this paper.

Game Reasoning Arena: A Framework and Benchmark for Assessing Reasoning Capabilities of Large Language Models via Game Play Evaluating Large Language Models with Grid-Based Game Competitions: An Extensible LLM Benchmark and Leaderboard

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T04:33:42.811925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:33:42.811925Z digest=sha256:52af06e5a9511464e0b0659c734f40ffbaa0866122f3ace8c9898501a686d4a9

Observation 347ec6d5-84a2-47d0-b02c-4a838aa54c0a · inbound

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play cites this paper.

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play Evaluating Large Language Models with Grid-Based Game Competitions: An Extensible LLM Benchmark and Leaderboard

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T12:36:12.257522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:36:12.257522Z digest=sha256:feb4748333ebfd835537b9c5f0c1b650f299fdbedfb07d66eb77edfd5dc4c540