Pith. sign in

Paper Citation Record · LEDGER

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play

As of 17 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 0 inbound Pith citation observations for arXiv:2605.22238.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.22238 v1

Coverage vector

measured 11 of 11 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-22T05:37:44.982932Z

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

11 of 11 outbound references displayed

  • verified exact9
  • verified fuzzy2
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6b3c7bfb-814c-4cef-86ca-bd27026cc3e8 · outbound

This paper cites Science , volume =.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play Science , volume =

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T05:56:09.458656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:ed68401277cced3d8f834ee740f9cc13f6b3b4dd20a929f30d50292709bec7cf

Observation b352fd9b-bde4-48db-9097-77c2079dc5c5 · outbound

This paper cites 2024 , url =.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play 2024 , url =

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T05:56:09.462337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:f1c0659f8cbdd868c9fc10ac4d3c5ab4c512fcede4eca153a331ccc7cf8d8dea

Observation 708a3513-9f52-4af9-83c9-a4c96515b264 · outbound

This paper cites Human-level play in the game of diplomacy by combining language models with strategic reasoning.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play Human-level play in the game of diplomacy by combining language models with strategic reasoning

Reference 10

Resolution
verified exact
doi, observed 2026-05-22T05:41:07.034526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:41d528c43ac6bf5710fcf49ac3a4c6422f6972c4176cadc6a5c5f0efd797e090

Observation aaa34980-a9de-4766-81ec-46c655554ffe · outbound

This paper cites GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:41:08.490213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:d3aef139135eb0e62fb65dd43ceb5d132db3d2c6f1d24026d3e74bbfe8f68430

Observation 2f52693b-87b6-43c8-ba42-e929196dd843 · outbound

This paper cites Strategic Reasoning with Language Models.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play Strategic Reasoning with Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:41:08.486282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:8e4c1757aa28a9f2b19e098c9d9306720076e5b61b0d36611a1d0effaf74ea4d

Observation 1fd2599b-1768-4a2f-878f-c60e6f7610e4 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play Measuring Massive Multitask Language Understanding

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:41:08.482224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:71757c33bdeba67f50321bf7013503680ae343f8ec910fbc0b80e55544b68e4b

Observation 4248d872-252a-4963-bc25-47c423a5bb63 · outbound

This paper cites Holistic Evaluation of Language Models.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play Holistic Evaluation of Language Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:41:08.473414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:4a9cbbdfa12104dc9e3974e1fee7a2635ba97bc2e5d455bd53e3c554136c3bdb

Observation 9c3e1682-da6d-4e13-831d-89dd2c7aa389 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play AgentBench: Evaluating LLMs as Agents

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:41:08.469034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:260e3d6ec7cc015f6b26bd27619e3994f809285404a525624d7c4566ce2b516f

Observation a385a6bf-19c5-4b7f-99fd-95d7e89096fc · outbound

This paper cites GAIA: a benchmark for General AI Assistants.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play GAIA: a benchmark for General AI Assistants

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:41:08.463536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:bbdc51c75ee4e9b7c00ebae35317c63596319ea3b62cefc71e2260ce3307fe43

Observation 850250d3-cf0a-41a9-ae37-e18e2fbeca8c · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:41:08.478010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:ec1243c3c125cddbb092f5a0f9e2e9faffe84d568b8dbd294289d3346825ba75

Observation 19e3369b-35d0-40d3-872e-3d0c9fd0433b · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:41:08.494410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-22T05:37:44.982932Z digest=sha256:bf75ef901d9cd9d17fd29aaa2a67b18ab0ccb115c5a8a1bf0da396c226118e39

Pith citing papers

No inbound Pith citation observations are available.