Pith. sign in

Paper Citation Record · LEDGER

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models

As of 17 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2506.10264.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10264 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:34:40.588180Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:46:04.569599Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T04:46:06.449061Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact1
  • verified fuzzy15
  • unresolved19
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5ef329c2-faa8-4c3c-91f6-9bdce7c3567b · outbound

This paper cites A Survey of Large Language Models.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models A Survey of Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:36.717806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:36.717806Z digest=sha256:2f3d08b83e0567140e8c9d0419f253a9c115ca2e0b6b62bcc43a561247ec8f90

Observation 5a0ed04e-5fa6-42f1-97a3-aafaf91785e6 · outbound

This paper cites GPT-4 Technical Report.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:36.819465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:36.819465Z digest=sha256:8f51926ce20ae3ce024a3618b9520ca34b173605fc82a948264b590abd137ee7

Observation 8fcf919a-7c90-4337-a613-36fc49f7c809 · outbound

This paper cites DeepSeek-V3 Technical Report.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models DeepSeek-V3 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:36.962474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:36.962474Z digest=sha256:0dd99aad8736f18e2e524cc3d9a64db2bb04c2befeefede5e722c58b44f7d910

Observation c5ffafba-0b18-46cb-9c61-3384db503749 · outbound

This paper cites From System 1 to System 2: A Survey of Reasoning Large Language Models.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models From System 1 to System 2: A Survey of Reasoning Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:37.111179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:37.111179Z digest=sha256:759f53bd26c98c8fa7deae09e432d9fa4512fff8362c5aa26f35cc2de8260dba

Observation 1fc1952e-d8ff-4904-a2a9-b9b9d374da4e · outbound

This paper cites Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:37.190945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:37.190945Z digest=sha256:1006b09ab3bc8d64a932157832e0e2c604e7fbbfd37ce19332e144dd1aa1dd4c

Observation f1a9a322-c273-4045-bfc8-e53b9d002f2b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:37.280407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:37.280407Z digest=sha256:f54606de32abe4a416be7565eda8fd29b0016a90b46b2f41a1bedd641e241c3a

Observation d739b67a-0cd2-485d-a83c-f8f98f421293 · outbound

This paper cites Large Language Models for Mathematical Reasoning: Progresses and Challenges.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Large Language Models for Mathematical Reasoning: Progresses and Challenges

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:37.342152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:37.342152Z digest=sha256:0640e7d0c62d74a1f122a943d38e8391a53cb2a7e79d165be6cd1c76c80a53a6

Observation e98680d2-4cf7-4ec3-b9b9-f7237ac17f62 · outbound

This paper cites Symbol-llm: Leverage language models for symbolic system in visual human activity reasoning,.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Symbol-llm: Leverage language models for symbolic system in visual human activity reasoning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:44.113562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T04:34:37.475335Z digest=sha256:263cce6face32a1057353c55523f398cefa95c28f69c10dc97d0b7db9314c59c

Observation a1320bd8-410b-40cf-a8b0-d6f97d3b1b38 · outbound

This paper cites Large language models as com- monsense knowledge for large-scale task planning,.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Large language models as com- monsense knowledge for large-scale task planning,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:44.033875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T04:34:37.563033Z digest=sha256:44a3b8ec89d4027137570fe663b8d2e0821979479e6f24840644dd08ffda3872

Observation 92f67507-c599-4a74-9854-40264fa754f0 · outbound

This paper cites LLM as a Mastermind: A Survey of Strategic Reasoning with Large Language Models.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models LLM as a Mastermind: A Survey of Strategic Reasoning with Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:37.674228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:37.674228Z digest=sha256:6dcd406dcc9e86837a4628cbe20e0b9463c081b64edded6098fcefdd35922703

Observation db5f9711-92a8-4130-83e3-98e13121d246 · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:37.773251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:37.773251Z digest=sha256:a7beb04a78b32964cd76057e7df5b9f57ecbd4251a2ada6e98f61fbc6831785c

Observation 9c92b787-ddad-496b-83e6-a9517e2cd855 · outbound

This paper cites AvalonBench: Evaluating LLMs Playing the Game of Avalon.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:37.906915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:37.906915Z digest=sha256:c266f1c416370f9b42dceb4cbe1d1f0e0aee5094903cf73f1babc8b78272930e

Observation c76d2130-52e1-4adf-bd7e-30f47eee7cde · outbound

This paper cites Large language models play starcraft ii: Benchmarks and a chain of summarization approach,.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Large language models play starcraft ii: Benchmarks and a chain of summarization approach,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:43.936074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T04:34:37.993162Z digest=sha256:687ba608d0c123a94a1a4982a3220ba066e72053390598b4b40522f5d0968562

Observation 548dc6b5-71ea-4142-bde3-ed7c49c19f2c · outbound

This paper cites Hierarchical Expert Prompt for Large-Language-Model: An Approach Defeat Elite AI in TextStarCraft II for the First Time.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Hierarchical Expert Prompt for Large-Language-Model: An Approach Defeat Elite AI in TextStarCraft II for the First Time

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:38.110383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:38.110383Z digest=sha256:06516a845dc740ec89d76395b04e9194bfa43a857e474eb17a66b4fab39b21cd

Observation d37a19b2-dfde-4166-84ee-8fc49b616ab4 · outbound

This paper cites GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:38.244643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:38.244643Z digest=sha256:a4b78b6b9d84d9239740ecbd3e04061f826f01166ccf23539605543e974ce4dc

Observation 4cef86c4-0035-4502-a347-cb4f3d6ce03d · outbound

This paper cites MAgIC: Investigation of Large Language Model Powered Multi-Agent in Cognition, Adaptability, Rationality and Collaboration.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models MAgIC: Investigation of Large Language Model Powered Multi-Agent in Cognition, Adaptability, Rationality and Collaboration

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:38.370753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:38.370753Z digest=sha256:c3811e66ad81c136ef5ae3b722df0b087f1e6e27e2d352468af6bf0875171c23

Observation 74aa5be8-8b80-47cf-a647-7f0c6a55d467 · outbound

This paper cites SmartPlay: A Benchmark for LLMs as Intelligent Agents.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:38.508070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:38.508070Z digest=sha256:101d90872b86e10cbf46c99a99db16d348a1a529204af96fd3734e612019d636

Observation 7a00ec7a-1e20-44a4-8a41-c6985d370f85 · outbound

This paper cites MinePlanner: A Benchmark for Long-Horizon Planning in Large Minecraft Worlds.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models MinePlanner: A Benchmark for Long-Horizon Planning in Large Minecraft Worlds

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:34:40.910207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T04:34:38.630962Z digest=sha256:c6a47940437c26b9aab2aced19bf64803d7ae83711dee99109b7cf3708e209d7

Observation 5cc67543-6e4f-4d34-ba79-1051858155a5 · outbound

This paper cites Benchmarking Agentic Workflow Generation.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Benchmarking Agentic Workflow Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:38.736915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:38.736915Z digest=sha256:0dc9da81f996fefadc7d10dea1edf6485277eb582a23c36bf00038af9618e7fa

Observation f60b63af-7eef-4021-a1fa-69ac4cacbf93 · outbound

This paper cites Opendeception: Bench- marking and investigating ai deceptive behaviors via open-ended interaction simulation,.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Opendeception: Bench- marking and investigating ai deceptive behaviors via open-ended interaction simulation,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:38.830551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:38.830551Z digest=sha256:a2a974d863cf78b5618096d2e8f108e777ef7e7b234522d53458590298b6fba5

Observation 9eef73d2-8dc6-4004-bb3c-9b50647cc91e · outbound

This paper cites Intelligent decision making technol- ogy and challenge of wargame,.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Intelligent decision making technol- ogy and challenge of wargame,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:43.767158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T04:34:39.036516Z digest=sha256:bb32455ba7720d01d53739fb573e70a32ffff215beda2b5fa7919630188661c7

Observation b0313ab1-a818-47fb-a7c0-fcbc6dfdbde4 · outbound

This paper cites an unresolved cited work.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:34:43.640547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T04:34:39.155687Z digest=sha256:19ea0f8f4990aced4508337dd77f1cd5d2a19b62678877db166cbeba6459e3d6

Observation b80d093f-01fe-4f2f-a1b2-9759cf70a94b · outbound

This paper cites A five- dimensional framework for authentic assessment,.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models A five- dimensional framework for authentic assessment,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:43.471678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T04:34:39.247694Z digest=sha256:0972fe465b606737f1fa92f402edf4d96f5e7eda5e26bbd18d611dfe8733b37d

Observation 95650192-16e4-4a14-86b5-250e7802eeca · outbound

This paper cites Appropriate criteria: Key to effective rubrics,.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Appropriate criteria: Key to effective rubrics,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:43.333857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T04:34:39.308215Z digest=sha256:53b041f140e735d59cbf37be97fc962ae620baeb278f9dfff577d82e1b8b2745

Observation 6915e39c-f832-4f67-9b9e-2119d88fea25 · outbound

This paper cites an unresolved cited work.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:34:43.206383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T04:34:39.468038Z digest=sha256:7c5c3a015b41d831fa5a266adc698a817f32bb9df4b75cf339e00261bcedd6bb

Observation f1778ec8-9ea0-43c4-aad4-b6c2a46a3eb5 · outbound

This paper cites A review of rubric use in higher education,.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models A review of rubric use in higher education,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:43.051297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T04:34:39.639505Z digest=sha256:686995b75f10dde92ab88c3e6505f0bc081f458df1d17776f8bb458c11606dcb

Observation 2d8f7a0b-7ce3-4ead-8e43-9d0e4c299ef8 · outbound

This paper cites Wiggins,Educative Assessment: Designing Assessments to Inform and Improve Student Performance.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Wiggins,Educative Assessment: Designing Assessments to Inform and Improve Student Performance

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:42.744037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T04:34:39.806281Z digest=sha256:7a04f5fd0f93155a4d11938addf204aab5099bbb3a25a27790640c97a11d70e7

Observation 0bd4715e-b19e-4b47-9f0b-8a57cf5c177b · outbound

This paper cites A revision of bloom’s taxonomy: An overview,.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models A revision of bloom’s taxonomy: An overview,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:42.504103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T04:34:39.897007Z digest=sha256:185e44088bfead4ca588f7d1bf3967ea51b28fb0180770d8afd131f81edd78b7

Observation 880d3f6e-76e3-425b-a96f-5da3de1616c7 · outbound

This paper cites an unresolved cited work.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:34:42.185146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T04:34:40.006617Z digest=sha256:9fe39f7cbf0a531054244fa872dc5d58d8bfe9efe0c6a92de68cfecf050e24b9

Observation 3f2c76a7-d340-4289-b2d8-643704ab78a3 · outbound

This paper cites Improving instruction and assessment via bloom’s taxonomy and descriptive rubrics,.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Improving instruction and assessment via bloom’s taxonomy and descriptive rubrics,

Reference 30

Resolution
malformed identifier
no resolver link, observed 2026-08-07T04:34:40.143962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:40.143962Z digest=sha256:3b4116f035f965c496b8e64b9f261816ea1fc5442af4df5b66f0407a29b85d4f

Observation ccde419f-35c4-4a70-a439-1e6ccf439aa2 · outbound

This paper cites Llm-rubric: A multidimensional, calibrated approach to auto- mated evaluation of natural language texts,.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Llm-rubric: A multidimensional, calibrated approach to auto- mated evaluation of natural language texts,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:41.962754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T04:34:40.231939Z digest=sha256:f2b8cdf507fec637242a0a6143b4a1fba6d0a17ceec60c1d41d8c1a0c5452c32

Observation 29726c48-a90a-492d-985b-6a95f506daa0 · outbound

This paper cites Gptscore: Evaluate as you desire,.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Gptscore: Evaluate as you desire,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:41.753901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T04:34:40.298283Z digest=sha256:e41c7bda038fca0bfe232d77b4a6e416460a44978fb5f6c06d78c8dac5ddf270

Observation 978657f9-f8ad-441e-87cd-bd2e3649bc8d · outbound

This paper cites G-eval: Nlg evaluation using gpt-4 with better human alignment,.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models G-eval: Nlg evaluation using gpt-4 with better human alignment,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:41.573731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T04:34:40.374306Z digest=sha256:5077b6544a38238e75967434ce51b91e05449b591da7cd600f5e242e458a4a90

Observation fa1583f1-75a5-41f0-b858-5fee8b9242b1 · outbound

This paper cites Large language models are not fair evaluators,.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Large language models are not fair evaluators,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:41.400228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T04:34:40.439948Z digest=sha256:372a1be9162e6e281400d9fe5689f993f90597cf11b0938008b864d3b01cd98f

Observation d32b7a67-3b61-43d7-9c60-5a1138982931 · outbound

This paper cites Realtoxicityprompts: Evaluating neural toxic degeneration in language models,.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Realtoxicityprompts: Evaluating neural toxic degeneration in language models,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:41.222654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T04:34:40.510465Z digest=sha256:1cc6e3460c7cd7721d917425c92d10c467bba7add3b7e9044f935aa23c63def3

Observation 7615c8bc-a834-4911-9077-33f9fefb4a88 · outbound

This paper cites Why we need new evaluation metrics for nlg,.

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models Why we need new evaluation metrics for nlg,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:41.082443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T04:34:40.588180Z digest=sha256:f4aeff7302c59c4cc65e1c689c593e96ed713ebd2c0f81148809a45f1659ef01

Pith citing papers

Observation 7f9b05cd-6ed6-4e5e-b6d7-4f31ed705399 · inbound

Multi-level Value Alignment in Agentic AI Systems: Survey and Perspectives cites this paper.

Multi-level Value Alignment in Agentic AI Systems: Survey and Perspectives WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models

Reference 148

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:46:06.453067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T04:46:04.569599Z digest=sha256:84601a08fe24f7ab9c918801555ed8fe6feb51cd3ebc92f69a983a43794be7f2