Pith. sign in

Paper Citation Record · LEDGER

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks

As of 18 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 1 inbound Pith citation observation for arXiv:2508.10428.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.10428 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:30:30.167398Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T21:02:26.081166Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T00:39:17.301001Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact2
  • verified fuzzy4
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 402af738-96bf-48ea-a77c-3c528f4d0aa5 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:28.708904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:28.708904Z digest=sha256:0842a0251ac3836bc9d90f3f96c345155daf0fcc39ed9be33fcb10aab776222b

Observation fcfd4c6f-7986-4566-bfbd-09aa7febcc19 · outbound

This paper cites write newline.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:28.792837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:28.792837Z digest=sha256:8f3eea2480d731b09e3f865d24876718355734727fdcd56a7d10e3c55c25e77c

Observation 3ab399f5-2654-49db-b347-e4bfaeaa3917 · outbound

This paper cites GPT-4 Technical Report.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:28.941922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:28.941922Z digest=sha256:4d046bc638f8d353174eebba6062e406bfdbd58e7bfbcf4f5460ecd8f6c2e6a5

Observation 6b936859-190d-4248-bc2a-732f483af256 · outbound

This paper cites K.; and Johnson, B.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks K.; and Johnson, B

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:30:30.442231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T20:30:29.110846Z digest=sha256:2b0276e5ce170f69cb9f76f1db6f0f2bfa94ff305fbb1b9ab83bdc509dfc8d1f

Observation d7755b31-5bda-4475-bf0b-d9d04ef19bcf · outbound

This paper cites DeepSeek-V3 Technical Report.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks DeepSeek-V3 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:29.216038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:29.216038Z digest=sha256:3ea94457d5fc87274c43d735cfbf76dafd1e58f95f3d52354912e6dc56c4dfb7

Observation a1b8d12e-1e9a-43ae-abf7-463dafdf8255 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:29.352392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:29.352392Z digest=sha256:0296e3b45a2c44e5d2e6dd17971261aab8bca49b3c43e27f4c9f65784e586cad

Observation a3304b1e-6bc5-4952-b5d3-00963259cf41 · outbound

This paper cites L.; Yao, S.; Chen, Y.; Shen, P.; Yu, H.; Zhang, H.; Zhang, X.; Dong, Y.; et al.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks L.; Yao, S.; Chen, Y.; Shen, P.; Yu, H.; Zhang, H.; Zhang, X.; Dong, Y.; et al

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:30:30.433855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T20:30:29.424757Z digest=sha256:0ff3fde7d737266fe9e46c95b1fd6810656c5b0fa54bb84c0092b4fadaaafb6c

Observation 36a70c36-a638-4ce1-a441-ead97bbd793d · outbound

This paper cites LLM-PySC2: Starcraft II learning environment for Large Language Models.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks LLM-PySC2: Starcraft II learning environment for Large Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-05T20:30:30.321803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T20:30:29.516705Z digest=sha256:571358fcdd60f4cffb19b1f8778c08a0486382045053772d714769f6914ba06e

Observation 505180f6-239f-4f55-a4c2-653ea13f0b53 · outbound

This paper cites AvalonBench: Evaluating LLMs Playing the Game of Avalon.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:29.614571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:29.614571Z digest=sha256:ab3d47e0677c7654dcecbeb1e094780c2f243c4bc429cced749ae5e4b50e72a0

Observation 8f376d00-af2a-400e-bd75-d07db199621e · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks AgentBench: Evaluating LLMs as Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:29.700721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:29.700721Z digest=sha256:7809bcdc6aa3e8d1b7ccfe64fd41307d9d66b0ecbb457788dd6b710a45f0171f

Observation fde970f7-9348-40e1-9558-6fcb47d8a8ff · outbound

This paper cites an unresolved cited work.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:30:30.425166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T20:30:29.807997Z digest=sha256:f14eaf27f13398502d0f2f5eef8f4b52d292f0519fa4422d615fc372852448ea

Observation f87afcd7-dcb0-4cce-b4f7-3ef77efe648e · outbound

This paper cites an unresolved cited work.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:30:30.415893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T20:30:29.895387Z digest=sha256:39e369b41b0d7511496a1c2de216e8deb2d3f43d409dac735ab9ad5e198adf80

Observation b2f2ef39-5c7f-4313-a5a7-9521dd196d9d · outbound

This paper cites an unresolved cited work.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:30:30.407634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T20:30:29.990030Z digest=sha256:c37e3dac26dd2a6d084de1b8379a53659fb8bc40804af4a4ad3a207be6aa2d23

Observation 7dc2bd4a-cc3e-44b4-ab76-a4ef4532e340 · outbound

This paper cites S.; Farquhar, G.; Foerster, J.; and Whiteson, S.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks S.; Farquhar, G.; Foerster, J.; and Whiteson, S

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.090360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.090360Z digest=sha256:65b3d401804f72dd97aedc761e726747ce85145600371bf71c17ba33efdf1929

Observation 0d1c389c-55fc-44b2-a709-f66e4b0f7af7 · outbound

This paper cites TurtleBench: A Visual Programming Benchmark in Turtle Geometry.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks TurtleBench: A Visual Programming Benchmark in Turtle Geometry

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.112631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.112631Z digest=sha256:1f8b7952e5d745a42e1f8af66c29fe5922f93f427197d2bcafdc8813c67bff56

Observation 0fb317f2-1174-470a-a596-89e4fb743e01 · outbound

This paper cites The StarCraft Multi-Agent Challenge.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks The StarCraft Multi-Agent Challenge

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.115627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.115627Z digest=sha256:12d8e533b7445a1da0165d07447377152e481dd19e7d5e464fa5927d2e88571e

Observation f48fbc19-1b65-4b94-acd6-9684496eeb38 · outbound

This paper cites Exploring and Improving the Spatial Reasoning Abilities of Large Language Models.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Exploring and Improving the Spatial Reasoning Abilities of Large Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-05T20:30:30.273375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T20:30:30.118412Z digest=sha256:d3a96b3c9110f0e8a156e42624606e354c52ef160635e63019e79b01916b4d92

Observation 92afc32c-3678-4e37-b21d-cb174bd43a38 · outbound

This paper cites Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.121068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.121068Z digest=sha256:8b31772e54b3fbf932ce2c47d9fc1a8481c0a80d5c2f0fd2bce25068e7023e4f

Observation 7390fba7-1b8b-42a2-843a-1b9e285ac15b · outbound

This paper cites an unresolved cited work.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.123990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.123990Z digest=sha256:64642d0d3be8f0375a2be2822157d825a43c637aec93b9f0420dcd72bd3e0c07

Observation e5d87f50-03aa-4644-a248-302c63543b03 · outbound

This paper cites Qwen3 Technical Report.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Qwen3 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.127668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.127668Z digest=sha256:9c6285123214275bea9fcbd1181e27f89e89bcb021c44e90cc3f61e548bfaa08

Observation 52da062c-f74c-4ed1-8a55-a8b24269c144 · outbound

This paper cites M.; Mathieu, M.; Dudzik, A.; Chung, J.; Choi, D.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks M.; Mathieu, M.; Dudzik, A.; Chung, J.; Choi, D

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:30:30.388120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T20:30:30.130712Z digest=sha256:7b036a0959ce23373c5041be71878999093b2296f9fe79d3a7b8e6b531625a45

Observation f63a0b5f-976e-4d20-be07-1e6fc91df16b · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.133129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.133129Z digest=sha256:0aef77acd800a4ca1f5c57c22279a389356f1deb9d267711e354d35609107659

Observation b3884b65-701e-440e-ad64-110431ccac56 · outbound

This paper cites Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.136060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.136060Z digest=sha256:c7df0773764c8723387b4c94f06d6809b65e382e37a29bc585d6e4a90328a910

Observation 78f0f4cc-a38c-47d8-95eb-2a376f230b70 · outbound

This paper cites Avalon's Game of Thoughts: Battle Against Deception through Recursive Contemplation.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Avalon's Game of Thoughts: Battle Against Deception through Recursive Contemplation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.139169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.139169Z digest=sha256:db3bb14a6a2b21b68a99b896d40f4e410f192e226bb1aac12707a6c8ccc51fab

Observation a81d78a3-00e7-4120-bc75-c523599c43e7 · outbound

This paper cites V.; Zhou, D.; et al.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks V.; Zhou, D.; et al

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.146973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.146973Z digest=sha256:877dad8ef770d92b5284022b73f26e91a9b4894d4bdcb3b3d2b7eef42f5280a5

Observation 7ec19a5a-b486-460d-bf68-d71f42a55245 · outbound

This paper cites Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.149845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.149845Z digest=sha256:3cccbe32e7ae94c764b4fc0c73402cd6af241985c5554a3d95700d28698aca56

Observation 17bad6cb-08d8-4fa7-ae26-f9c03a337197 · outbound

This paper cites an unresolved cited work.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.152711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.152711Z digest=sha256:b2ab6299661bc8d827c6ec739b24fd798b1faa18024ca4a8c609aa95dd789d28

Observation decb996e-fb87-4d6e-bca9-4778c7a89e27 · outbound

This paper cites an unresolved cited work.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.156489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.156489Z digest=sha256:59756a45314fb8c10d2418db6949303c3855f813f19d020c8124cbdc3a695dca

Observation 35b59c4f-6776-4b7f-b325-341e092b3b3b · outbound

This paper cites an unresolved cited work.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.159198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.159198Z digest=sha256:4e3c90a12d873e628e00dc161e52b9627959711a6585257815c5083f99c5493f

Observation 7053f5d5-dea0-40bc-9503-8717edc46fbd · outbound

This paper cites Advancing LLM Reasoning Generalists with Preference Trees.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Advancing LLM Reasoning Generalists with Preference Trees

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.162000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.162000Z digest=sha256:63277553b4402ce979a1d2e8a45bb059cc9c4d696d69cc3d68a64be939ad71d3

Observation 6a41886f-e875-406b-9aac-2dd5df8f76f6 · outbound

This paper cites Y.; Ju, J.; Nguyen, A.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Y.; Ju, J.; Nguyen, A

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:30:30.357438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T20:30:30.164783Z digest=sha256:6604418b6d7f685ada1a5ccd29c04acd1808d7c6d797e3fd0e35976efd9a53e7

Observation 86275c19-63dd-4d84-b96b-ff9cbb9321a8 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.167398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.167398Z digest=sha256:0b82ac47d6d8ef3e455e33a32b776d998952953473c2c031c2067951b940f00b

Pith citing papers

Observation 0b0044aa-d816-46ff-82d3-be3a138f2e0f · inbound

RTSGameBench: An RTS Benchmark for Strategic Reasoning by Vision-Language Models cites this paper.

RTSGameBench: An RTS Benchmark for Strategic Reasoning by Vision-Language Models SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:17.302344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T21:02:26.081166Z digest=sha256:b80421649d7cc2d03ed5161d036656a2e98a76c6cc0c3ccf9337cd1a78bfdf6b