Pith. sign in

Paper Citation Record · LEDGER

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training

As of 8 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2509.06053.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.06053 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:36:02.619219Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact2
  • verified fuzzy11
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 76069711-87f4-4052-8c95-f8b4ad411aac · outbound

This paper cites Reinforcement learning in robotics: A survey.The International Journal of Robotics Research, 32(11):1238–1274, 2013.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Reinforcement learning in robotics: A survey.The International Journal of Robotics Research, 32(11):1238–1274, 2013

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T04:36:01.718482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:36:01.718482Z digest=sha256:10fd209c38e0ad7b731ad19953793a882b578e0c928ec14a8cbc8f8ecd01fb9c

Observation 8adbcd51-4595-448a-a98f-249794217e54 · outbound

This paper cites Sim-to-real robot learning from pixels with progressive nets.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Sim-to-real robot learning from pixels with progressive nets

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:36:03.881420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T04:36:01.757271Z digest=sha256:2112099d72db501ebb4e39bd2eea223707bf779003f35a9731851be64cdafcfd

Observation a4d57b60-a009-40e3-86e1-db4827171bc6 · outbound

This paper cites Google research football: A novel reinforcement learning environment.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Google research football: A novel reinforcement learning environment

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:36:03.873689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T04:36:01.838723Z digest=sha256:bf69938bb2dbd8527b1bac6cea7310df87f1cbdf1359bc91d821792f86c5a936

Observation 60035789-2acb-46d4-939c-dc77ac2ab32a · outbound

This paper cites A comprehensive review of multi-agent reinforcement learning in video games.IEEE Transactions on Games, 2025.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training A comprehensive review of multi-agent reinforcement learning in video games.IEEE Transactions on Games, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:36:03.866070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T04:36:01.932070Z digest=sha256:c103ec5d66d6e4285ddbb958a56fc802406bae41d159a37ede1234bde9c08e94

Observation 7aa75364-9c6c-42ee-89bc-0c6262147ced · outbound

This paper cites Robustness and sample complexity of model-based marl for general-sum markov games.Dynamic Games and Applications, 13(1):56–88, 2023.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Robustness and sample complexity of model-based marl for general-sum markov games.Dynamic Games and Applications, 13(1):56–88, 2023

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:36:03.858739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T04:36:01.986521Z digest=sha256:7e6ff18e22516feaab578df8471514a5fe30c49a4c54b38921c05f10fe7c76b3

Observation d96436a9-e1dc-4079-80a9-0716bb5a3cd9 · outbound

This paper cites Multi-agent reinforcement learning for autonomous driving: A survey.CoRR, 2024.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Multi-agent reinforcement learning for autonomous driving: A survey.CoRR, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:36:03.850909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T04:36:02.046727Z digest=sha256:89f634cb48ec952035987eb7d69c8f53d53ad391dd0c37d6c077021ca7795008

Observation cfcdfffe-0dfd-4d28-8cba-ca5f3d0de4ed · outbound

This paper cites Deep reinforcement learning for autonomous driving: A survey.IEEE transactions on intelligent transportation systems, 23(6):4909–4926, 2021.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Deep reinforcement learning for autonomous driving: A survey.IEEE transactions on intelligent transportation systems, 23(6):4909–4926, 2021

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T04:36:02.077841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:36:02.077841Z digest=sha256:9b08a3e64c6f48cf559cd16733a78615b6c94539d7dedbe0481064658f3f9aa5

Observation 9825a211-cfb7-454e-98c8-5c2db3c02ff8 · outbound

This paper cites Equilibrium selection for multi-agent reinforcement learning: A unified framework.arXiv preprint arXiv:2406.08844, 2024.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Equilibrium selection for multi-agent reinforcement learning: A unified framework.arXiv preprint arXiv:2406.08844, 2024

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-08-05T04:36:03.079984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T04:36:02.132762Z digest=sha256:ce219004a667842cab22de3b3c4b1a069a360f538fa6227175dde7bd786374b1

Observation 9987f2c2-1689-4a8e-8760-18d0312108a2 · outbound

This paper cites Emergent reciprocity and team formation from randomized uncertain social preferences.Advances in neural information processing systems, 33:15786–15799, 2020.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Emergent reciprocity and team formation from randomized uncertain social preferences.Advances in neural information processing systems, 33:15786–15799, 2020

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:36:03.837935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T04:36:02.184973Z digest=sha256:e357cfed8142fe92cce1c410719852a9cd81dc000506045b09029dccd5c46f72

Observation 32ae72e4-5c42-4870-9f13-c7450bd67df1 · outbound

This paper cites Taxai: A dynamic economic simulator and benchmark for multi-agent reinforcement learning.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Taxai: A dynamic economic simulator and benchmark for multi-agent reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:36:03.772790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T04:36:02.269117Z digest=sha256:1687a6908b9fb9cc0fe0e12ba5119aa2d842cfca9ff926621cbf2da49b499747

Observation eb2d90d7-842f-419e-8e66-3481a3a2bc57 · outbound

This paper cites A new approach to solving smac task: Generating decision tree code from large language models.arXiv e-prints, pages arXiv–2410, 2024.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training A new approach to solving smac task: Generating decision tree code from large language models.arXiv e-prints, pages arXiv–2410, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:36:03.676990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T04:36:02.321874Z digest=sha256:6580b32fee3e36d5f4129985a537a93352cc702447d9dc0817317269b4d3ea89

Observation fee254cb-faf1-426e-88f5-9317392408f4 · outbound

This paper cites ADRD: LLM-Driven Autonomous Driving Based on Rule-based Decision Systems.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training ADRD: LLM-Driven Autonomous Driving Based on Rule-based Decision Systems

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-05T04:36:02.769162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T04:36:02.379064Z digest=sha256:88af3822655c7eb1e5a142b8b150babf8aa2a7238188a55cb4adb7085b9ed7a5

Observation 9929797c-8d60-4e4f-b4bb-40cab986b4bd · outbound

This paper cites Language models speed up local search for finding programmatic policies.Transactions on Machine Learning Research, 20(X), 2024.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Language models speed up local search for finding programmatic policies.Transactions on Machine Learning Research, 20(X), 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:36:03.500218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T04:36:02.461787Z digest=sha256:12a9d739e8e61ed97cc33a4d503b6176d92426747be7a40f9589fa867156cfaa

Observation 55979992-ad1f-4405-abe3-e55b1f3cd0d8 · outbound

This paper cites Synthesizing programmatic reinforcement learning policies with large language model guided search.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Synthesizing programmatic reinforcement learning policies with large language model guided search

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:36:03.328505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T04:36:02.507355Z digest=sha256:db3deaccbfd32424293c933ef04b5f0e69a0eac8d0fd4efbd9e36b46bb684604

Observation 8af0bf21-4f54-4731-88bc-7bfe7d8d2ad2 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T04:36:02.563015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:36:02.563015Z digest=sha256:f2e357acbb587697a6edac5c579d95556e49a304dbf8d897e09c6cea13b0c846

Observation 261bd7e9-b332-4bd3-b498-d0d70c55ef7a · outbound

This paper cites React: Synergizing reasoning and acting in language models.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training React: Synergizing reasoning and acting in language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:36:03.210132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T04:36:02.619219Z digest=sha256:4261a7da350199063252a4d0849f9bcd3f56385649e36e4bc78ecd74eec757a3

Pith citing papers

No inbound Pith citation observations are available.