Pith. sign in

Paper Citation Record · LEDGER

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training

As of 20 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2509.06053.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.06053 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:36:02.619219Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact2
  • verified fuzzy11
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 76069711-87f4-4052-8c95-f8b4ad411aac · outbound

This paper cites Reinforcement learning in robotics: A survey.The International Journal of Robotics Research, 32(11):1238–1274, 2013.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Reinforcement learning in robotics: A survey.The International Journal of Robotics Research, 32(11):1238–1274, 2013

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T04:36:01.718482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:36:01.718482Z digest=sha256:7813330aee0b0b98977c7161918e2ef4a0e566c883849301593bdba22656cce4

Observation 8adbcd51-4595-448a-a98f-249794217e54 · outbound

This paper cites Sim-to-real robot learning from pixels with progressive nets.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Sim-to-real robot learning from pixels with progressive nets

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:36:03.881420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T04:36:01.757271Z digest=sha256:6c03f4c03b5515756dc381504ff548b64da7f423939248c6a8ef4f1cbed24d3c

Observation a4d57b60-a009-40e3-86e1-db4827171bc6 · outbound

This paper cites Google research football: A novel reinforcement learning environment.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Google research football: A novel reinforcement learning environment

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:36:03.873689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T04:36:01.838723Z digest=sha256:7ead6c2614add0314db3c26a05e5b3f50cadc05bf9ff386d1bf9f75d80d3aa53

Observation 60035789-2acb-46d4-939c-dc77ac2ab32a · outbound

This paper cites A comprehensive review of multi-agent reinforcement learning in video games.IEEE Transactions on Games, 2025.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training A comprehensive review of multi-agent reinforcement learning in video games.IEEE Transactions on Games, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:36:03.866070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T04:36:01.932070Z digest=sha256:66654f354beeb1e8a2d0d15d9b34b4fefa14664620c95dee99a2c7be7f30bc66

Observation 7aa75364-9c6c-42ee-89bc-0c6262147ced · outbound

This paper cites Robustness and sample complexity of model-based marl for general-sum markov games.Dynamic Games and Applications, 13(1):56–88, 2023.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Robustness and sample complexity of model-based marl for general-sum markov games.Dynamic Games and Applications, 13(1):56–88, 2023

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:36:03.858739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T04:36:01.986521Z digest=sha256:bb12c4cd872020576788b3e1c0ecabae1f66ee72b5bcb6b060548e8a7aebe377

Observation d96436a9-e1dc-4079-80a9-0716bb5a3cd9 · outbound

This paper cites Multi-agent reinforcement learning for autonomous driving: A survey.CoRR, 2024.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Multi-agent reinforcement learning for autonomous driving: A survey.CoRR, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:36:03.850909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T04:36:02.046727Z digest=sha256:d7e3e03025f7af3f69a86455af9aaeab235bc72b4f8220651e2fb0ed5cf443ed

Observation cfcdfffe-0dfd-4d28-8cba-ca5f3d0de4ed · outbound

This paper cites Deep reinforcement learning for autonomous driving: A survey.IEEE transactions on intelligent transportation systems, 23(6):4909–4926, 2021.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Deep reinforcement learning for autonomous driving: A survey.IEEE transactions on intelligent transportation systems, 23(6):4909–4926, 2021

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T04:36:02.077841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:36:02.077841Z digest=sha256:ac460aee7ab14cf2815eadb8efdf56617be8727e086ae50242ed6930003210e6

Observation 9825a211-cfb7-454e-98c8-5c2db3c02ff8 · outbound

This paper cites Equilibrium selection for multi-agent reinforcement learning: A unified framework.arXiv preprint arXiv:2406.08844, 2024.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Equilibrium selection for multi-agent reinforcement learning: A unified framework.arXiv preprint arXiv:2406.08844, 2024

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-08-05T04:36:03.079984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T04:36:02.132762Z digest=sha256:5c342286f9e034f607b6e59e88cefb5f483bc08bcc67bf8850d77680803b84cc

Observation 9987f2c2-1689-4a8e-8760-18d0312108a2 · outbound

This paper cites Emergent reciprocity and team formation from randomized uncertain social preferences.Advances in neural information processing systems, 33:15786–15799, 2020.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Emergent reciprocity and team formation from randomized uncertain social preferences.Advances in neural information processing systems, 33:15786–15799, 2020

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:36:03.837935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T04:36:02.184973Z digest=sha256:a014ab105fa44dcadd4c6a5714f9b509d37d8404c004817c6f14d54f5c38b120

Observation 32ae72e4-5c42-4870-9f13-c7450bd67df1 · outbound

This paper cites Taxai: A dynamic economic simulator and benchmark for multi-agent reinforcement learning.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Taxai: A dynamic economic simulator and benchmark for multi-agent reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:36:03.772790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T04:36:02.269117Z digest=sha256:78200c10bc674e181484effab7cb3e32560aa4220738cf881a819ec9744cf93f

Observation eb2d90d7-842f-419e-8e66-3481a3a2bc57 · outbound

This paper cites A new approach to solving smac task: Generating decision tree code from large language models.arXiv e-prints, pages arXiv–2410, 2024.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training A new approach to solving smac task: Generating decision tree code from large language models.arXiv e-prints, pages arXiv–2410, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:36:03.676990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T04:36:02.321874Z digest=sha256:83b369244e5b8bfeb4a0602fd1e4e3cf2774089cf781e225bdbf3c774367ed30

Observation fee254cb-faf1-426e-88f5-9317392408f4 · outbound

This paper cites ADRD: LLM-Driven Autonomous Driving Based on Rule-based Decision Systems.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training ADRD: LLM-Driven Autonomous Driving Based on Rule-based Decision Systems

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-05T04:36:02.769162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T04:36:02.379064Z digest=sha256:7b224c6899f7faa3ecd4dab8a8a9953c98c6f6e26936071b5405cd30609f11d2

Observation 9929797c-8d60-4e4f-b4bb-40cab986b4bd · outbound

This paper cites Language models speed up local search for finding programmatic policies.Transactions on Machine Learning Research, 20(X), 2024.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Language models speed up local search for finding programmatic policies.Transactions on Machine Learning Research, 20(X), 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:36:03.500218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T04:36:02.461787Z digest=sha256:ed333781bd003ab1068817ad356394ef7af7979a52133aeadf912af5dec7b716

Observation 55979992-ad1f-4405-abe3-e55b1f3cd0d8 · outbound

This paper cites Synthesizing programmatic reinforcement learning policies with large language model guided search.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Synthesizing programmatic reinforcement learning policies with large language model guided search

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:36:03.328505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T04:36:02.507355Z digest=sha256:b4f17aebf689e00ae15f904a50b42f4d7c5234fe6b3debf067a629fff3c4557e

Observation 8af0bf21-4f54-4731-88bc-7bfe7d8d2ad2 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T04:36:02.563015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:36:02.563015Z digest=sha256:33b8fc63d61389fdd600d98c7319b92642cc04cf121ad21e60abbc71aa24e154

Observation 261bd7e9-b332-4bd3-b498-d0d70c55ef7a · outbound

This paper cites React: Synergizing reasoning and acting in language models.

PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training React: Synergizing reasoning and acting in language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:36:03.210132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T04:36:02.619219Z digest=sha256:78e90b83201440c0228830fc49b9a98a52c1305de6170b6ec9dbc469e3f42b01

Pith citing papers

No inbound Pith citation observations are available.