Pith. sign in

Paper Citation Record · LEDGER

Leveraging Procedural Generation to Benchmark Reinforcement Learning

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:1912.01588.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1912.01588 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:59:40.090855Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

171
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 253eddd0-5bf8-4c84-9f40-fa56f9e8ab76 · inbound

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution cites this paper.

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:12:31.413770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-16T08:12:30.984870Z digest=sha256:8018ee9e05d016939eb62f87f016e618c54cbaef7c0b0b4aa26a850c13ed6a68

Observation 104aae9c-0c71-449a-9de8-c9c97c00416d · inbound

Proximal Policy Distillation cites this paper.

Proximal Policy Distillation Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-23T23:13:36.980164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T23:09:42.421333Z digest=sha256:38afb1b721f142e0eaa5c6444324e2127fd0a166e1ba0a5136d4a56b6acf8ebe

Observation b42d011d-d382-4b29-ae7a-c14f333f4633 · inbound

An Open-Source Software Toolkit & Benchmark Suite for the Evaluation and Adaptation of Multimodal Action Models cites this paper.

An Open-Source Software Toolkit & Benchmark Suite for the Evaluation and Adaptation of Multimodal Action Models Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T04:59:40.090855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:59:40.090855Z digest=sha256:324b7c35675373a9131d6e0867f179815594b30d61c8db5d882cdf527b83a535

Observation 2d7ab01c-84ce-482a-990a-51ff1f76c9ce · inbound

Dyn-O: Building Structured World Models with Object-Centric Representations cites this paper.

Dyn-O: Building Structured World Models with Object-Centric Representations Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:02.762041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:02.762041Z digest=sha256:b3006db7e11f32ded6e6342e9486759bad4f622d2a6dfdc381d6a8724b09b6e8

Observation 6d28784e-5ce8-4fe5-8a07-7470b597d00c · inbound

Preemptive Solving of Future Problems: Multitask Preplay in Humans and Machines cites this paper.

Preemptive Solving of Future Problems: Multitask Preplay in Humans and Machines Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:42:07.614826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T06:38:33.969927Z digest=sha256:14b5f7fd94c729c821631a3fc8d442ac81d9be121f0dc6f461a701de3e35ff0c

Observation be0497ab-3a38-4fa9-a49f-2eeb5dbd12c0 · inbound

Assistax: A Multi-Agent Hardware-Accelerated Reinforcement Learning Benchmark for Assistive Robotics cites this paper.

Assistax: A Multi-Agent Hardware-Accelerated Reinforcement Learning Benchmark for Assistive Robotics Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T12:38:46.903783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:38:46.903783Z digest=sha256:d4e7eee956a8ee8661456464cbafca77f1a3d586b54ffdbdbe971ae775ed7c24

Observation e1284f15-6fcf-4e2e-89e0-e9ca4c5edb42 · inbound

Benchmarking Partial Observability in Reinforcement Learning with a Suite of Memory-Improvable Domains cites this paper.

Benchmarking Partial Observability in Reinforcement Learning with a Suite of Memory-Improvable Domains Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 1992

Resolution
unresolved
no resolver link, observed 2026-08-06T10:34:35.473241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:34:35.473241Z digest=sha256:6ab3155a841747701d40e1daa21572777ca97b5823415d8ddfaf306b0e3fab73

Observation 41bf3f85-9e3b-4c37-bee0-20063b01e320 · inbound

On the Sample Efficiency of Inverse Dynamics Models for Semi-Supervised Imitation Learning cites this paper.

On the Sample Efficiency of Inverse Dynamics Models for Semi-Supervised Imitation Learning Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T05:21:35.522064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:21:35.522064Z digest=sha256:42a79483e0885e6fc22549167e9d42195211b9010cae211c403edd6629fc218f

Observation ff30be29-828b-4205-8356-f5a39941ec16 · inbound

Understanding Goal Generalisation in Sequential Reinforcement Learning cites this paper.

Understanding Goal Generalisation in Sequential Reinforcement Learning Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:50:20.671717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-25T04:49:50.034743Z digest=sha256:678bf78783497bba4f85ef5fc9a0b0fa1422279e99a20f07bb1635cf0b0f7b07

Observation 70afcc4e-51ae-4410-9c78-348c1d0cdaef · inbound

CLAW: Learning Continuous Latent Action World Models via Adversarial Latent Regularization cites this paper.

CLAW: Learning Continuous Latent Action World Models via Adversarial Latent Regularization Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:36:29.552457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T09:55:00.402411Z digest=sha256:4581616b1eaa19246201edd3ce7c75ee9a014008fd6cb3c1a04f6c10127be2fe

Observation d063ccc5-8c68-409d-b95c-a9aaa4d03dc4 · inbound

Reinforcement Learning Foundation Models Should Already Be A Thing cites this paper.

Reinforcement Learning Foundation Models Should Already Be A Thing Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:39:05.044993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T21:50:00.974590Z digest=sha256:e3af7dcacf6dc2469d4bbb6eaf7419f28a300606f2f6a1e98b559951d286bb06

Observation d81c660d-93d4-4273-92f3-c62ff223fc21 · inbound

Reasoning Core: Designing Broad Procedural Data for Completion-Supervised Reasoning Training cites this paper.

Reasoning Core: Designing Broad Procedural Data for Completion-Supervised Reasoning Training Leveraging Procedural Generation to Benchmark Reinforcement Learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T04:19:48.168776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:19:48.168776Z digest=sha256:ff23609ccd3e0cd73e50466ec81b3289dc32bd3df73084ac179985e2d0db30cd