Pith. sign in

Paper Citation Record · LEDGER

What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2006.05990.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2006.05990 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T04:36:12.619468Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

105
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 11bf0bf6-346d-4a68-8682-77a61759fe70 · inbound

What Matters in Learning from Offline Human Demonstrations for Robot Manipulation cites this paper.

What Matters in Learning from Offline Human Demonstrations for Robot Manipulation What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-13T08:51:55.997703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T08:51:55.826747Z digest=sha256:6f12fbc99cba610dcc575b4236b86cf4803299eadb0b5c71fbfd45f2e9d58f0d

Observation edd1cefa-a79c-401e-a686-fd5a82942d12 · inbound

The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models cites this paper.

The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:08:52.801639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T07:08:52.723528Z digest=sha256:d37e12c467507efdbf0064c264ac8756acadd752e4af94f4d0106f514ded06db

Observation 5e7bc265-bb22-4746-9f1c-6b5223fd8b4b · inbound

Mastering Diverse Domains through World Models cites this paper.

Mastering Diverse Domains through World Models What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:08:21.980352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T09:08:21.677362Z digest=sha256:bc4f62ad94df188a456f18de4e22a57ef18bf93bca7c90fc92e3a84b381d5289

Observation 0c67352e-c4d3-4b5f-b882-fb53c819e4ef · inbound

Simulating Errors in Touchscreen Typing cites this paper.

Simulating Errors in Touchscreen Typing What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T04:36:12.619468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:36:12.619468Z digest=sha256:0b814bf9ac6bfcb5c3cb10f7a9527005212e02dac9879a778a9871b09448db06

Observation 87b35a59-608e-447d-9477-3e97bb1daa7f · inbound

MuJoCo Playground cites this paper.

MuJoCo Playground What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T23:34:10.862569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:34:10.862569Z digest=sha256:a9e0b7d8c8a9b07ae42935860538fb18dec8aa65f3f094c887b3a5e8efe34fea

Observation b4b230da-84bc-4768-b4c3-268d49f2c8c0 · inbound

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming cites this paper.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.205546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.205546Z digest=sha256:0c7c637253077867ecf33ab44258928dbb860822579bd3e6225af13b45be8dbf

Observation 58df91f7-f6d0-43f2-914c-f803515f43e7 · inbound

Magistral cites this paper.

Magistral What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:22.151310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:22.151310Z digest=sha256:df5a0ded52aac656333ab6a7ebe0eeff3be5471d5640554e5ccffa697bed21fc

Observation 9bbad98f-46c2-43b4-bc7a-bd421c171fc0 · inbound

Communicating Smartly in Molecular Communication Environments: Neural Networks in the Internet of Bio-Nano Things cites this paper.

Communicating Smartly in Molecular Communication Environments: Neural Networks in the Internet of Bio-Nano Things What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T22:48:46.211558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:48:46.211558Z digest=sha256:79b7839f94568735dd94e5dcc7051d2f039899f545ff84eda11a30125e4e620f

Observation 05e8926f-c64c-4013-be31-a7adcfd16cf7 · inbound

On the Effect of Regularization in Policy Mirror Descent cites this paper.

On the Effect of Regularization in Policy Mirror Descent What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:44.262283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:18:44.262283Z digest=sha256:d13fb2a0948545abc3af286f19aaeecdf987eeae7a58044960e2bcb4328b2034

Observation 0922a89a-ec0c-4f0d-984e-0897a8430c28 · inbound

Learning human-to-robot handovers through 3D scene reconstruction cites this paper.

Learning human-to-robot handovers through 3D scene reconstruction What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:26.684738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:26.684738Z digest=sha256:90d2876a6acd7ed2838ca8da40cad6d5a94e2bed3bb025e4bc017f770801b5e0

Observation b0bd07b4-47d3-45fb-8dd6-c316ec3a8241 · inbound

Online Training and Pruning of Deep Reinforcement Learning Networks cites this paper.

Online Training and Pruning of Deep Reinforcement Learning Networks What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:05:16.167535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:05:16.167535Z digest=sha256:410f01814bea228331275fc4525bcf7c923fdd404e5aae49912342e16f612aa2

Observation 4d1af3aa-8b4b-4dea-865c-f087b1bfce4d · inbound

Differentiable Weightless Controllers: Learning Logic Circuits for Continuous Control cites this paper.

Differentiable Weightless Controllers: Learning Logic Circuits for Continuous Control What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-03T19:16:26.948543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:16:26.948543Z digest=sha256:e0d8723f023d1ff3b5fd8398ac0384ebb1da44bd62ccf81daeecabae218a8433

Observation 16ea3926-2720-442f-ae65-02428ba04317 · inbound

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments cites this paper.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:20.216901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:20.216901Z digest=sha256:47811463cf9b7535df809c2e634fed255bf0e0c08bdd00c0900694b78e788200

Observation b9348a80-ac48-439d-9a34-652d31809979 · inbound

From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space cites this paper.

From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:41:04.216628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T12:50:57.603403Z digest=sha256:569794f3ccbd7b50a6a2942de1e12018286a2d8b1dcb29dcf8d32b5f9f4217ec

Observation 0897607c-7adc-46df-8bc8-f1172f451575 · inbound

Low-Rank Adaptation for Critic Learning in Off-Policy Reinforcement Learning cites this paper.

Low-Rank Adaptation for Critic Learning in Off-Policy Reinforcement Learning What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:27.889717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T02:55:22.577012Z digest=sha256:0b0f4916b4182d6e8a97c041c237b6de27fb0ff8f4fd4d61ad40b9fa532f761d

Observation 810aace4-d58a-49c6-85b6-0fb17d21ba16 · inbound

Application of Deep Reinforcement Learning to Event-Triggered Control for Networked Artificial Pancreas Systems cites this paper.

Application of Deep Reinforcement Learning to Event-Triggered Control for Networked Artificial Pancreas Systems What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:41:16.498617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T14:51:30.263897Z digest=sha256:49e6d284fd6d06683aeae7122a1e6198dadc040ee11cdfdcbb5e27b778bff58e

Observation ec7d7e36-c90a-4688-b726-47c03e42b11c · inbound

Application of Deep Reinforcement Learning to Event-Triggered Control for Networked Artificial Pancreas Systems cites this paper.

Application of Deep Reinforcement Learning to Event-Triggered Control for Networked Artificial Pancreas Systems What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:47:41.630296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T17:46:20.907461Z digest=sha256:655eda323aabea9665805f1d018b60ee9997c8f67609a2698b6174cd6ba978b9

Observation e5544742-6a92-4df8-8429-0affd16ea488 · inbound

Hint Tuning: Less Data Makes Better Reasoners cites this paper.

Hint Tuning: Less Data Makes Better Reasoners What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:26.341552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T00:54:13.146373Z digest=sha256:0477289585a2529a00c4e6311046ec43e5bee7cb905514be23b4f432ccce4087

Observation 2306338e-254c-4b5e-a92e-c19e3e962391 · inbound

Hint Tuning: Less Data Makes Better Reasoners cites this paper.

Hint Tuning: Less Data Makes Better Reasoners What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:35:07.063312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T23:34:43.785312Z digest=sha256:1089c21f43d59eb5665c4c0894cbdce72c1b07732b4d7a0cd33078ab8dd94721

Observation fd05a4da-3bbf-4276-a8e5-fcb36ddeedbf · inbound

TuniQ: Autotuning Compilation Passes for Quantum Workloads at Scale for Effectiveness and Efficiency cites this paper.

TuniQ: Autotuning Compilation Passes for Quantum Workloads at Scale for Effectiveness and Efficiency What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:52:08.830825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T02:50:20.040271Z digest=sha256:43368dce8891893710c81e83a302620f616e3ed5db6638548e200f98519080c6

Observation ed702539-4593-4c4c-9d30-38c779010c53 · inbound

TOPPO: Rethinking PPO for Multi-Task Reinforcement Learning with Critic Balancing cites this paper.

TOPPO: Rethinking PPO for Multi-Task Reinforcement Learning with Critic Balancing What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:52:05.744552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T01:48:13.679862Z digest=sha256:ea00724432994c63fbd63e21e281dd68fae0b457dc1acb0e627fe8570d7c150d

Observation e8ccc653-d395-4959-978f-9d755b9a10bb · inbound

Randomized Advantage Transformation (RAT): Computing Natural Policy Gradients via Direct Backpropagation cites this paper.

Randomized Advantage Transformation (RAT): Computing Natural Policy Gradients via Direct Backpropagation What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 98

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:48:17.499793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T12:44:29.147095Z digest=sha256:dcb76c1234d1827a8a4c70daaabe331ca875640d64bb23a29526aaabddc0b4f7

Observation 4a2c263d-a5cb-48b1-aca4-b3c0ef4abb7b · inbound

When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks? cites this paper.

When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks? What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:42:49.605925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T23:38:33.520625Z digest=sha256:6a157f5972fd57d7a07c3c6c2cc57f9b1b0e18244cede2683a3574391d69a5b7

Observation 4249dc2e-3a72-431a-b874-9af188867e0a · inbound

When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks? cites this paper.

When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks? What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-29T05:43:07.739085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T05:42:55.457904Z digest=sha256:5f3d362e4df05b728e9b1527098a7773b9c2d21fa0f9ffe9557b4c73f027d6e7

Observation ea047044-d8ca-484c-88d2-a6d70d7ccaf6 · inbound

PowerOPD: Stabilizing On-Policy Distillation with Bounded Power Transformation cites this paper.

PowerOPD: Stabilizing On-Policy Distillation with Bounded Power Transformation What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:48:46.310519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T03:39:37.096675Z digest=sha256:c04eaea0afb356dedd3015050a0e9203d9453289663bd36cbb820634acc6b347

Observation 69add4d6-dddd-4560-a164-723cc1443895 · inbound

Learning to Trigger: Reinforcement Learning at the Large Hadron Collider cites this paper.

Learning to Trigger: Reinforcement Learning at the Large Hadron Collider What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:39:45.862272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T08:35:33.377206Z digest=sha256:2c83426f986fe45ff7016860a1e92dd1ab017ea55fad6e6178ecf46b185a3cac

Observation ae2dde90-abce-4a0e-ae84-b3abf3babac2 · inbound

Learning to Trigger: Reinforcement Learning at the Large Hadron Collider cites this paper.

Learning to Trigger: Reinforcement Learning at the Large Hadron Collider What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:54:39.123007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T10:20:52.408426Z digest=sha256:5f99f9a67d733d27a4f37a552cbecc89b52f89985d6a8d109044bea4e4061bc5

Observation c078e5ea-f243-46fc-a542-efdd9bed50b0 · inbound

The Importance of Encoder Choice:A Tabular-Image Study cites this paper.

The Importance of Encoder Choice:A Tabular-Image Study What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 168

Resolution
verified exact
local_arxiv, observed 2026-07-10T19:07:34.906671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-10T19:03:32.353393Z digest=sha256:ab071cf1bf94837021990af754c289e9958ac04f4e6a7c36373ab02dc7078110