Pith. sign in

Paper Citation Record · LEDGER

On-Policy RL with Optimal Reward Baseline

As of 19 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 12 inbound Pith citation observations for arXiv:2505.23585.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23585 v2

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:46:50.980711Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:07:10.581791Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:59:40.759572Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2dcd3541-28e0-4613-beeb-c5457c7d466b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

On-Policy RL with Optimal Reward Baseline DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:49.967939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:49.967939Z digest=sha256:b407c75b69717230c3181597d504cab08487fa88583727428f7866274b5600c3

Observation 2e642235-367e-4ce6-96e0-6e50fc2734c9 · outbound

This paper cites Approximately optimal approximate reinforcement learning.

On-Policy RL with Optimal Reward Baseline Approximately optimal approximate reinforcement learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:51.636308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:46:50.133721Z digest=sha256:dbe9170cb1a6b3e7d2f6b7afd5586713be0bd865290f86cd8e430c9ecb9f4537

Observation 19022d9d-dd2a-46a8-b3ea-bad1333745d1 · outbound

This paper cites DeepSeek-V3 Technical Report.

On-Policy RL with Optimal Reward Baseline DeepSeek-V3 Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:50.292119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:50.292119Z digest=sha256:a08f7d1f8026fb9a6b38aa0a4ddfc20421f0587162f9c15e9678cba3508d4065

Observation 5bb3f326-21a2-4b99-bbd4-319545fc2b19 · outbound

This paper cites American invitational mathematics examination - aime.

On-Policy RL with Optimal Reward Baseline American invitational mathematics examination - aime

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:51.488207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:46:50.380147Z digest=sha256:754db6f4225f516898615dfa680240c59d321511e6a4a17cd05cc04278992adb

Observation d49e6d92-b1f3-46eb-8e7b-fd09ebf946af · outbound

This paper cites GPT-4 Technical Report.

On-Policy RL with Optimal Reward Baseline GPT-4 Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:50.477043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:50.477043Z digest=sha256:47acfb6dc66de4a88cbabba345631948c06f37dd9544c47a9d8bc78ca8b393a1

Observation bf99ce1d-558d-41a9-be08-e083b247bb7c · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

On-Policy RL with Optimal Reward Baseline HybridFlow: A Flexible and Efficient RLHF Framework

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:50.703608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:50.703608Z digest=sha256:8ba50539cea43834e318ec4705c62df6fb1b83782e6cbce502edfc2283078a6e

Observation c5fc5994-bdba-4e76-ba57-475127ed913e · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

On-Policy RL with Optimal Reward Baseline Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:50.798376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:50.798376Z digest=sha256:2b5904eafb05927ec9941f44bcf139041b685ffc6d9686f5ce7030d5fb92bef1

Observation 0f4e0d30-1fee-407c-b9f1-0db0ecc45304 · outbound

This paper cites Qwen3 Technical Report.

On-Policy RL with Optimal Reward Baseline Qwen3 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:50.980711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:50.980711Z digest=sha256:ec27ef32581c9bc2c0da0db88668729a9ec3baa825dee955094f5dd629b1d0aa

Observation 24f923a4-be90-4d91-9fbb-4fed684d2973 · outbound

This paper cites Neural text generation with unlikelihood training.

On-Policy RL with Optimal Reward Baseline Neural text generation with unlikelihood training

Reference 1992

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:51.316058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T12:46:50.893301Z digest=sha256:2114a5c7fbcf1c654d3526a8793395226d8a6e254bfc0572783b293d56a72056

Observation c3930687-7926-4244-a3bb-9abbaa758984 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

On-Policy RL with Optimal Reward Baseline DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:50.635625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:50.635625Z digest=sha256:7041c7146b8d9ee863035270598dd4d63c9d33466b8daf180763567e7112bdef

Observation a37069ee-a91e-4bf0-8578-d051882fec5f · outbound

This paper cites Proximal Policy Optimization Algorithms.

On-Policy RL with Optimal Reward Baseline Proximal Policy Optimization Algorithms

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:50.540130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:50.540130Z digest=sha256:335c060ee457986fced406fedf5883f7bae5529cae18d556e86aa180f83b8d22

Observation 956640e0-8b55-45e4-96bf-058de5bdc983 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

On-Policy RL with Optimal Reward Baseline Gemini: A Family of Highly Capable Multimodal Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:49.833777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:49.833777Z digest=sha256:eb0ea81203eac476d87818705e5376d9da041717c4fffc9d93f06ac8490f3ae5

Observation c10971a2-1dd0-4e53-b562-bd4a54572573 · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,.

On-Policy RL with Optimal Reward Baseline Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:49.876159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:49.876159Z digest=sha256:3dc5d3d4f3fc2c0d0bd2c7475fa2dd5c74e924706ae665f82b01442bdc85b6b4

Observation afc7fb7b-9b0f-46ce-9d99-f4240d1b0608 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

On-Policy RL with Optimal Reward Baseline Understanding R1-Zero-Like Training: A Critical Perspective

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:50.225237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:50.225237Z digest=sha256:42cf46d490f31c37c9c7d6d236455dc931d3a8afe4ed65e3c8ac8bbc4652f1a9

Observation 370b104d-8aa6-45cb-a7fc-fc06aed6c4b3 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

On-Policy RL with Optimal Reward Baseline REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:50.044130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:50.044130Z digest=sha256:014180833d52e76fd82a424144858e41f39efbb410417895b7b2e0bb7227df90

Pith citing papers

Observation 2734aa74-a458-4f4d-93c9-0945d4426584 · inbound

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities cites this paper.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities On-Policy RL with Optimal Reward Baseline

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:10.581791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:10.581791Z digest=sha256:89567e0265f5381c87e22b0d68b0a2ff932efcb1a29b6b07e7b29680b5104e31

Observation 8793b3f4-1be9-4745-a4f2-e769ce5826b0 · inbound

Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective cites this paper.

Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective On-Policy RL with Optimal Reward Baseline

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:41:03.216079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T07:37:12.501489Z digest=sha256:ac4c5bdaf36eafa187a67680fb2c98655da5c9ab7e125dac49ead57be2c0a6ce

Observation 9db8ea4c-ef83-401e-a53e-06dc3fb98d35 · inbound

Policy Improvement Reinforcement Learning cites this paper.

Policy Improvement Reinforcement Learning On-Policy RL with Optimal Reward Baseline

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:48:22.858481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T22:47:05.132020Z digest=sha256:0d4b82d2442a8a581895f23395462afb503fcdb8de5cb56010df4d8ee74cb26f

Observation 0d264321-b1b6-4599-8e4a-c783024c61a8 · inbound

Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning cites this paper.

Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning On-Policy RL with Optimal Reward Baseline

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:27.278990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-07T07:00:32.206081Z digest=sha256:62e4e24e9f554308a675dc44ee255bd17ac469145bc6836a3c21a314d07c4ff3

Observation 752d6683-b8cb-479f-b559-908239a3875c · inbound

Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning cites this paper.

Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning On-Policy RL with Optimal Reward Baseline

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:49:14.929225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T23:47:53.282259Z digest=sha256:7893ead834b0775ad433361ad7b236338b05002cb4c68734796de0ba9ebcf173

Observation 32df8d68-545c-4bb8-aee5-2a2c4803d5fa · inbound

Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR cites this paper.

Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR On-Policy RL with Optimal Reward Baseline

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:10.479798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-09T15:39:54.001327Z digest=sha256:b10b1b7453123a05fece344dabaac43c5029d35205923b2623f8ec214b0b165b

Observation bb1413e4-7d89-463c-8777-db47a92cb050 · inbound

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization cites this paper.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization On-Policy RL with Optimal Reward Baseline

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:42:02.898305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:7400fa2e4f48f54e8f0d6c00ad2638d234c74bcc0b70efa8fe238247cbaa00d6

Observation 65121ee2-c018-4af9-8868-4743fd40aca8 · inbound

Holder Policy Optimisation cites this paper.

Holder Policy Optimisation On-Policy RL with Optimal Reward Baseline

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:12:22.799052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-13T06:08:28.855671Z digest=sha256:011b55a21acec16bab24512ef39036ad6c08427ca03afbfc5fca8465cb70f937

Observation 6d678c24-3e73-4b0a-95e6-a526d8b673f7 · inbound

Holder Policy Optimisation cites this paper.

Holder Policy Optimisation On-Policy RL with Optimal Reward Baseline

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:01:23.127974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:b3d5258fe24eafec49e49918103417a63efe1badf172a4c8565642df0eab074f

Observation 7d1f1c99-71d2-4524-900f-75fea0fce1c7 · inbound

Keep Policy Gradient in Charge: Sibling-Guided Credit Distillation for Long-Horizon Tool-Use Agents cites this paper.

Keep Policy Gradient in Charge: Sibling-Guided Credit Distillation for Long-Horizon Tool-Use Agents On-Policy RL with Optimal Reward Baseline

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:07:48.294529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T10:28:29.586968Z digest=sha256:6cd24d1008ed6bf536bddafba97b30e6e4036472aefd83495610a62c7d1c8942

Observation 7fb79ea2-5234-42bd-9488-99fc41b2e7a0 · inbound

Keep Policy Gradient in Charge: Sibling-Guided Credit Distillation for Long-Horizon Tool-Use Agents cites this paper.

Keep Policy Gradient in Charge: Sibling-Guided Credit Distillation for Long-Horizon Tool-Use Agents On-Policy RL with Optimal Reward Baseline

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:55:31.197198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T07:48:15.475103Z digest=sha256:996e970c36c729c69eee049b1b4d248a6a25fbca5153df88b868ec041c910fa7

Observation feac46a2-371a-4fc9-8797-50a800d304e6 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning On-Policy RL with Optimal Reward Baseline

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.760859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:1894e2b714fcadd321c603f9921043ebe367e8acaebceab58e29802c4be1ffcb