Pith. sign in

Paper Citation Record · LEDGER

On-Policy RL with Optimal Reward Baseline

As of 20 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 12 inbound Pith citation observations for arXiv:2505.23585.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23585 v2

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:46:50.980711Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:07:10.581791Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:59:40.759572Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2dcd3541-28e0-4613-beeb-c5457c7d466b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

On-Policy RL with Optimal Reward Baseline DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:49.967939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:49.967939Z digest=sha256:eab7fc23298e4b23fc76e588c86bd2852c881d577eb6a7e38bd1ddc65b18bece

Observation 2e642235-367e-4ce6-96e0-6e50fc2734c9 · outbound

This paper cites Approximately optimal approximate reinforcement learning.

On-Policy RL with Optimal Reward Baseline Approximately optimal approximate reinforcement learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:51.636308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:46:50.133721Z digest=sha256:1282e3665e177437cb6c46873e23f66f986cacec7fbff34e36d42685a9b49733

Observation 19022d9d-dd2a-46a8-b3ea-bad1333745d1 · outbound

This paper cites DeepSeek-V3 Technical Report.

On-Policy RL with Optimal Reward Baseline DeepSeek-V3 Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:50.292119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:50.292119Z digest=sha256:bfe1b2fd2f2a9015be3206fbf4e49f2538a3b977e48947c060cbc6ce7ae39a49

Observation 5bb3f326-21a2-4b99-bbd4-319545fc2b19 · outbound

This paper cites American invitational mathematics examination - aime.

On-Policy RL with Optimal Reward Baseline American invitational mathematics examination - aime

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:51.488207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:46:50.380147Z digest=sha256:8e1e13e051b267ee8071cabc5b33d5bd65feabdca4446850c2540d5f6b57b98d

Observation d49e6d92-b1f3-46eb-8e7b-fd09ebf946af · outbound

This paper cites GPT-4 Technical Report.

On-Policy RL with Optimal Reward Baseline GPT-4 Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:50.477043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:50.477043Z digest=sha256:0b2ee45310ef72db98d62f3b9c9051e21064fb74c8b72589c69ecf9684e0b9da

Observation bf99ce1d-558d-41a9-be08-e083b247bb7c · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

On-Policy RL with Optimal Reward Baseline HybridFlow: A Flexible and Efficient RLHF Framework

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:50.703608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:50.703608Z digest=sha256:2e54ede9ab16e88108e42e705fb3fa2d1920a29ab14f57598de04110a3b7e555

Observation c5fc5994-bdba-4e76-ba57-475127ed913e · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

On-Policy RL with Optimal Reward Baseline Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:50.798376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:50.798376Z digest=sha256:0c5d72c3cfcae9b0db3189fd12814fc59353a3d24726f73587b448ae82c3f867

Observation 0f4e0d30-1fee-407c-b9f1-0db0ecc45304 · outbound

This paper cites Qwen3 Technical Report.

On-Policy RL with Optimal Reward Baseline Qwen3 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:50.980711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:50.980711Z digest=sha256:a1d2da66d98b4bf2e4f33234daed6b3b07284a9da7fd557c85ec78750597100d

Observation 24f923a4-be90-4d91-9fbb-4fed684d2973 · outbound

This paper cites Neural text generation with unlikelihood training.

On-Policy RL with Optimal Reward Baseline Neural text generation with unlikelihood training

Reference 1992

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:46:51.316058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:46:50.893301Z digest=sha256:02cd6d40f5b03c558f678e47994bcbedef136c7eb74c7f84cbbb3763530bfc4f

Observation c3930687-7926-4244-a3bb-9abbaa758984 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

On-Policy RL with Optimal Reward Baseline DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:50.635625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:50.635625Z digest=sha256:de5988f3d450527626c1a54a0ac7a5bdaca5017e28dcbf48e521c861b33c5d01

Observation a37069ee-a91e-4bf0-8578-d051882fec5f · outbound

This paper cites Proximal Policy Optimization Algorithms.

On-Policy RL with Optimal Reward Baseline Proximal Policy Optimization Algorithms

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:50.540130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:50.540130Z digest=sha256:77b0de3701d4cf899e6691b53908128c3250bde20fc193c0a887c3f0f2190c7f

Observation 956640e0-8b55-45e4-96bf-058de5bdc983 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

On-Policy RL with Optimal Reward Baseline Gemini: A Family of Highly Capable Multimodal Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:49.833777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:49.833777Z digest=sha256:dca1cbdd9bbc8cfbc99481dcf06bd3e99b7e4b3fdbd57bfbfbefd31291aa8372

Observation c10971a2-1dd0-4e53-b562-bd4a54572573 · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,.

On-Policy RL with Optimal Reward Baseline Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:49.876159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:49.876159Z digest=sha256:491a8e772c2f69bde0328330cea5258af586c85e5c0d1268c745bfa009ca79ce

Observation afc7fb7b-9b0f-46ce-9d99-f4240d1b0608 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

On-Policy RL with Optimal Reward Baseline Understanding R1-Zero-Like Training: A Critical Perspective

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:50.225237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:50.225237Z digest=sha256:89131b206107b9c22b45f9a03e366f43c5e4784a5d0f72432178babcbf80c730

Observation 370b104d-8aa6-45cb-a7fc-fc06aed6c4b3 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

On-Policy RL with Optimal Reward Baseline REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:50.044130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:50.044130Z digest=sha256:5da237a286f18bb2ce80e8b0a8321e1ad970eb6d964cf1858c068a89b3ae77ae

Pith citing papers

Observation 2734aa74-a458-4f4d-93c9-0945d4426584 · inbound

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities cites this paper.

UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities On-Policy RL with Optimal Reward Baseline

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:07:10.581791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:07:10.581791Z digest=sha256:f5f3c892c0b29aa24e7595c3770970583323c682be64ce6ec1f36ee9665fa5de

Observation 8793b3f4-1be9-4745-a4f2-e769ce5826b0 · inbound

Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective cites this paper.

Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective On-Policy RL with Optimal Reward Baseline

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:41:03.216079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T07:37:12.501489Z digest=sha256:a055c4bc33291bac7072fffcc3ea3e713274a659e1a777202d35255e9182e7fa

Observation 9db8ea4c-ef83-401e-a53e-06dc3fb98d35 · inbound

Policy Improvement Reinforcement Learning cites this paper.

Policy Improvement Reinforcement Learning On-Policy RL with Optimal Reward Baseline

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:48:22.858481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T22:47:05.132020Z digest=sha256:d2be808ac3ab40c5da456c7d87033774201dde6e44ea83ddb9a6359fd8bd7947

Observation 0d264321-b1b6-4599-8e4a-c783024c61a8 · inbound

Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning cites this paper.

Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning On-Policy RL with Optimal Reward Baseline

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:27.278990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-07T07:00:32.206081Z digest=sha256:d99aa9474f44464f4baf4aba9a8fe9e603d35744d0ea1f0af8beb6207058a943

Observation 752d6683-b8cb-479f-b559-908239a3875c · inbound

Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning cites this paper.

Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning On-Policy RL with Optimal Reward Baseline

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:49:14.929225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T23:47:53.282259Z digest=sha256:fe7f306cba1dcdc922d95b5d4420ca50258ea7864198956b338430afe8c2470e

Observation 32df8d68-545c-4bb8-aee5-2a2c4803d5fa · inbound

Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR cites this paper.

Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR On-Policy RL with Optimal Reward Baseline

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:10.479798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-09T15:39:54.001327Z digest=sha256:2f437d536153e90e2e33607263ffe7a7750c028768f002487007bc8d0aa1daaf

Observation bb1413e4-7d89-463c-8777-db47a92cb050 · inbound

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization cites this paper.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization On-Policy RL with Optimal Reward Baseline

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:42:02.898305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:cfc710f2b81a0a95428fa489366c1bddb16e8bd823ed4454ecb5d5c32d6d0b6c

Observation 65121ee2-c018-4af9-8868-4743fd40aca8 · inbound

Holder Policy Optimisation cites this paper.

Holder Policy Optimisation On-Policy RL with Optimal Reward Baseline

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:12:22.799052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-13T06:08:28.855671Z digest=sha256:b7b58644a649fc9ebb86c7a362e69e9c20154cdb442408cc3351fe02a822f3c1

Observation 6d678c24-3e73-4b0a-95e6-a526d8b673f7 · inbound

Holder Policy Optimisation cites this paper.

Holder Policy Optimisation On-Policy RL with Optimal Reward Baseline

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:01:23.127974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:bc4259b40d13770515a47a6e0660a55160333cfd8a285dcd852060f36eb18908

Observation 7d1f1c99-71d2-4524-900f-75fea0fce1c7 · inbound

Keep Policy Gradient in Charge: Sibling-Guided Credit Distillation for Long-Horizon Tool-Use Agents cites this paper.

Keep Policy Gradient in Charge: Sibling-Guided Credit Distillation for Long-Horizon Tool-Use Agents On-Policy RL with Optimal Reward Baseline

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:07:48.294529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T10:28:29.586968Z digest=sha256:9e9705dbb66764642638a54955012cac979f9f8a5cc7e15c7ceefd058e23f9cd

Observation 7fb79ea2-5234-42bd-9488-99fc41b2e7a0 · inbound

Keep Policy Gradient in Charge: Sibling-Guided Credit Distillation for Long-Horizon Tool-Use Agents cites this paper.

Keep Policy Gradient in Charge: Sibling-Guided Credit Distillation for Long-Horizon Tool-Use Agents On-Policy RL with Optimal Reward Baseline

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:55:31.197198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-01T07:48:15.475103Z digest=sha256:308202940c6fc64ac47ab87f8b521efef98464214d1d61e26481dd4a4fa67002

Observation feac46a2-371a-4fc9-8797-50a800d304e6 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning On-Policy RL with Optimal Reward Baseline

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.760859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:02ee88068bfb55c5558d48dd246292187599a95cdc0845fe086c1bb19bfe3d2c