Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:46:50.980711Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 12 inbound Pith citation observations for arXiv:2505.23585.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:46:50.980711Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T14:07:10.581791Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T07:59:40.759572Z
15 of 15 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2dcd3541-28e0-4613-beeb-c5457c7d466b · outbound
On-Policy RL with Optimal Reward Baseline DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e642235-367e-4ce6-96e0-6e50fc2734c9 · outbound
On-Policy RL with Optimal Reward Baseline Approximately optimal approximate reinforcement learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 19022d9d-dd2a-46a8-b3ea-bad1333745d1 · outbound
On-Policy RL with Optimal Reward Baseline DeepSeek-V3 Technical Report
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bb3f326-21a2-4b99-bbd4-319545fc2b19 · outbound
On-Policy RL with Optimal Reward Baseline American invitational mathematics examination - aime
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d49e6d92-b1f3-46eb-8e7b-fd09ebf946af · outbound
On-Policy RL with Optimal Reward Baseline GPT-4 Technical Report
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf99ce1d-558d-41a9-be08-e083b247bb7c · outbound
On-Policy RL with Optimal Reward Baseline HybridFlow: A Flexible and Efficient RLHF Framework
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5fc5994-bdba-4e76-ba57-475127ed913e · outbound
On-Policy RL with Optimal Reward Baseline Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f4e0d30-1fee-407c-b9f1-0db0ecc45304 · outbound
On-Policy RL with Optimal Reward Baseline Qwen3 Technical Report
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24f923a4-be90-4d91-9fbb-4fed684d2973 · outbound
On-Policy RL with Optimal Reward Baseline Neural text generation with unlikelihood training
Reference 1992
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c3930687-7926-4244-a3bb-9abbaa758984 · outbound
On-Policy RL with Optimal Reward Baseline DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a37069ee-a91e-4bf0-8578-d051882fec5f · outbound
On-Policy RL with Optimal Reward Baseline Proximal Policy Optimization Algorithms
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 956640e0-8b55-45e4-96bf-058de5bdc983 · outbound
On-Policy RL with Optimal Reward Baseline Gemini: A Family of Highly Capable Multimodal Models
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c10971a2-1dd0-4e53-b562-bd4a54572573 · outbound
On-Policy RL with Optimal Reward Baseline Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afc7fb7b-9b0f-46ce-9d99-f4240d1b0608 · outbound
On-Policy RL with Optimal Reward Baseline Understanding R1-Zero-Like Training: A Critical Perspective
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 370b104d-8aa6-45cb-a7fc-fc06aed6c4b3 · outbound
On-Policy RL with Optimal Reward Baseline REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2734aa74-a458-4f4d-93c9-0945d4426584 · inbound
UloRL:An Ultra-Long Output Reinforcement Learning Approach for Advancing Large Language Models' Reasoning Abilities On-Policy RL with Optimal Reward Baseline
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8793b3f4-1be9-4745-a4f2-e769ce5826b0 · inbound
Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective On-Policy RL with Optimal Reward Baseline
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9db8ea4c-ef83-401e-a53e-06dc3fb98d35 · inbound
Policy Improvement Reinforcement Learning On-Policy RL with Optimal Reward Baseline
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0d264321-b1b6-4599-8e4a-c783024c61a8 · inbound
Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning On-Policy RL with Optimal Reward Baseline
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 752d6683-b8cb-479f-b559-908239a3875c · inbound
Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning On-Policy RL with Optimal Reward Baseline
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 32df8d68-545c-4bb8-aee5-2a2c4803d5fa · inbound
Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR On-Policy RL with Optimal Reward Baseline
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bb1413e4-7d89-463c-8777-db47a92cb050 · inbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization On-Policy RL with Optimal Reward Baseline
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 65121ee2-c018-4af9-8868-4743fd40aca8 · inbound
Holder Policy Optimisation On-Policy RL with Optimal Reward Baseline
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6d678c24-3e73-4b0a-95e6-a526d8b673f7 · inbound
Holder Policy Optimisation On-Policy RL with Optimal Reward Baseline
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7d1f1c99-71d2-4524-900f-75fea0fce1c7 · inbound
Keep Policy Gradient in Charge: Sibling-Guided Credit Distillation for Long-Horizon Tool-Use Agents On-Policy RL with Optimal Reward Baseline
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7fb79ea2-5234-42bd-9488-99fc41b2e7a0 · inbound
Keep Policy Gradient in Charge: Sibling-Guided Credit Distillation for Long-Horizon Tool-Use Agents On-Policy RL with Optimal Reward Baseline
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation feac46a2-371a-4fc9-8797-50a800d304e6 · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning On-Policy RL with Optimal Reward Baseline
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.