Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T10:01:26.786756Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2601.11960.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T10:01:26.786756Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation de93cd5b-a255-4ae4-b4af-eb7e298cce7c · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and R \' e mi Munos
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df6e4778-5269-45c0-88f6-44e35b543106 · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Reasoning with Exploration: An Entropy Perspective
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83810eff-5631-467e-9a1c-b3d5faa660be · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Christiano, Jan Leike, Tom B
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4f750a5-f8d4-4b08-a41a-cae44bc892b2 · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Training Verifiers to Solve Math Word Problems
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a4df2cc-5786-4c62-8279-14453832d4d8 · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08bdac19-a887-43f9-8756-62dddc72b42e · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe7394cb-3bf2-42cc-95bf-a360fdece666 · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3718b81-1ec8-407e-83b2-b33a2b6ebd14 · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning DeepSeek-V3 Technical Report
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4359e8d4-22d2-4f64-ae2b-4eb44359f6ea · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ddb480e-8dd8-4ffb-9799-f6aab7625168 · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 896f061b-00f4-4a0e-95e1-d87943c4a162 · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecc0eea7-c38b-4c39-be4b-8c2c7c7efa47 · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cefa712-b75c-4fbc-ab67-0b56b131c00d · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77d9dbf8-a841-4e7d-ba04-9821845fcff6 · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e5abbef-0fb4-472b-920e-6c71c43a1b5f · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26f60728-75b5-4ae3-aa3c-b533ba7f0d97 · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e03b6bc5-cfe7-43a8-b400-aeb42e093ad6 · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31f5a61a-9a6b-4c37-9878-502cfd63b686 · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Park, Junsu Kim, Gyeongman Kim, Jinyoung Jo, Sean Choi, Jaewoong Cho, and Ernest K
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 072e22b8-610e-4e70-9e52-1140402e5bdd · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Qwen2.5 Technical Report
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73f7241f-5bfc-422f-b655-5e7e3e9525d8 · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Proximal Policy Optimization Algorithms
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e73d007-56cc-40a1-ab2c-bab8052de6f1 · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 123152b8-da22-41d5-a73c-f9116eb03e7f · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f259c5e4-0b05-431d-b06d-860e347777f2 · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 507e3a6b-9f24-490b-b33b-6d11de912508 · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Sutton and Andrew G
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e54eb95-f327-4220-9e32-1e2f4eafbb16 · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03613c47-f757-4d02-85a3-20eb907510f6 · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31b0c102-880e-4755-bbaa-8d84bca1bb8a · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Entropy-Regularized Token-Level Policy Optimization for Language Agent Reinforcement
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd69d07a-9e81-4deb-8017-5450a211bc63 · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68a64716-ac48-493d-98da-ceb42c2f96f6 · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1989704-4b62-4235-ab9d-3fb24ae7f43d · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 924bd161-e955-4f97-9a89-d39d01e7d3bb · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef88384c-f0aa-4c12-8df2-2224cc34eff7 · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Qwen3 Technical Report
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb28cef4-4463-444b-999f-a9bd72234041 · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b43a749-ca41-4c1a-841d-7204f0597a61 · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69cc710c-9a75-43bd-8724-32f641e3ce03 · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning online" 'onlinestring :=
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe36e1f0-98b3-4da3-a948-14f4ff5733e9 · outbound
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning write newline
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.