Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T00:16:33.590510Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2608.01418.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T00:16:33.590510Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
14 of 14 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0bc010ef-3952-458c-97ad-d505ee5e4dc3 · outbound
Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning Efficient RL Training for LLMs with Experience Replay
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54bc4f66-0d05-4dd2-92e8-85cace2ffe15 · outbound
Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6e93afe7-d628-42b9-88b1-9e5f64e49a58 · outbound
Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f9d414d-f5c0-47c8-a418-2ca9f08cbd35 · outbound
Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning A step back: Prefix importance ratio stabilizes policy optimization.arXiv preprint arXiv:2601.22718,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4040a79f-ea87-4fa8-a154-8efd0e91c9fe · outbound
Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning Token-Level Policy Optimization: Linking Group-Level Rewards to Token-Level Aggregation via Sequence-Level Likelihood
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation af6dbf80-5fb8-4ea1-acdf-ba5e722f0163 · outbound
Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 89dc932e-66f2-4cf1-a2f5-cd6350771e67 · outbound
Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning Group Sequence Policy Optimization
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b56770f-0b21-449f-b0ff-bf889c5f8085 · outbound
Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning LX t=1 CtAβ t tX k=1 zk # = LX k=1 Eπβ
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e90a00ca-d151-4699-b0d1-7924275ab189 · outbound
Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning LX t=k Aβ t (st, at) sk, ak # =E πθ
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e3301e55-5ea8-47b7-96d1-c6c14728768f · outbound
Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning LLMs can learn to reason via off-policy RL.arXiv preprint arXiv:2602.19362,
Reference 2000
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f872b028-9c2e-4057-b909-59df10be4880 · outbound
Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a330717-f8e5-4aed-824f-a0c464e7ad91 · outbound
Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9fca2eb-90a0-4632-8a2c-1b9a3f07ad54 · outbound
Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning 9 Homayoun Honari, Roger Creus Castanyer, Michael Przystupa, Michael Noukhovitch, Pablo Samuel Castro, and Glen Berseth
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 626811b2-1f25-40a2-801e-a98222fc130d · outbound
Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.