Pith. sign in

Paper Citation Record · LEDGER

Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2608.01418.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01418 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:16:33.590510Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact4
  • verified fuzzy2
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0bc010ef-3952-458c-97ad-d505ee5e4dc3 · outbound

This paper cites Efficient RL Training for LLMs with Experience Replay.

Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning Efficient RL Training for LLMs with Experience Replay

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T00:16:33.534858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:16:33.534858Z digest=sha256:f24d78c4f6cc4fb17e7aefb6ef14696bae9f7c31550d347c03d4665fae45a79f

Observation 54bc4f66-0d05-4dd2-92e8-85cace2ffe15 · outbound

This paper cites Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning.

Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-06T00:16:34.000797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:16:33.548470Z digest=sha256:39d3c900ca04c22cc1c25fc3e3fb842a3c46802b3278339611ad2b7ec85fa7f9

Observation 6e93afe7-d628-42b9-88b1-9e5f64e49a58 · outbound

This paper cites DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training.

Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T00:16:33.553112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:16:33.553112Z digest=sha256:8dcd198b3f1a7995ce8438eeab987e590c4138f5d4cc6599193aa8436f8a3687

Observation 3f9d414d-f5c0-47c8-a418-2ca9f08cbd35 · outbound

This paper cites A step back: Prefix importance ratio stabilizes policy optimization.arXiv preprint arXiv:2601.22718,.

Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning A step back: Prefix importance ratio stabilizes policy optimization.arXiv preprint arXiv:2601.22718,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T00:16:33.557592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:16:33.557592Z digest=sha256:22d0fc80e359a5da3cdce2e727f5471493ff747d513a6a9e7955493de247c6b0

Observation 4040a79f-ea87-4fa8-a154-8efd0e91c9fe · outbound

This paper cites Token-Level Policy Optimization: Linking Group-Level Rewards to Token-Level Aggregation via Sequence-Level Likelihood.

Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning Token-Level Policy Optimization: Linking Group-Level Rewards to Token-Level Aggregation via Sequence-Level Likelihood

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-06T00:16:33.864755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:16:33.561853Z digest=sha256:dadb5fdbc8a827b984b475d15705ba8ac239e4180228c422b0cdaba3aaec7dbb

Observation af6dbf80-5fb8-4ea1-acdf-ba5e722f0163 · outbound

This paper cites Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective.

Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-06T00:16:33.645888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:16:33.577721Z digest=sha256:1b4f425e924cefc9548338961665d368e596bca280a58aaac176bb6b6ca0710c

Observation 89dc932e-66f2-4cf1-a2f5-cd6350771e67 · outbound

This paper cites Group Sequence Policy Optimization.

Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning Group Sequence Policy Optimization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T00:16:33.581880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:16:33.581880Z digest=sha256:4d7a5dc7e44e328baf7abaedb01bf7a5a1377ff8e080dbb6380043fd1cec5c3f

Observation 0b56770f-0b21-449f-b0ff-bf889c5f8085 · outbound

This paper cites LX t=1 CtAβ t tX k=1 zk # = LX k=1 Eπβ.

Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning LX t=1 CtAβ t tX k=1 zk # = LX k=1 Eπβ

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:16:34.138954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:16:33.585855Z digest=sha256:3ce016dbb9017b2e46dfce15ca614948f12510e3050944e07c8ba4b62dfef1c1

Observation e90a00ca-d151-4699-b0d1-7924275ab189 · outbound

This paper cites LX t=k Aβ t (st, at) sk, ak # =E πθ.

Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning LX t=k Aβ t (st, at) sk, ak # =E πθ

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:16:34.124938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:16:33.590510Z digest=sha256:4f5ce50036424c1347e53343ba973ddf67b3696cbcc77041a833fbac67fb7cd3

Observation e3301e55-5ea8-47b7-96d1-c6c14728768f · outbound

This paper cites LLMs can learn to reason via off-policy RL.arXiv preprint arXiv:2602.19362,.

Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning LLMs can learn to reason via off-policy RL.arXiv preprint arXiv:2602.19362,

Reference 2000

Resolution
unresolved
no resolver link, observed 2026-08-06T00:16:33.565799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:16:33.565799Z digest=sha256:53e9c0a49107f3000b2c9ea44f216aaba6b41328cd8395c9f00c6380adcda573

Observation f872b028-9c2e-4057-b909-59df10be4880 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-06T00:16:33.569569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:16:33.569569Z digest=sha256:12861f6c7179da21609827f31a0797a3d9f4ed0a1b9536d1a714bc722b842d1d

Observation 0a330717-f8e5-4aed-824f-a0c464e7ad91 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T00:16:33.573986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:16:33.573986Z digest=sha256:47105a64122c08c95fad1de100fcc5559d425d72c0f2f2a56fa7a1ff37c7db55

Observation f9fca2eb-90a0-4632-8a2c-1b9a3f07ad54 · outbound

This paper cites 9 Homayoun Honari, Roger Creus Castanyer, Michael Przystupa, Michael Noukhovitch, Pablo Samuel Castro, and Glen Berseth.

Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning 9 Homayoun Honari, Roger Creus Castanyer, Michael Przystupa, Michael Noukhovitch, Pablo Samuel Castro, and Glen Berseth

Reference 2025

Resolution
verified exact
raw_fallback, observed 2026-08-06T00:16:34.081238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:16:33.544313Z digest=sha256:35a9dbd902d237e005a10fe261947de90c13c665e48555f8d56d7644c8226783

Observation 626811b2-1f25-40a2-801e-a98222fc130d · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-06T00:16:33.539499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:16:33.539499Z digest=sha256:a7fd546201a8698cd6253d8cc2d71315c8989e101ed2f926f0cd6cfc4ac5b0d2

Pith citing papers

No inbound Pith citation observations are available.