Pith. sign in

Paper Citation Record · LEDGER

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

As of 9 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 4 inbound Pith citation observations for arXiv:2607.07508.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.07508 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-09T08:51:10.098370Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:16:33.548470Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T00:16:33.996094Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact18
  • verified fuzzy1
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 533f0531-72c7-4a80-be28-209459ab1997 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.437175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:68ca875c26562f81e80873db6fe0dabe7d3bd9e587bd1634027e50d5ae14f1ab

Observation 2d71f88d-b416-485d-b658-4db627090fae · outbound

This paper cites MathArena: Evaluating LLMs on Uncontaminated Math Competitions.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning MathArena: Evaluating LLMs on Uncontaminated Math Competitions

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.422874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:240febb7d7df99e99fa5145292dfe57a2129b8f098d563b3eff233f9c110bf4b

Observation b1535a6b-7bb4-4c04-9275-05960d480bab · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Training Verifiers to Solve Math Word Problems

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.425617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:8e9402b091dbc93be7d197682acf4d3b3171fe8db37442f3ee4b78582e5a6538

Observation 92f02aa6-2291-45cd-95a5-382a9e360cfa · outbound

This paper cites AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.428422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:8ca45f9f9f0038e1081714e6a15820831ecc54cf5a86a648c972a25b5a5d81e2

Observation 62ee85c2-6cbe-4527-a66b-b2854a66f22e · outbound

This paper cites GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.399622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:3c9d8dc3b37d0e0b05abd904e0ea4c469bedca861e62eea7b019274b1669fe5c

Observation 3a1b20da-e3a2-4b93-8472-8b2f65e442fb · outbound

This paper cites Acme: A Research Framework for Distributed Reinforcement Learning.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Acme: A Research Framework for Distributed Reinforcement Learning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.431374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:1f35a0beb3e4215e99662e3632d9a90842f778bd0d25736cf44c4ff19e10beea

Observation cb616be2-e8c4-4786-b318-de0a9c8e7843 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.405180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:47bfb38ed12e612b17f7bb7a8bcc21e557006d33e66ffe0d0b04bce0fe1d7bde

Observation b458b75f-a787-4470-b38d-a48f0b94318b · outbound

This paper cites Let's Verify Step by Step.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Let's Verify Step by Step

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.446011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:fef9c01690714595f7537ef603e588a5c001dfe5f86ace9eb7159c201e047fcc

Observation de2b4abc-161c-4bc5-8aa5-27fba8e177cd · outbound

This paper cites Part ii: Roll flash–accelerating rlvr and agentic training with asynchrony.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Part ii: Roll flash–accelerating rlvr and agentic training with asynchrony

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-09T08:56:06.419907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:2e31d7710557486bd7f34c50b9b519fcb6c612623107e48c0b661b7635e9ab7c

Observation fd9748e5-9bdf-404b-b861-64a172b03968 · outbound

This paper cites an unresolved cited work.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-07-09T08:56:06.846916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:e18fbbed4a7a2d728ed4074af53d7cdcba78deb59c8367deec1d24688c93ff15

Observation 3f38ed09-ba6e-4358-b6de-8abe62e05649 · outbound

This paper cites V olodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning V olodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T08:56:06.848602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:240315865db50e252e0e43579fefba2627608b64317937b9d8be43ce35c35043

Observation 868684ee-bdfe-4db4-b959-e46c40df7f67 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning WebGPT: Browser-assisted question-answering with human feedback

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.448589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:f343eba2fddb2540f35057f8a3e6e6ea1b1dcbb0c96d95cf4beb669195b4b8a9

Observation cb2a4a22-1319-451c-813d-68375c26c8ff · outbound

This paper cites Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.402413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:2c8b638572b10c493609dc99c5bbe48e4c2b94d76505f2eedd4d3e456a7b0baa

Observation b6880e6f-ff2c-4e2b-9cc3-b29ec19eb916 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.416635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:ddce4d302933b28c989c08fad0dc6534f8ba0f06f06a421a368460c47aa6a310

Observation f0b09ed3-9e6a-430e-bab5-aed6b585c98f · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.439895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:839b445dc6e9ae754abaa6e3f9f2da1bd21f144234390f57052f8849502bea65

Observation 4c3d12bc-f903-4de7-8813-003e0a82207d · outbound

This paper cites Every step evolves: Scaling reinforcement learning for trillion-scale thinking model.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Every step evolves: Scaling reinforcement learning for trillion-scale thinking model

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-09T08:56:06.443232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:a8d878db2e3f4e8d74304548262d602ae173333255660f5d3a275886a0bfeee5

Observation 1d802991-aa15-4f98-a56c-7d128d623fa6 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.408245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:ca63feb6a95569ad936f46f4b917b790af04447e89e0c79577e89423d1909288

Observation 1627bf83-e4d9-4f99-87bf-caf7d7137fad · outbound

This paper cites Mobilerl: Online agentic reinforcement learning for mobile gui agents, 2025 a.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Mobilerl: Online agentic reinforcement learning for mobile gui agents, 2025 a

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T08:56:06.451773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:5c6d7b43b51c1bef343f03fdc2f4fd37390a4edf3b8d9116aac3b42b934981a2

Observation eb48c571-9088-411c-987c-19551079569c · outbound

This paper cites Qwen3 Technical Report.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Qwen3 Technical Report

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.411146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:8285ff80a1c033a283195768b4b16c836eb4e2c8b2a6d10dd56be83c60fdcbad

Observation f2edcb82-d888-4486-af58-067824bc5ac7 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.413615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:79c18726d0d6c99711f94fcb3f31b1842af159c66a2eaad37deb9e49fe5a93db

Observation 4e0b0ac7-cdff-448e-a7e1-ef20f930a758 · outbound

This paper cites Group Sequence Policy Optimization.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Group Sequence Policy Optimization

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.434205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:066a086fb26cb2ded837514d4afc5e363a05a3f7cc61ed280cb85d119b172ade

Pith citing papers

Observation 884b4557-7dad-476e-b828-55a4aa438cee · inbound

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning cites this paper.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:32.203143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:32.203143Z digest=sha256:1303971a0dfbdb105015f400fddb2bf6814ec50098f31229aa334f029582c302

Observation cc26dc60-50b0-4b78-87d0-2c2eac0d56a0 · inbound

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications cites this paper.

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

Reference 106

Resolution
unresolved
no resolver link, observed 2026-08-01T03:38:25.886104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:38:25.886104Z digest=sha256:88abdd5056b1dfe29cc5e62eb617e71e7ed0a7e46479cecbb3cc113e26d6a69b

Observation c3e750cc-d962-4755-b375-d5cd91e43c79 · inbound

Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents cites this paper.

Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

Reference 147

Resolution
unresolved
no resolver link, observed 2026-07-31T14:04:46.052248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T14:04:46.052248Z digest=sha256:8dea9d94f52c3c8bbf8b094bfac31fe8d22647078c136f4e652b8f5847827080

Observation 54bc4f66-0d05-4dd2-92e8-85cace2ffe15 · inbound

Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning cites this paper.

Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-06T00:16:34.000797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:16:33.548470Z digest=sha256:92df31f74908e67f153808aebf774512bc2feb917356157eb80b0fa340775563