Pith. sign in

Paper Citation Record · LEDGER

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction

As of 17 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2605.12070.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.12070 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T05:57:49.286939Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T08:56:36.611575Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact26
  • verified fuzzy13
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 4a5e0c94-92bd-403e-8da9-d4993c90e2a6 · outbound

This paper cites Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:42:33.922725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:fcf808a5c1f8241b610c6ccf8a07bb744bb6741a8438d204ca8191fcfde76302

Observation 86d40475-e663-4cba-b2b3-22c2e2a98b27 · outbound

This paper cites Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:42:33.920908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:e0d5ac389c80939d2ef396ab35913abdee237e368a2308c982dcf6f70031f554

Observation 58d05ae4-e8f2-4793-8c7d-7813d40ae441 · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:23.557756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:174626c8782da8ca042807f045379c8d470f35a6450f36b9ffcc4bc59af324d2

Observation 243d28b7-4cab-4dcb-aa7b-30b7bef80327 · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:23.573116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:e289ce1a3018a54c4849bfbeccab85db1d20b0073d9680d1e92dda1873dcc834

Observation d4f29b51-a9d4-4994-8318-68e2048ae54d · outbound

This paper cites Agentic Reinforced Policy Optimization.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Agentic Reinforced Policy Optimization

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:57:12.173630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:2b7b4b45de64f2e40a950ca3818507776114474b4aff6a10ff13f55e260a1a64

Observation d6efb9a3-b5a1-4520-8e78-302bf9fe19b1 · outbound

This paper cites Areal: A large-scale asynchronous reinforcement learning system for language reasoning.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Areal: A large-scale asynchronous reinforcement learning system for language reasoning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:42:33.918869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:dcadfb4846ce2cd6f7aed0b91629b1693aa51f83e020330a00a376904fb8b742

Observation 80e444e8-3257-47cd-9d51-1af09203b22a · outbound

This paper cites RL-VLA$^3$: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction RL-VLA$^3$: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:23.570402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:c0169c0ee15c7c2c3310aa8262520ad3b592d38651cd39dbb3980f29a54dd82b

Observation 23f2fcd8-c8fc-4b4e-bc46-400f0d170df4 · outbound

This paper cites Vitabench: Benchmarking llm agents with versatile interactive tasks in real-world applications.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Vitabench: Benchmarking llm agents with versatile interactive tasks in real-world applications

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.564409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:8366a5ded1dfb29f0b54ac575e0085151fdc27afd6097da966f0ba3142b798c7

Observation a77987df-108c-4545-9349-22f3e287bc09 · outbound

This paper cites Batch size-invariance for policy optimization.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Batch size-invariance for policy optimization

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:42:33.924421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:1ab7b4ea6592773d0153d9a8c464edcc6e310053342eada3bde0ac13cb7d86d2

Observation 6083ae74-36f4-4064-aea4-5312a80d8551 · outbound

This paper cites Stable asynchrony: Variance-controlled off-policy rl for llms.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Stable asynchrony: Variance-controlled off-policy rl for llms

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.561246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:3ecce814d2540808b5f262634fbb6b79b021ab37f45de1b5025f89f6a59a4ee4

Observation b3444ef2-1355-4175-a5ed-9daa6d8c2ab2 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Efficient memory management for large language model serving with pagedattention

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:42:33.902836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:9449e251558447ed22d63668a7c5996a89c015539a42d636c68d87a2b12176db

Observation 9934be26-66fd-4b6b-9e49-ce933009aa73 · outbound

This paper cites A-3PO: Accelerating Asynchronous LLM Training with Staleness-aware Proximal Policy Approximation.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction A-3PO: Accelerating Asynchronous LLM Training with Staleness-aware Proximal Policy Approximation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-08-13T02:25:25.422196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:2b8015d1b98873211685f8902404a7a7ab473f68bb392eb638857f12e94349e4

Observation b829d819-d35e-4b0c-9cd7-5b8de5cde411 · outbound

This paper cites When speed kills stability: Demystifying RL collapse from the training-inference mismatch.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction When speed kills stability: Demystifying RL collapse from the training-inference mismatch

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:42:33.900937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:878b0f5135b0a45399c159cb27e3ddff2b47a382958cad5015bc5588f934741f

Observation 736074fc-6e75-4b52-a98a-b7676343e11d · outbound

This paper cites Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.522294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:784b872da4034677877fdcfe7151f474b15315682e617e71a79128ab1f3a0d85

Observation b2d56bc3-d8b7-4202-9c32-dc064c6545e0 · outbound

This paper cites Rethinking the Trust Region in LLM Reinforcement Learning.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Rethinking the Trust Region in LLM Reinforcement Learning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-27T02:05:11.603431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:cd9eb4d1ec743197cd83d259bba328ee727bd763d2d01ae14c8f74b7adac6fce

Observation 80cabeb2-26d8-4292-8949-9e574954a27f · outbound

This paper cites Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.512700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:5bbd5b4632ebc1489e27d34d60a0495b43a4ebd5fe2d08376b199dba7dcaf533

Observation da8b7d41-10d5-40c8-9ec7-8343a3ae2312 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Proximal Policy Optimization Algorithms

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:23.502320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:597a511e1f85b5aecc98bedec0f7adf9b7f6e53c78371af04c9748c151f64b68

Observation c1cfbc40-68d0-45cd-9e44-0d264e6ef316 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:23.505475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:b1584ace0ca3078ebe8739cab6a2715337e62a80897b67849e81f5fbff85ce17

Observation a3b1edf2-5521-487b-a66f-f39d305372c9 · outbound

This paper cites VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:23.516131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:1ea8b9ab06ea283c4456a3800f704471dadfad2c27bb91e46043125467cfc0f7

Observation 22f492b9-3b1c-4b79-992f-c61670db853c · outbound

This paper cites Laminar: A scalable asyn- chronous rl post-training framework.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Laminar: A scalable asyn- chronous rl post-training framework

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.508737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:f532b01e077c54b4fba45ce73741fa2d06a456bf06195aa5ffbc98734338b016

Observation fd249724-412f-4d96-8d1a-171b34535e21 · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Hybridflow: A flexible and efficient rlhf framework

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:42:33.915476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:391882ee3048fd14f67ac080dc0188c37adb6ce8cdf272e3431b214ed541ce2d

Observation 63cb8be7-490c-480a-9b90-15f5bd778d5b · outbound

This paper cites Klear-reasoner: Advancing reasoning capability via gradient-preserving clipping policy optimization.arXiv preprint arXiv:2508.07629.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Klear-reasoner: Advancing reasoning capability via gradient-preserving clipping policy optimization.arXiv preprint arXiv:2508.07629

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.549528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:7d304d172375e8ac310ec13d697417cd89540b97f0780a2190d75de384b24f7a

Observation b20ee37d-8e70-47fa-b22a-ab0bdb940594 · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Kimi K2.5: Visual Agentic Intelligence

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:23.540340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:84b51062d63c873057dc0255ce77a43006e843ef0ee05165389a9d41ebb68236

Observation 4dab0159-f0c5-4a6b-b257-86db61318b35 · outbound

This paper cites Every step evolves: Scaling reinforcement learning for trillion-scale thinking model.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Every step evolves: Scaling reinforcement learning for trillion-scale thinking model

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.546490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:f7139490cece4ec3d9a8da004e8f52fb5fd25a1a1316097ccecc89e64a508c9e

Observation 3665664e-ce1f-4476-8860-bfcb504be7e3 · outbound

This paper cites Ernie 5.0 technical report.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Ernie 5.0 technical report

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.543107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:51207bdb85a34e5bc65f1219e566b50063c8168720e50e04382036a5913e18af

Observation f3660f3f-d005-46a3-a023-10a31f34b7e2 · outbound

This paper cites Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.534275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:4c0907288a6e70507e8af4a4be82141302329565217898b525ac696909afac75

Observation 1b87c563-85bb-43ec-9f24-4f665cfaa99a · outbound

This paper cites Let it flow: Agentic crafting on rock and roll, building the rome model within an open agentic learning ecosystem.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Let it flow: Agentic crafting on rock and roll, building the rome model within an open agentic learning ecosystem

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.537635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:bd85e67b923e3284ed7cb299f585790d3dfcd533c8fbd1b7bb9da118c216b867

Observation b029efc2-fed6-4bf5-823d-cfee5f79ba8a · outbound

This paper cites MiMo-V2-Flash Technical Report.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction MiMo-V2-Flash Technical Report

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:23.524895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:0f798a8ed9162ba718151368ed21909c843208df105c68cad20531694594a238

Observation 07b90231-bb85-44e8-94df-75d358c4b7cb · outbound

This paper cites Your efficient rl framework secretly brings you off-policy rl training, August 2025.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Your efficient rl framework secretly brings you off-policy rl training, August 2025

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:42:33.910212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:ccb3e230f7572fb5990ebef2478e123c1bb0287a8b78b0470296acdf852a3f92

Observation 32303e9a-21cb-447b-b9b0-b2788502cca5 · outbound

This paper cites Your efficient rl framework secretly brings you off-policy rl training, august 2025.URL https://fengyao.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Your efficient rl framework secretly brings you off-policy rl training, august 2025.URL https://fengyao

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:42:33.917217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:af930ecf5c39fc50b8c00c74896f756d69cb5814d06f433b93194429e3718ed1

Observation 26ad966d-6e42-4bc6-b296-acd744d36442 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:23.530861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:b73253d307fbb1fc9913bfcf7606401493dbd66d2afab964ef651ca6b5e9e2c1

Observation 141c3492-816f-4804-b0a7-14d2c7e4c070 · outbound

This paper cites GLM-5: from Vibe Coding to Agentic Engineering.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction GLM-5: from Vibe Coding to Agentic Engineering

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:23.554770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:bccad7372997da96553aa99e3dcb7fe56e976f8377b6c74ff1d575ad1f265c5f

Observation 99f9e4b1-c3e1-4ede-a81c-539796e9fa94 · outbound

This paper cites The landscape of agentic reinforcement learning for llms: A survey.Transactions on Machine Learning Research.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction The landscape of agentic reinforcement learning for llms: A survey.Transactions on Machine Learning Research

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:42:33.912014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:4b365ed6c64ba00333df26e8c50780b7dcf10e3394910edb475a64708ccca805

Observation 8b232ab9-5f0a-47da-a982-e7ab8a5792f7 · outbound

This paper cites Small leak can sink a great ship–boost rl training on moe with icepop!.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Small leak can sink a great ship–boost rl training on moe with icepop!

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:42:33.908495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:47afb979b69249ea703cfd299d7ec81c84b8813ea1a8e35911b38f1347398a4d

Observation c7831df4-df25-4f4c-a67b-36852ae4c24b · outbound

This paper cites Stabilizing reinforcement learning with llms: Formulation and practices.arXiv preprint arXiv:2512.01374, 2025a.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Stabilizing reinforcement learning with llms: Formulation and practices.arXiv preprint arXiv:2512.01374, 2025a

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.496418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:3bfed9e76d6d550cd8f1d7dce5ca7944ede179cc68683e927a93a0bf1a7b1411

Observation 9238bbf6-a52f-46fe-991e-3fdacb4541c9 · outbound

This paper cites Group Sequence Policy Optimization.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Group Sequence Policy Optimization

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:23.552263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:756a7ef92fafe85de92d6602cf13fbf3cf1303da2411b3e57b92bef7ba8916ec

Observation 806f8306-3f5d-4164-bd7e-59b441558a1c · outbound

This paper cites Prosperity before collapse: How far can off-policy rl reach with stale data on llms?.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Prosperity before collapse: How far can off-policy rl reach with stale data on llms?

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.499553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:c1634ed0023640b79a5cc3018a9d7b38ea4f4342c8adef548f940f762a3ebf02

Observation b9b6f318-4fad-404f-ac68-15a895208590 · outbound

This paper cites Sglang: Efficient execution of structured language model programs.Advances in neural information processing systems, 37:62557–62583.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Sglang: Efficient execution of structured language model programs.Advances in neural information processing systems, 37:62557–62583

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:42:33.906563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:a3035033e4ed093ed414a0f90e3ab7a5c899206b4b728dfd5e1e8a4cc633d3e4

Observation 5a4380e2-768b-4371-8664-b6fae4aaf663 · outbound

This paper cites Limitations.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Limitations

Reference 39

Resolution
malformed identifier
raw_fallback, observed 2026-05-13T09:42:33.904611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:24a6bc5c1df41276276612991ef53be26ade857cbfcc931d7e8e863887670824

Observation 66624dd4-4111-4234-8af2-517fe92cdfea · outbound

This paper cites Therefore IRB approval is not applicable.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Therefore IRB approval is not applicable

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:42:33.913667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:b42ad683f02cc8b96bc27bc3468381d16cf92cfcd4ffa9f4b571f67db4d88b27

Pith citing papers

Observation f891b82f-e12e-4294-8fd7-2d1871075466 · inbound

From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation cites this paper.

From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-02T08:58:21.176076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-02T08:56:36.611575Z digest=sha256:978699a732349c34993513be5d899b209eb93f9da753109b91be0c046448652b