Pith. sign in

Paper Citation Record · LEDGER

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning

As of 18 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2607.18722.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.18722 v3

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T14:36:34.322864Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e6413e05-596f-4d37-be35-57411ae27a57 · outbound

This paper cites Staleness in fully asynchronous RL.https://appliedcompute.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Staleness in fully asynchronous RL.https://appliedcompute

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:31.831428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:31.831428Z digest=sha256:fa4841d72a47920f369f8eab63a392c6fc59d8eb11a5cb4fcc4c315f31b55bbb

Observation b6cb9c57-9e63-4c53-bf68-528e5929e1b8 · outbound

This paper cites Olmo 3.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Olmo 3

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:31.937887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:31.937887Z digest=sha256:4a822e796d4f96a1b9dca6032ffa20ce398473d94699a9e234991dc718084e33

Observation b8b5cc90-592b-426c-b6ce-460b27712920 · outbound

This paper cites Kpop: Taming training–inference mismatch in reinforcement learning with adaptive masking regions, May 2026.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Kpop: Taming training–inference mismatch in reinforcement learning with adaptive masking regions, May 2026

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:32.058621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:32.058621Z digest=sha256:b705cb0da2304e86ee70701c3130eca5a4f99603f6bbf97c0e84dc2dda9dfabe

Observation 884b4557-7dad-476e-b828-55a4aa438cee · outbound

This paper cites Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:32.203143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:32.203143Z digest=sha256:e23c54b27fe90f04770a67af3d1ebbf535a0058a129e3dcaa722c74b9f32fcf0

Observation 64398b4c-f0f4-4f83-89fe-d0e1a2fdda98 · outbound

This paper cites Stabilizing rlvr via token-level gradient diagnosis and layerwise clipping, 2026.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Stabilizing rlvr via token-level gradient diagnosis and layerwise clipping, 2026

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:32.304147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:32.304147Z digest=sha256:f55cfdf011c32bee203405f05eefc61117427200ae2004e38a7682d70e5dff1d

Observation a46f3cb1-3564-4238-98fa-791a20249e55 · outbound

This paper cites Approximately optimal approximate reinforcement learning.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Approximately optimal approximate reinforcement learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:32.449198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:32.449198Z digest=sha256:cb32128636c81b43a26696278fd39a76939b58fa10b8cab59039f746b3cd5d20

Observation 2cae68bc-4ccf-4684-b76f-fa09566df595 · outbound

This paper cites Stabilizing MoE reinforcement learning by aligning training and inference routers.arXiv preprint arXiv:2510.11370, 2025.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Stabilizing MoE reinforcement learning by aligning training and inference routers.arXiv preprint arXiv:2510.11370, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:32.634733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:32.634733Z digest=sha256:bf5f24e306455c98e459d5d634ada4249c7e7006766ec1d71089a70c0e687da3

Observation 3c811da1-d109-4b48-9278-e72bb60c6bfb · outbound

This paper cites Rethinking the trust region in LLM reinforcement learning.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Rethinking the trust region in LLM reinforcement learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:32.845639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:32.845639Z digest=sha256:d9e4f9372965bc7719b1b333de58678c30b87e2683e130928f6015bb5a509f8a

Observation 738ca189-9408-4045-8f74-1be796585e9f · outbound

This paper cites Trust Region Policy Optimization.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Trust Region Policy Optimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:32.958394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:32.958394Z digest=sha256:dac1ca31c12da0d40d51210f476735dd5fa8f4a29f4d9a52495239449d61b3f7

Observation beec5b53-d5e2-4e73-ba9e-d12b235f85f8 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:33.084469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:33.084469Z digest=sha256:2221b00679784b2a6a5e88bca9401901b98614f02e53f410cf133e34c82e9c1d

Observation 0d9e1a7f-c816-4895-ab79-cb4bbabb0842 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:33.192530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:33.192530Z digest=sha256:e871d2f9f4f610f1eb460e708f0ca3175f045979cf94087c34ecb85125bd0c08

Observation be9955a2-bf4f-4a54-a983-4bc8373cd7d4 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:33.391991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:33.391991Z digest=sha256:5f7824d33d8922a85a4eabac02bb3fc91059d8d07a72cb56ae5d2a91307d1454

Observation 5d1bd3b8-31c6-4612-af43-5d8203e0db37 · outbound

This paper cites Qwen3 Technical Report.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Qwen3 Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:33.544666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:33.544666Z digest=sha256:dbaab0dc099fa75d027ac54e2c985ac8b6998ff8e501e4602a6cea8db477952e

Observation b288e0a7-fbd5-4981-9ec2-f356ff9bcba6 · outbound

This paper cites Rethinking the Divergence Regularization in LLM RL.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Rethinking the Divergence Regularization in LLM RL

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:33.642069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:33.642069Z digest=sha256:7ef73207619bad51607a82a5393017157efcd5612205da51376b0183b876c5cf

Observation 2da0c69c-9f37-48ac-9d82-5b7cfe2ab2e5 · outbound

This paper cites DAPO: An open-source LLM reinforcement learning system at scale.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning DAPO: An open-source LLM reinforcement learning system at scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:33.791105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:33.791105Z digest=sha256:7078131e59915cf650ac89494c1a2cbf0eaab80166d0b90408dfab98f7bb5bd4

Observation 1e6148d5-599a-405e-8982-b49caa7c7880 · outbound

This paper cites Small leak can sink a great ship–boost rl training on moe with icepop!, Sep 2025.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Small leak can sink a great ship–boost rl training on moe with icepop!, Sep 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:33.909133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:33.909133Z digest=sha256:7b9646ac99f9b4ef78e7f669c6b30bca5c21c75d8ca79d0a07db14b86bf63480

Observation 12399fa0-bf49-48f4-9c37-add1d4c02137 · outbound

This paper cites Stabilizing reinforcement learning with llms: Formulation and practices, 2025.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Stabilizing reinforcement learning with llms: Formulation and practices, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:33.995325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:33.995325Z digest=sha256:458c7cefe035e775c28e41aa1e43fe4f77374a5005bf30d0cf9ae00b797cd2cf

Observation 21e3219c-4068-4cba-a11e-089608d7fb36 · outbound

This paper cites Group Sequence Policy Optimization.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning Group Sequence Policy Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:34.082604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:34.082604Z digest=sha256:a9c69a33346460a92a3544e81f62552d0ba61170c9d3eb215c5880e5ba56c196

Observation b8d1cb8a-44e8-445d-8086-f3b1bc2763ca · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning SGLang: Efficient Execution of Structured Language Model Programs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:34.155714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:34.155714Z digest=sha256:b3a1a6fbcb9ee40cef9e1184d8aa4aedc0527cfb1516644364b2d260646f04be

Observation b3d1f6f7-a2c5-44b6-a836-00045af49020 · outbound

This paper cites X yt µ(yt |s t) π(yt |s t) µ(yt |s t) −1 # = 2ξ TX t=1 Est∼µ h 2DTV(µ(· |st)∥π(· |st)) i = 4ξE y∼µ.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning X yt µ(yt |s t) π(yt |s t) µ(yt |s t) −1 # = 2ξ TX t=1 Est∼µ h 2DTV(µ(· |st)∥π(· |st)) i = 4ξE y∼µ

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:34.238953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:34.238953Z digest=sha256:b7cc6923d7ee0e45418c0b2ce07a31f26adc56ddbd2dc9a9945233ee575b015e

Observation fd5ebb47-0db7-44c1-b13d-b11f3608725d · outbound

This paper cites This is consistent with a sampledDTV gate improving stability on the reported stack, but it does not explain the remaining gap toSAT-GSPO w/ R3 or even toSAT-GRPO w/ R3.

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning This is consistent with a sampledDTV gate improving stability on the reported stack, but it does not explain the remaining gap toSAT-GSPO w/ R3 or even toSAT-GRPO w/ R3

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:34.322864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:36:34.322864Z digest=sha256:4ae2fb96cfc34436126595a75556fd1913ffa1d9a7cb35d7afcbaf48f09dbfba

Pith citing papers

No inbound Pith citation observations are available.