Pith. sign in

Paper Citation Record · LEDGER

AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning

As of 23 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2608.00455.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.00455 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:23:02.622661Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy14
  • unresolved10
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 87ac05cc-0bc8-47c7-9ac4-658f77604162 · outbound

This paper cites Real-time reinforcement learning for composer.

AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning Real-time reinforcement learning for composer

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:23:03.199496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:23:02.524661Z digest=sha256:d2c0a2b60639a63a626bc348eb0d2b050ca052a846daaeceb0b673a72bba1834

Observation 27628c50-99db-4e9e-b6b1-99b8c1f46a50 · outbound

This paper cites Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents.

AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-15T15:23:03.021885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:23:02.533029Z digest=sha256:9c8e29a60f8b8b4d92455a0be8729274c0909cb69c4e0a6b98ccfdbb36969c3f

Observation 9376c251-19d5-4556-a222-771bab938de5 · outbound

This paper cites ProRL Agent: Rollout-as-a-service for RL training of multi-turn LLM agents.

AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning ProRL Agent: Rollout-as-a-service for RL training of multi-turn LLM agents

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T15:23:02.537534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:23:02.537534Z digest=sha256:4598ea478303d8e1483541191b6806de0ca136307bc0fcbf3fa582ae25fb0548

Observation 27e87b0f-4d8e-408e-ae44-4e2dffbba69f · outbound

This paper cites RollArt: Disaggregated Multi-Task agentic RL training at scale.

AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning RollArt: Disaggregated Multi-Task agentic RL training at scale

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:23:03.177185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:23:02.541308Z digest=sha256:b1c1829c5583ec5e0f8df104bf100bb6856d9fadb3ca401497cc7347d563a000

Observation c31e978f-7f91-4a81-b236-2660151a972d · outbound

This paper cites Rollout-training co-design for efficient llm-based multi-agent reinforcement learning.

AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning Rollout-training co-design for efficient llm-based multi-agent reinforcement learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T15:23:02.545196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:23:02.545196Z digest=sha256:a5e85c6562084fd029bda55da542b2207a9f14d07c27c6c5cead8ef940dda997

Observation cfd0eaa0-e3c6-4adf-8550-14ff79b74b66 · outbound

This paper cites Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL.

AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T15:23:02.549141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:23:02.549141Z digest=sha256:1f1c490d82a8d33b0e3b59b1284926251324257e1a13c20eb3fbf16f2b882663

Observation aff6792a-6898-4c3f-ad61-9fd9613697b0 · outbound

This paper cites AuroraRL: Fast, Fault-Tolerant, and Cost-Efficient Reinforcement Learning over Decentralized Network.

AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning AuroraRL: Fast, Fault-Tolerant, and Cost-Efficient Reinforcement Learning over Decentralized Network

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T15:23:02.884699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:23:02.553385Z digest=sha256:42f04316e91e3a9745828eaa823797125a5c9ac7d5290dad8c51b5e614626774

Observation c4fd0578-3dc6-4fac-bc5f-d59d9226abbb · outbound

This paper cites ByteCheckpoint: A unified checkpointing system for large foundation model development.

AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning ByteCheckpoint: A unified checkpointing system for large foundation model development

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:23:03.165793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:23:02.557094Z digest=sha256:1fbfa1b734643800a8c54010cf3b4b668ef8780d69f3399ab0e7f39737c6ec0e

Observation 8fc109bf-49fe-41e2-b220-4cf006fa3b73 · outbound

This paper cites HybridFlow: A flexible and efficient RLHF framework.

AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning HybridFlow: A flexible and efficient RLHF framework

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:23:03.154803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:23:02.560505Z digest=sha256:852adaec75819bf63d118cdd334992d8018304597d998b08a5be64ad03d553e9

Observation d4fa7514-51fd-4a6f-85b5-a417ef8a5703 · outbound

This paper cites Openrlhf: An easy-to-use, scalable and high-performance rlhf framework.

AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning Openrlhf: An easy-to-use, scalable and high-performance rlhf framework

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:23:03.143619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:23:02.563927Z digest=sha256:6b11114ce956dd8d2f6de6d3b7afa634ef7480e09e5dc53abe2e8f6184369593

Observation bd125550-762a-42f9-85cd-1566e98f3969 · outbound

This paper cites NeMo-Aligner: Scalable Toolkit for Efficient Model Alignment.

AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning NeMo-Aligner: Scalable Toolkit for Efficient Model Alignment

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T15:23:02.567717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:23:02.567717Z digest=sha256:7c462116158ed702dbd7f83f849baf0f60e728a8e84117878a6af35e094ea091

Observation aa50c4a5-d853-4991-90d0-6dc8df4729a0 · outbound

This paper cites Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library.

AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T15:23:02.572250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:23:02.572250Z digest=sha256:0f1d58d659ad92e467ebfe46054ab0b3f3031bf6f16f79fe32a271fb71eab521

Observation b506897d-387b-46a3-81e0-ef940679a0b9 · outbound

This paper cites Areal: A large-scale asynchronous reinforcement learning system for language reasoning.

AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning Areal: A large-scale asynchronous reinforcement learning system for language reasoning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:23:03.131802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:23:02.576111Z digest=sha256:b105af54bc48183ea0b6920465698d660f06c28b9ff500cc4e2cd7536e13aaef

Observation ce30214d-06b0-46ee-b380-92a1c2411c92 · outbound

This paper cites DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training.

AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T15:23:02.579866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:23:02.579866Z digest=sha256:a8b4ff9b3a12a15739a519deca658f0a2a0ae2914354de0782089f9ef6bc737c

Observation 403052df-cb3d-4144-90fd-302660ec2b96 · outbound

This paper cites Laminar: A scalable asynchronous RL post-training framework.

AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning Laminar: A scalable asynchronous RL post-training framework

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T15:23:02.584658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:23:02.584658Z digest=sha256:e8919b0f7462d1c838ee20c0e69200d0f2fa3f5f1dc5547a2a77495d46351217

Observation 8e42b59f-1c56-403c-a6ca-62037aa51f5f · outbound

This paper cites Weave: Efficient co-scheduling for disaggregated RL post-training.

AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning Weave: Efficient co-scheduling for disaggregated RL post-training

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:23:03.119730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:23:02.588181Z digest=sha256:66ce75d7c38f554a9f86d89692b6c1534de11f3aa0591a2454f6315d67808d7a

Observation a9c310d6-661b-4110-982f-e8536a440bad · outbound

This paper cites ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL.

AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T15:23:02.591636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:23:02.591636Z digest=sha256:d87150ae8501b8111f277db1f144eb835e0e4eec2c9268239ed94697b929aeb1

Observation 143163f9-60b5-40fb-ad8b-46674e79b591 · outbound

This paper cites Efficient large-scale language model training on GPU clusters using Megatron-LM.

AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning Efficient large-scale language model training on GPU clusters using Megatron-LM

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:23:03.107261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:23:02.595352Z digest=sha256:10ea4c4ee3e64b92281e073d5d589ab618e85374a076728e0e276bb802b3cdf5

Observation cd610275-790d-4e3d-b699-41817f960c22 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T15:23:02.598663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:23:02.598663Z digest=sha256:87f7677c48b5f7395cff93a95989036872ee993b94022ad1ab34db2696f5edce

Observation f107140e-5967-44c9-aed3-e0e26b20d03b · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning Gonzalez, Hao Zhang, and Ion Stoica

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:23:03.095600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:23:02.602183Z digest=sha256:9c2f960f49460f5e576ef698057784b9cc487f7d523f79ff66308a5237bb9eea

Observation 63c92526-3054-498f-b7c7-842eeb2e804c · outbound

This paper cites Gonzalez, Clark Barrett, and Ying Sheng.

AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning Gonzalez, Clark Barrett, and Ying Sheng

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:23:03.084282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:23:02.605639Z digest=sha256:36e74a1dcf2d043932813e813cdebfbdc4ff6bba02f2857f29706f56470177ab

Observation e0d55182-edf2-40af-95f0-ab1ef602376d · outbound

This paper cites Universal checkpointing: A flexible and efficient distributed checkpointing system for large-scale DNN training with reconfigurable parallelism.

AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning Universal checkpointing: A flexible and efficient distributed checkpointing system for large-scale DNN training with reconfigurable parallelism

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:23:03.071907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:23:02.608865Z digest=sha256:cb3f2ec52d921d4f53df6956765ef701faaced431bad04aaee92fcfc07271470

Observation 1307dfa1-5828-4323-884a-207f0631f272 · outbound

This paper cites TensorHub: Scalable and Elastic Weight Transfer for LLM RL Training.

AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning TensorHub: Scalable and Elastic Weight Transfer for LLM RL Training

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T15:23:02.612335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:23:02.612335Z digest=sha256:6f6a91685e202d07ff734ccda09fd58269f39665a31218ae80dc5399fa4e1b16

Observation 12da2c2c-ee1b-4b57-b944-1223a9d627d0 · outbound

This paper cites slime documentation: Delta weight sync.

AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning slime documentation: Delta weight sync

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:23:03.060359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:23:02.615881Z digest=sha256:1f4e9ceaba72d258cba00c0b69ebb6a47c964cc7d7523d5177f58e19a7352b90

Observation 013515be-e6f6-43d8-ac3c-12f64398677f · outbound

This paper cites Decoupled weight decay regularization.

AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning Decoupled weight decay regularization

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:23:03.047711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:23:02.619090Z digest=sha256:dae6bdec6c8bf415c43cdf07d18da7460631488089ed3a2ae02e309a91076832

Observation 5496ea2d-e461-463b-b753-21e702b42580 · outbound

This paper cites B MCore and Hugging Face Parameter Layouts This appendix explains why AREAL-DTE performs change detection after conversion rather than directly in the optimizer-native MCore layout.

AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning B MCore and Hugging Face Parameter Layouts This appendix explains why AREAL-DTE performs change detection after conversion rather than directly in the optimizer-native MCore layout

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:23:03.034644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:23:02.622661Z digest=sha256:1dc290fbced537ecefcf994a54e1bfd9a2b514c1b6a959dc6b8fcbcb7d1c30d2

Observation 5abc359d-5e05-44a1-a157-4732438a5572 · outbound

This paper cites an unresolved cited work.

AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning Unresolved cited work

Reference 2026

Resolution
parse uncertain
raw_fallback, observed 2026-08-15T15:23:03.187794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:23:02.529401Z digest=sha256:5ad4d10e11fd6125974dd7e589b5036b5e626df08463e72670a404515c8c33d3

Pith citing papers

No inbound Pith citation observations are available.