Pith. sign in

Paper Citation Record · LEDGER

StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 35 inbound Pith citation observations for arXiv:2504.15930.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.15930 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:18:15.353219Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 88c77a74-8c21-474b-910f-2e31b5772013 · inbound

DeepCEE: Efficient Cross-Region Model Distributed Training System under Heterogeneous GPUs and Networks cites this paper.

DeepCEE: Efficient Cross-Region Model Distributed Training System under Heterogeneous GPUs and Networks StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:15.353219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:15.353219Z digest=sha256:799a26833eda85ff22cbe6be19741efd524a01447cbd4b244af856854943ecf6

Observation 59953177-8abe-4da9-bdfe-30c73e0f5279 · inbound

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning cites this paper.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 70

Resolution
malformed identifier
arxiv_id, observed 2026-05-15T14:24:22.029099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:90b08af5a534901c2bc73a6f2e50bc611829debfc955d1142ddebaae63abc3ac

Observation 4f11a428-3c5d-49a9-bcf4-d8d0d4410b22 · inbound

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library cites this paper.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:33.073265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:33.073265Z digest=sha256:ac03c59c60c2d212d65da5107eb6c5cd15cd478c705752276e6bd71e1638fe77

Observation 09a90acc-3fd0-4efd-8eda-2e08a99ff9bb · inbound

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models cites this paper.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:24.023367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:24.023367Z digest=sha256:77a251ccb0f389d4e495e57e2e02b196b1abf2f307531ae05c73e4ebdd29ab48

Observation eaef4a1d-b396-4be5-aee8-b1ce2379af6e · inbound

AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training cites this paper.

AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T20:50:27.252236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:50:27.252236Z digest=sha256:deb3189c50adb40bc0997abf09b4f23c2cca02b53335deee8b663e6283798b10

Observation 74cff2a9-d30f-490e-bb5e-b9933ca563d1 · inbound

RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs cites this paper.

RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:30:55.166777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T05:29:46.136115Z digest=sha256:042f7bba91ed13dc31caa02f92e153bf23ba480de8b58881f7d23d4bc845bf3e

Observation d7209fb2-2243-47dc-b1ec-681f4017e71a · inbound

Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning cites this paper.

Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:40:14.565630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T20:38:30.169363Z digest=sha256:b5db58d51b1572aa65799909eb0e8385b7c0a42ef11c56440b92de0c50e5983e

Observation 5f79b4f8-9d80-424e-abdd-656aedd4537d · inbound

HetRL: Efficient Reinforcement Learning for LLMs in Heterogeneous Environments cites this paper.

HetRL: Efficient Reinforcement Learning for LLMs in Heterogeneous Environments StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:23:36.618823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-16T22:21:26.271796Z digest=sha256:182a0659b0812756df6c27f0bf811bd204b20d0d9fe79a91d890b34500cb81f9

Observation 940e6faa-2206-42ea-b86d-21bf3ee3ce7c · inbound

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training cites this paper.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:56.161461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:56.161461Z digest=sha256:88f1331658ae7a069e6ef7357bdfb95961979df3ee435d720cc9ac7d8483d074

Observation 850b703d-a73a-4f39-aed9-60976c229ce6 · inbound

ECHO-2: A Large-Scale Distributed Rollout Framework for Cost-Efficient Reinforcement Learning cites this paper.

ECHO-2: A Large-Scale Distributed Rollout Framework for Cost-Efficient Reinforcement Learning StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T05:32:41.063975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:32:41.063975Z digest=sha256:a4518403a002eccfcf19ac699b47f1a41d68d094279b162d4f0ba3117a777bfb

Observation 6764a43b-da9e-48b3-9f05-f6a9823d679f · inbound

TensorHub: Scalable and Elastic Weight Transfer for LLM RL Training cites this paper.

TensorHub: Scalable and Elastic Weight Transfer for LLM RL Training StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:11:00.579153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:16:33.020610Z digest=sha256:524a2d710babdf4fa9d8ea3274c907937b4686ce2d88ca412088c5b71afc7561

Observation 8e99c259-8bfb-473f-ac32-523e7d0678b0 · inbound

JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training cites this paper.

JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:11:11.455784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T06:36:33.914193Z digest=sha256:9f929dbe1a29f6d9e9e38a7c11c2e68b9fff2d8ff8e6e09a2db15f71dfbd9416

Observation 7d8401ba-16d8-40da-a466-2bac472ba821 · inbound

DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training cites this paper.

DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:46:26.358485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T13:49:26.560459Z digest=sha256:814f1273474b484c97b6969e2cb9d48fb456eab8ca06964d715604d88435616c

Observation 3dac3891-cd52-429e-baf4-ac9292a3446e · inbound

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL cites this paper.

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:31:16.458025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T05:14:14.168753Z digest=sha256:fe40dfc809d1fd20e430407b3361e000ae82edc19f5ac95680bdef232f6a962c

Observation 3a096d9f-ad4a-47de-855b-debc9c8ae7bf · inbound

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL cites this paper.

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:39:53.330649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T08:39:31.911497Z digest=sha256:176d7db057a2fa2bdcfe561942c6e79f7114a268beeb976241b5d45f560ca54b

Observation f13a6e96-ffb6-4569-ba9e-afb95c171d0d · inbound

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models cites this paper.

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:00:55.308346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-11T02:02:41.411795Z digest=sha256:b1ed0fc15ffa340b73cc3de1b21eb8b08e146d719fe001efff5ae372d3c6e328

Observation cac17c85-9c38-4b1f-9928-8f512c323582 · inbound

FlashEvolve: Accelerating Agent Self-Evolution with Asynchronous Stage Orchestration cites this paper.

FlashEvolve: Accelerating Agent Self-Evolution with Asynchronous Stage Orchestration StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:51:43.127104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:46:59.400834Z digest=sha256:26aca2a0f522e9b9dcdb68d9a94b2711a72cf0bbb4b3e1745a1d82a5fd779275

Observation a4915de7-acd6-45b8-b6a4-9a4118b9a47a · inbound

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning cites this paper.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:46:14.859743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:9f378c994cff5cd663c54ce8e086cfec75f3dfafcbaec7268d07bf40900673ae

Observation 2b31c08b-f0c0-41bb-b1a6-12f9e9f58af4 · inbound

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning cites this paper.

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:01:30.777660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:23:13.071107Z digest=sha256:50cd7a475a2c2a10897e1533443797b104559305234c3668e4baea278dfc118a

Observation d060c8be-386b-4cbb-be5f-fad2cb4be93d · inbound

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs cites this paper.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:13:43.778925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:c48b2ea347d833069c0137b81e36e4a3bfa5cdd7195768e0de8c1a57f961ebf5

Observation 7e9154e6-835d-4752-9f43-fafcf454e4b3 · inbound

How Off-Policy Can GRPO Be? Mu-GRPO for Efficient LLM Reinforcement Learning cites this paper.

How Off-Policy Can GRPO Be? Mu-GRPO for Efficient LLM Reinforcement Learning StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:43:19.734238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T13:38:45.835754Z digest=sha256:01cc06691c45e15a45d992236660ba18d6529e966e8c738b95ff28f86b7f6e1e

Observation acafb593-ecf7-4218-a7bc-8153d4c3e4aa · inbound

Accelerating Long-Tail Generation in Synchronous RLHF Training via Adaptive Tensor Parallelism cites this paper.

Accelerating Long-Tail Generation in Synchronous RLHF Training via Adaptive Tensor Parallelism StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:55:11.586454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T00:54:27.551174Z digest=sha256:85043a238e07754d279b7720cec6fe425ff7cd4f725d3c4d0933895747067b3a

Observation bafa7e87-d6b2-494a-ad89-159bcb56b9fe · inbound

SiDP: Memory-Efficient Data Parallelism for Offline LLM Inference cites this paper.

SiDP: Memory-Efficient Data Parallelism for Offline LLM Inference StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:13:17.817129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T10:07:06.141581Z digest=sha256:a6c418718827bc1f3f168fc71de16b543411ca3b853c754b9a29e2a74eb7700e

Observation 313d458a-1224-4b4d-8c51-f82bee97f89a · inbound

DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning cites this paper.

DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:42:50.074123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T23:25:30.618655Z digest=sha256:f78ccf80fa4cc030e746fe711c495528508a5f1c7c59a417640130c18079d925

Observation 3066e0d2-19d8-45e1-908d-b0e073ee6e79 · inbound

Libra: Efficient Resource Management for Agentic RL Post-Training cites this paper.

Libra: Efficient Resource Management for Agentic RL Post-Training StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:27.173391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T11:02:00.385932Z digest=sha256:2da55ad1ff0bc212367824b3e4f0e00d836b178d9022d164dd2020d3da8d16ee

Observation 48418179-a97c-45e8-bc5c-a39dbcb6ede1 · inbound

Rollout-Level Advantage-Prioritized Experience Replay for GRPO cites this paper.

Rollout-Level Advantage-Prioritized Experience Replay for GRPO StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-07-02T06:06:41.021530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T07:43:21.284574Z digest=sha256:625dedd81afcf8a08fb67f4d77122326a741eef971b591ba8ac2ceef3a6526ea

Observation eebaf89b-21f8-47ef-92d7-2c8ebb41663d · inbound

AsyncWebRL: Efficient Asynchronous Reinforcement Learning for Multi-Step Visual Web Agents cites this paper.

AsyncWebRL: Efficient Asynchronous Reinforcement Learning for Multi-Step Visual Web Agents StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T11:56:55.515434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T02:45:25.694378Z digest=sha256:b7d81c4b34afd62812f18c870e9675a400523a1c1b7e7163fee49ce0601f9bb5

Observation 1b028528-a2ae-46a6-a595-a4b14ed67838 · inbound

Harnessing Routing Foresight for Micro-step-level MoE load balancing in RL Post-training cites this paper.

Harnessing Routing Foresight for Micro-step-level MoE load balancing in RL Post-training StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-07-03T12:58:08.698480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T08:35:16.435272Z digest=sha256:86bf13c74d55bacef665833915c0f3ef806fcdfd67d96e58890d4b085381df5f

Observation d389e0db-7c88-45cd-a650-a2682c1d0d57 · inbound

AsyncOPD: How Stale Can On-Policy Distillation Be? cites this paper.

AsyncOPD: How Stale Can On-Policy Distillation Be? StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 34

Resolution
malformed identifier
arxiv_id, observed 2026-07-04T16:29:56.831788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T00:40:57.112748Z digest=sha256:eee87c5fc528ddb7e4ce5646b83d04335d600661379794b7fd8e281c33203f42

Observation bad8377f-0b26-4e14-a931-a2b5eb54ac00 · inbound

Moebius: Serving Mixture-of-Expert Models with Seamless Runtime Parallelism Switch cites this paper.

Moebius: Serving Mixture-of-Expert Models with Seamless Runtime Parallelism Switch StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-04T14:19:54.318908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T04:06:13.426379Z digest=sha256:c78714a8f7b2ddfad9e2c661c32addf6f8d367dbc2d8114730fdc0a6c933b12b

Observation ee440a99-b36c-4fa5-8ae7-eb967396bc06 · inbound

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents cites this paper.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T06:06:40.582867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-02T05:58:39.107614Z digest=sha256:f7c3a5936ee27f9f0fa5a29bd6fdecf0d0d7f44833a72f7075acff2983e816fb

Observation 082f5c95-f61d-4dd9-8c06-0462846c29b6 · inbound

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents cites this paper.

Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:48:49.509444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-03T18:47:46.719344Z digest=sha256:0af9ef2d4735e55f7fc2ebb43d7ae8de65044de3b352c5919198bf94265951ff

Observation b5221918-d94c-451b-b18c-ec0212a48ec4 · inbound

Bidirectional Resource Scheduling for Disaggregated and Asynchronous RL Post-Training cites this paper.

Bidirectional Resource Scheduling for Disaggregated and Asynchronous RL Post-Training StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-13T04:42:35.589143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T04:42:35.589143Z digest=sha256:671df1687d9039deb99d92c1436a963be7f105286bc2af07f27932cc114dac0e

Observation f5568b77-114e-43db-aeff-8c01158870ee · inbound

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning cites this paper.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:31.823521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:31.823521Z digest=sha256:da76c67acde5a6dffcbc2f15ae0707ecefdc7af28a03235c5787e23c7c6f3e56

Observation bc987d12-9c0f-4ba4-b4f1-b067647d1df2 · inbound

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization cites this paper.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:17.765327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:17.765327Z digest=sha256:db9b650f10b9e5341c37ace0a80cd325215a587a881816d5fb15ed717a7c2350