Pith. sign in

Paper Citation Record · LEDGER

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

As of 10 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 14 inbound Pith citation observations for arXiv:2505.24034.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24034 v2

Coverage vector

measured 74 of 74 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:45:30.237588Z

measured 88 of 88 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:19:26.241394Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

74 of 74 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved60
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 07a64a7c-f311-4425-bb11-eeedc2081e68 · outbound

This paper cites write newline.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:23.770147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:23.770147Z digest=sha256:338181e9e36a2f888a4d1249e114615a03833c12229697beea8d774d0baa0625

Observation 5de36d5c-a3cd-4aa9-90c9-651ae928d77b · outbound

This paper cites write newline.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:23.846339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:23.846339Z digest=sha256:c9d9247f645cbf306bd8bf77850a5889002ab58ba97e6fa97077d08ab286da71

Observation d22a7033-1274-48b7-8a45-d975a962a426 · outbound

This paper cites @esa (Ref.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training @esa (Ref

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:23.965479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:23.965479Z digest=sha256:0e101a10440029490d20639babef79ae5b68a7fbb96252d0082a63732f9b1490

Observation db16a862-1b80-41a8-9618-463fc606670b · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:24.134180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:24.134180Z digest=sha256:9f12ea5f0e44101f6562767a8abe2b790bbe3ee634d9f0828ccd3987dd40cc2e

Observation ead16030-8ce3-407b-bd40-d766dac12bad · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:35.730570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:45:24.271441Z digest=sha256:5fe00bae6d2fada19c598b3a90af5543e6e42ec068c1fc26b5f5dc16bcbc71bc

Observation 45ff0517-193d-4dfa-b1a2-c1db99266a2e · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:24.334354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:24.334354Z digest=sha256:ad29e9547cde207c7008175bb7ba147d042af6d7919e713229a18b65a6d8fb6f

Observation f59b1836-af27-4c24-b324-e07d94e6bd63 · outbound

This paper cites Claude: Training helpful and harmless ai assistants.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Claude: Training helpful and harmless ai assistants

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:35.556691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:45:24.414307Z digest=sha256:957f8084b5ff74435048ba2e385ba81f12df101e40b5b4c484a35d7b2ac6da44

Observation 989c91bd-78b9-4040-b10a-bb580ba78bd7 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:24.561905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:24.561905Z digest=sha256:e248cfd22258cab419e632387f1a0f7faeee5c053d8d1aaeb16d4b3ed65ae9a8

Observation a392faf4-81bc-4a94-be1f-07e9a06b1d61 · outbound

This paper cites Dota 2 with Large Scale Deep Reinforcement Learning.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Dota 2 with Large Scale Deep Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:24.616677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:24.616677Z digest=sha256:73daa658d78ac835ffdcdb5b56bc63dda9860b52a79451510ebb2204a747068e

Observation 0217a153-b800-41d8-b2dc-6a03a2008b5f · outbound

This paper cites D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:24.670104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:24.670104Z digest=sha256:65362780c5a5d981b810b2ecdf0537dd45a5f03641941cf7d9445962d4cb0f64

Observation 9eca5e5f-1329-4d86-8666-7011d054086a · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:35.436611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:45:24.737131Z digest=sha256:6744484052f4b2fd52bf88b81543fd321cf62f793339d80bc4bd4f7f1da74b24

Observation e5c9a32b-14a1-4940-bb45-0c8e9d9623c9 · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:35.249352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:45:24.830347Z digest=sha256:9e4baff7d67b2ff5c92b70ed4ea1ec535c2eab45bc2cb77ca772ddb31ec80da5

Observation b94aa1bd-1675-437b-b240-bc9492c67b0d · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Training Verifiers to Solve Math Word Problems

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:24.928868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:24.928868Z digest=sha256:681991e07ccbb8ba162c4c372f3a58ab88b7d1618a99b79cef4457d3d4515c10

Observation 11693588-6706-486d-be96-698d8782eec0 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:25.046427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:25.046427Z digest=sha256:5d9f6c9e5425ad938c80366daab6ab2397b8aad234aef4860c2587985c6795c0

Observation f5e39292-d015-48f7-834a-6d8d9c80605f · outbound

This paper cites RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:25.186917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:25.186917Z digest=sha256:4628c6f0d347e8cc359d4ecc6aee2f6f9204f858a9982471ee9cbe554bdf3a2d

Observation e708a75c-3b4d-428f-943b-9da9ad7ccb01 · outbound

This paper cites Competitive Programming with Large Reasoning Models.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Competitive Programming with Large Reasoning Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:25.443666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:25.443666Z digest=sha256:47546cca45fcbf020e6216d76116120099b9abc6618ca50c9aa780d071c5ba3c

Observation b256365c-8f1c-482a-b6ee-372103d1a48e · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:35.084914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:45:25.543781Z digest=sha256:2e56528a29913a33cdc44e68ce9cb6058b99f1516fa25984fbd5348ff03d4379

Observation 1aff602b-b69d-462f-8792-7d66865a4e5d · outbound

This paper cites RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:25.658504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:25.658504Z digest=sha256:e1e77df19cef250caf14f58c5f094cf95fdfd216355cf477c463c942710da59c

Observation 41064d76-6cec-4431-99ca-ac9028e20ae3 · outbound

This paper cites Bard: Conversational ai by google.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Bard: Conversational ai by google

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:34.891145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:45:25.746481Z digest=sha256:72a9738a5bd34286241adc5edbfcda2ab30f597153486f126afdc3383ef4c3a8

Observation 73fccd74-5a99-4ea2-acf7-9c161fd76513 · outbound

This paper cites Deliberative Alignment: Reasoning Enables Safer Language Models.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:25.860418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:25.860418Z digest=sha256:774cec920b8680625c14bdd8c79203561f6ce0642781787189f5661d5a7c78c4

Observation 5f7c0837-28d7-4e25-9eba-81eca24608ef · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Measuring Mathematical Problem Solving With the MATH Dataset

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:25.995232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:25.995232Z digest=sha256:cb5833937e39e33c01620aa0066b5d118da63181438815e3d18a709a99e7b9f4

Observation 2128733f-0124-4d36-8e3d-4cdc5e9d507b · outbound

This paper cites Deep Learning Scaling is Predictable, Empirically.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Deep Learning Scaling is Predictable, Empirically

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:26.107786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:26.107786Z digest=sha256:1f66085828990f246f224a9ed4dd06f329a952196dbff4c5be032f4488a7ebd5

Observation e19f7d38-bb19-49ea-a33f-163005cb279f · outbound

This paper cites Training Compute-Optimal Large Language Models.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Training Compute-Optimal Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:26.228541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:26.228541Z digest=sha256:ae12024aa018b0e274d9cce1ab0195c5c56f02c1f61b884755279076f014aa7e

Observation 59cc2812-bd55-49a7-89d5-b1bd41a21eeb · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:34.707529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:45:26.309987Z digest=sha256:b419d8958ba83fc471e5b5f0266cdc0a65ed3562c409222c0c3b2d7c83006bf9

Observation c25dc515-d9e1-49fa-a8cb-a386905f912e · outbound

This paper cites OpenAI o1 System Card.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training OpenAI o1 System Card

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:26.395468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:26.395468Z digest=sha256:b3e892c62d2ed7e55cef20e9e7c8a2cfc4fa2eaf6eff9ac290c6460499fb2f02

Observation 2d545a37-25e9-4084-93c2-11c1aa9bb837 · outbound

This paper cites and Abbeel, P.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training and Abbeel, P

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:34.582388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:45:26.446359Z digest=sha256:90b0dcad8ceb54c794cedc7e0dec23dd9fe0e1e46ad11d65c7707e9014a52432

Observation 52141ded-4384-4fbb-bf54-6370bb16020e · outbound

This paper cites and Langford, J.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training and Langford, J

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:34.404114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:45:26.503995Z digest=sha256:d778b2fc202cb49aeed771820fcad47cc31773e9846baab4404fec3b7a00050f

Observation c25ab529-3f3d-466c-8630-6f9e046841cd · outbound

This paper cites Scaling Laws for Neural Language Models.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Scaling Laws for Neural Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:26.558951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:26.558951Z digest=sha256:f24922e628e163bf5ea522900d1152fc90533607e56107a27642fd01a9667cbb

Observation 3000b53c-46e7-4d71-b723-93897e3a6536 · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:34.275455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:45:26.611239Z digest=sha256:00573b8cf5b32c1e6566c95d9c72b57e7f2195ba920cc7441f296b207666dd48

Observation 5139bbd4-11cc-4431-be00-5e305202204e · outbound

This paper cites VinePPO: Refining Credit Assignment in RL Training of LLMs.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:26.663718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:26.663718Z digest=sha256:c8a629d1392a57d44a113a9460a7d4ad1ddb3e24d0e3372e3f543c2c87515924

Observation 4d97c8bf-a410-47c1-abd8-07a60b8739b7 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:26.712582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:26.712582Z digest=sha256:ca34ad33f5429e716d3ad0e31a3b948f577ca4d30a5e6afc939eefdae5524003

Observation 5ed95391-f60b-442b-a38c-da2cd65d4524 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Adam: A Method for Stochastic Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:26.768253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:26.768253Z digest=sha256:ddcc7de18b080ea7fa88ffe959ef527717e075688c5151840b202d632e62c49a

Observation 85c1a2e6-39ea-4fae-a9e7-0fe0007ea9c1 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:26.809054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:26.809054Z digest=sha256:2a4eea4905f5df5d9a29ab2d3e3923f6469fb869998a48097d629309609648bd

Observation baf38e29-692f-47ad-a007-0a6dfa798140 · outbound

This paper cites G., Park, J.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training G., Park, J

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:34.179070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:45:26.854355Z digest=sha256:63fa7e65074ada39ef94b6f125a08d2744e8432f16d26c928ad94bae26cd4ded

Observation 8d3eb950-d117-4390-b03a-00a18af5e0d2 · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:34.060963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:45:26.901233Z digest=sha256:fac27d91fa08ed9b93aefa651d90d9d89211b3883dc643b4254dd5bc2eb040f5

Observation 629e7fe6-a2ae-4098-9814-c7fc615ab6e3 · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:33.950050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:45:26.959962Z digest=sha256:e530725903b065c7c7d3b5b732a34fdf252f5f8cddbf08df5f585eeb2afdd559

Observation b2b022ba-afae-4eea-b518-3aae00c6c6f1 · outbound

This paper cites Let's Verify Step by Step.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Let's Verify Step by Step

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:27.040318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:27.040318Z digest=sha256:7d445f3fa0def4b1ca0e5f3b0ef86dbd32891dd009a57519a6b6e9bce34e3a62

Observation 759d7f3d-bc02-4068-8885-51aa739c7f6e · outbound

This paper cites The Llama 3 Herd of Models.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training The Llama 3 Herd of Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:27.098290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:27.098290Z digest=sha256:41081c40ee2476c68458da029224f50f8468ff77a95a065f7f9d2f56e7b15129

Observation 2ac4ad63-b16f-4bfc-9372-3147a3f164c1 · outbound

This paper cites Decoupled Weight Decay Regularization.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Decoupled Weight Decay Regularization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:27.144979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:27.144979Z digest=sha256:5d990d61d7b41a403471133d7190c38ffadb97dcda65c46ec2df5cdc30b7c7c3

Observation 753d3c29-85d2-4bb2-8f2b-8768ebdde0c3 · outbound

This paper cites P., Paprocki, M., C ert\' i k, O., Kirpichev, S.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training P., Paprocki, M., C ert\' i k, O., Kirpichev, S

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:27.201064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:27.201064Z digest=sha256:4d80073b38c72783fb015a0cafa1f3637da4e629b41bc2797d09870543f2028f

Observation 1b625d55-d532-4228-b7c6-1d331bebadd3 · outbound

This paper cites P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:33.815726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:45:27.243910Z digest=sha256:192e3049df42ce3a0b62b7a68ced1164ad8aafc5df732d3b55a5eb28cac4a123

Observation 2ae4da01-07d6-4862-b8f6-c9fd40a544b0 · outbound

This paper cites I., and Stoica, I.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training I., and Stoica, I

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:33.696311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:45:27.297791Z digest=sha256:7d939c56a487e313a9f4149307fea015e3efd5021022a0f6c0eb962f2ab4a00f

Observation eca310e8-d82c-44cf-aaf6-aff43d74a587 · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:33.572033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:45:27.353704Z digest=sha256:ef366f6a516f83366c566bcfd1c24718f647bcb29eee4883f6b8176a94d3f396

Observation e0f34c75-af43-4eb7-afbc-39fa7ab6c496 · outbound

This paper cites B., Singh, V., Lin, M., Gimelshein, N., Desmaison, A., and Yang, E.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training B., Singh, V., Lin, M., Gimelshein, N., Desmaison, A., and Yang, E

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:33.440164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:45:27.404939Z digest=sha256:4a3e95f1d1793d65c2ec13322b4f81d021c1778daf546c5dc5f83d18be364960

Observation 9c42c152-03f6-416f-aa9e-44c0921b7bfc · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:33.302036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:45:27.448474Z digest=sha256:2bc2fac177125fb13a2d64d408bd6e20516264c844f5a91f352aca318f10a063

Observation e224529b-dd60-40be-972f-44b209a252f9 · outbound

This paper cites GPT-4 Technical Report.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training GPT-4 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:27.526241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:27.526241Z digest=sha256:d5de7802af71be0e3d452e25f5225529e77df8476bf6f0ef0933855c598a45d0

Observation 49944a65-5e2f-4e27-9d42-68d438bdb65e · outbound

This paper cites Training language models to follow instructions with human feedback.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Training language models to follow instructions with human feedback

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:27.595282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:27.595282Z digest=sha256:a95c94fd170d17d26fdcad3bf11523d17877c4d3c99ccb03e0ab4bdb2f9c9a3e

Observation 2169cf50-ffc1-47c5-9777-6ed001b2056f · outbound

This paper cites and et al.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training and et al

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:33.167654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:45:27.679104Z digest=sha256:ea5edbaae7d8274ffa7ff4586882321916e9c51256037ae869414e0647fdfc5f

Observation e2749790-728d-4965-9cf0-94d6d3fc1d7c · outbound

This paper cites S., and Singh, S.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training S., and Singh, S

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:33.056767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:45:27.725997Z digest=sha256:848b83f5dc8ab4a3532cc84938f545f3d22ade1d2fbb3b2e0edaf19d2f34502e

Observation b2f0289c-d1e7-4be0-a362-40bab67b28cf · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:32.959344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:45:27.813391Z digest=sha256:71f33aa5b3ed74a0ca9b40c9b8154039a77517e0f455b0bb152ef22177d745af

Observation 085e1425-1e98-4dbf-82c4-9f6528c057f5 · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:32.826777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:45:27.880640Z digest=sha256:e2ac44d86b60157516a98b4c84844d4e5d65a63b9d0f2279fa13898c28aa89cf

Observation 497bfe96-31bf-49da-b6e3-5dea5b5e6a9c · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:32.650731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:45:27.990465Z digest=sha256:ca0cce86c471e60ccaa4d892a34d1706eca5c38be86ef91b3641822ab939be78

Observation 2dabefd9-2f45-4bb8-86fe-7f6d9ef369d8 · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:32.433370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:45:28.113659Z digest=sha256:5936223307fff8d017de529ac21dbc9b047a345ab471877fb796fab43f6ad20c

Observation 0c3b2671-921c-4580-ba77-6cf25f206ba7 · outbound

This paper cites Proximal Policy Optimization Algorithms.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Proximal Policy Optimization Algorithms

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:28.227786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:28.227786Z digest=sha256:74a380d0b6aaddce6313f212a45783b9f31935dde3097ca4a797a31a0edf62b1

Observation b69a40a9-1ab4-438c-9e68-a9c91c601b4f · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:28.335605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:28.335605Z digest=sha256:d37e9efd55c9701196e331b4968085cf6c76c0b3d2ffe71fc60d5425eccf51bb

Observation 84f342ad-16dc-4c18-a48c-25f7d2a1a718 · outbound

This paper cites S., Aithal, A., and Kuchaiev, O.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training S., Aithal, A., and Kuchaiev, O

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:32.250872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:45:28.486086Z digest=sha256:cc2dd5633ed741147518e931d6740227463ecca246baa1a55b5992ce5ebe530c

Observation a682b063-6ddc-4529-8a0f-2a2c1fb2c606 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training HybridFlow: A Flexible and Efficient RLHF Framework

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:28.605509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:28.605509Z digest=sha256:9d292633ba8b51c6d0ad334046945ea064a96a069d2b8c718b15d535ac29a9a2

Observation 0d72c3d2-2a49-4dfb-8db5-d2060873a8e9 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:28.702500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:28.702500Z digest=sha256:c854e1e8e8878a9aa0a0046d0591ac7951924d148d3092d7ee159fd9339a8feb

Observation bd313a9f-5794-4dbb-9f22-8b1be12cc3de · outbound

This paper cites Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:28.836184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:28.836184Z digest=sha256:65be147f93383c2585e7a1a90d290d3b6d921f9f83eba379346437cfe8e74900

Observation 29d07c9a-94ee-4274-a160-df653b7a835d · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:31.985189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:45:28.918413Z digest=sha256:2facc9ccb70a8689686a52a480693e7ede64dcb5786561bb0cb64983aabaa8db

Observation 2591cb05-32c7-4872-831e-d37cb3368112 · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:31.792103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:45:28.973764Z digest=sha256:cf7a1ee57b89b992c6558a471977298acb8a0c069c6cd68af800af65aa56f173

Observation 40392732-00a9-4837-8ca2-61d1d0092168 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Gemini: A Family of Highly Capable Multimodal Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:29.052265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:29.052265Z digest=sha256:2915feeb5edccca80eb76e3e19899fd0060cf802b068884cbfa685ddbe157f80

Observation a060fd00-2d23-4b38-929f-5030ff16084f · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:31.620568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:45:29.206789Z digest=sha256:13fe874e3de965f94e43ebfda0524e88b30649ef6e4b40282a09c63edc82dfba

Observation 7239647a-83cb-4ef0-bb04-9746ced809ca · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:29.254014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:29.254014Z digest=sha256:30c08431d0d135beb6a14fed6f7cb6591ff4fd9177410d89f0cff8ca120f79ba

Observation 0e9e40be-f8b0-41bc-9414-010bce4df230 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Solving math word problems with process- and outcome-based feedback

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:29.295485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:29.295485Z digest=sha256:259f4c385f13017eb99ce757de94607612432187a12fd84ba0fcf7832dd3ccad

Observation a39cd7e0-210c-44cd-9ec5-768f76a8034e · outbound

This paper cites M., Dudzik, A., Huang, A., Georgiev, P., Powell, R., et al.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training M., Dudzik, A., Huang, A., Georgiev, P., Powell, R., et al

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:31.458469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:45:29.338170Z digest=sha256:cfcb410224b5bbf338dbe6fcb79f8286b52dbce5c036b2d4e60f3d8d12092916

Observation 81515745-7875-4b0b-8e5c-5ced99f458a4 · outbound

This paper cites Sample Efficient Actor-Critic with Experience Replay.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Sample Efficient Actor-Critic with Experience Replay

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:29.387646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:29.387646Z digest=sha256:b56c6a04422053eb2b22e5157c5156827c765b41d1ce11f7e190c4e2ff5a9904

Observation 8f639335-f82c-48d5-a3be-fdd1be02f552 · outbound

This paper cites SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:29.562365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:29.562365Z digest=sha256:3fc4fe9941bcc9ee68eccb2ce291e26f78966ddbd6176b13fd9c683996268f10

Observation 9bb2f068-38ab-4814-b125-249b1d0199b6 · outbound

This paper cites An Adaptive Placement and Parallelism Framework for Accelerating RLHF Training.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training An Adaptive Placement and Parallelism Framework for Accelerating RLHF Training

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:29.683650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:29.683650Z digest=sha256:19abfe007f17fbcc133277ca325abba204dfc6e78f9696c71577eccbb41e46ba

Observation 0486456e-f3ee-40a1-9652-95381310eb25 · outbound

This paper cites A., Jin, D., Peng, K., Han, E., Nie, S., Zhu, C., Zhang, H., Zhou, W., Zeng, Z., He, Y., Mandyam, K., Talabzadeh, A., Khabsa, M., Cohen, G., Tian, Y., Ma, H., Wang, S., and Fang, H.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training A., Jin, D., Peng, K., Han, E., Nie, S., Zhu, C., Zhang, H., Zhou, W., Zeng, Z., He, Y., Mandyam, K., Talabzadeh, A., Khabsa, M., Cohen, G., Tian, Y., Ma, H., Wang, S., and Fang, H

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:31.278750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:45:29.826054Z digest=sha256:12df937b5184d41140863623b3fdc64896b8785dbcaf6dcc66fba6886ffdd22c

Observation 46341aad-f537-46d2-8202-9159ac88dfac · outbound

This paper cites Y., Ruwase, O., Rajbhandari, S., Wu, X., Awan, A.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Y., Ruwase, O., Rajbhandari, S., Wu, X., Awan, A

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:31.128954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:45:30.031519Z digest=sha256:8097a6d7a94eda93adc89448b9dc77e07f1c869464c8f326d5d0f19e626ce609

Observation 7010431b-99de-49de-b0f4-33e42bccc211 · outbound

This paper cites When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:30.128496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:30.128496Z digest=sha256:dff4f2dd2fd6f70ad1f1804ed4bf9cd526727ad94d0c541d07f8e3d0e88301fb

Observation 8648f7e9-c63d-4282-9317-11333bb0cad1 · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:30.901240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:45:30.191698Z digest=sha256:becb671203f3dc11469feecea163628568318bf2b154495e670f98727dd302f1

Observation 4944aca5-0cd3-4705-aef9-659817159aaf · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Fine-Tuning Language Models from Human Preferences

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:30.237588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:30.237588Z digest=sha256:4f2592158f30d52869d4a1259b333027dc8c29e0ad8fefd1c04e9cd2254ad932

Pith citing papers

Observation e74338b5-eb36-4ba3-884a-ddd1729123f4 · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

Reference 133

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:32:01.297713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:930077976c67de84b6d91b1454abdb5eae5aaef7080e84bd4a957af3070c5ec4

Observation 62aa2849-4bc7-4bce-953f-b36661f1c2ce · inbound

Magistral cites this paper.

Magistral LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:26.241394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:26.241394Z digest=sha256:d323bdd0d3513e7fa98a8d1633a33266d4caf6bcf83a1018cd5b590383f61f16

Observation 77243c07-76f2-453d-878e-2288b1fa8aa7 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

Reference 200

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:43.979849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:43.979849Z digest=sha256:31b1b97011133978ee75c4fceb2f18167600a8b240598244acc385363c61ba7b

Observation 317ba9ef-ed90-4c32-8c98-33dd0c7d27ea · inbound

HetRL: Efficient Reinforcement Learning for LLMs in Heterogeneous Environments cites this paper.

HetRL: Efficient Reinforcement Learning for LLMs in Heterogeneous Environments LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:23:36.641772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-16T22:21:26.271796Z digest=sha256:a03fef6aa94972a94dfd407d612034d9b775e1a77e28c177552441507cf9523a

Observation 4e20ecf2-a14a-4eec-9edc-ab8728858d75 · inbound

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training cites this paper.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:54.516468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:54.516468Z digest=sha256:4e488a9f3bdbca92db59400fa0f600bf905ba5596b2e7bfe30488d52d80807ce

Observation ee4ba978-e8d1-489f-b615-3c4137c3991b · inbound

TensorHub: Scalable and Elastic Weight Transfer for LLM RL Training cites this paper.

TensorHub: Scalable and Elastic Weight Transfer for LLM RL Training LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:11:00.631266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:16:33.020610Z digest=sha256:7c9222f96cb23b391aec33e03badad0d0da57ec78c11220bc0babd0297ff333c

Observation e6ba9bdf-05e5-4cc8-8ccf-873d988d8145 · inbound

DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training cites this paper.

DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:46:26.425646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T13:49:26.560459Z digest=sha256:4bb8bb017c573c266236ac703e6b6e5476b79299f5af2780e38dd2a1bd2e70c7

Observation 3521626d-9bfb-47d7-924e-d54365baccd4 · inbound

When to Stop Reusing: Dynamic Gradient Gating for Sample-Efficient RLVR cites this paper.

When to Stop Reusing: Dynamic Gradient Gating for Sample-Efficient RLVR LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:33:07.364766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T07:31:23.194325Z digest=sha256:18d93593ac77ef0efc862e48e7b3ae1ac280194c81408ba60c493a7176bb7c4e

Observation 3ca7aa93-bd09-4601-9abd-9e0159acd6aa · inbound

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training cites this paper.

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:53:55.792604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T19:47:09.817243Z digest=sha256:f5181af5fc57157115543c516855251c3d77b2103af1be6a994ad82d792e06f3

Observation cda6d2b5-4d4e-4fcb-9c45-e253884a7721 · inbound

Libra: Efficient Resource Management for Agentic RL Post-Training cites this paper.

Libra: Efficient Resource Management for Agentic RL Post-Training LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:27.151232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T11:02:00.385932Z digest=sha256:9c786fd1b2af4ea0cb1c5bd944f43927b8b6944d35fa23155bb3f9f22094b4c4

Observation 1a77efd9-f915-4fbb-b87c-7876560eb8d5 · inbound

Rollout-Level Advantage-Prioritized Experience Replay for GRPO cites this paper.

Rollout-Level Advantage-Prioritized Experience Replay for GRPO LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-07-02T06:06:41.077555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T07:43:21.284574Z digest=sha256:18e4a722e3a2cd963813fc5b08215ca2f8bafdbcb289aafedefe106c54e4bcbf

Observation 9394f3ea-be4d-41d8-a70f-3022c8445dcc · inbound

AsyncWebRL: Efficient Asynchronous Reinforcement Learning for Multi-Step Visual Web Agents cites this paper.

AsyncWebRL: Efficient Asynchronous Reinforcement Learning for Multi-Step Visual Web Agents LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T11:56:55.543867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T02:45:25.694378Z digest=sha256:399b07c8ae43dd0251cb70a633b13e4ae197a20c4f9f3213f3ff21bb45ed796d

Observation 2f99b87d-9796-49cc-ab6b-9a492c5824bc · inbound

Sparrow: Sparse Rollout for Stable and Efficient Long-context RL of Large Language Models cites this paper.

Sparrow: Sparse Rollout for Stable and Efficient Long-context RL of Large Language Models LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:07:26.505309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T19:10:53.882876Z digest=sha256:e50d5e89f1f69cb4268b7ac2645b427bdfea72ed0107a1a50dfb2bc6d27c652c

Observation a9cd52f0-5201-422d-9aed-8d6667281aaa · inbound

Harnessing Routing Foresight for Micro-step-level MoE load balancing in RL Post-training cites this paper.

Harnessing Routing Foresight for Micro-step-level MoE load balancing in RL Post-training LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-07-03T12:58:08.711158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T08:35:16.435272Z digest=sha256:aad03db518fcc926706b10bcd2c137a88e4eedd5a09a8f3b39b4c29539325b7d