Pith. sign in

Paper Citation Record · LEDGER

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

As of 13 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 14 inbound Pith citation observations for arXiv:2505.24034.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24034 v2

Coverage vector

measured 74 of 74 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:45:30.237588Z

measured 88 of 88 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:19:26.241394Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

74 of 74 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved60
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 07a64a7c-f311-4425-bb11-eeedc2081e68 · outbound

This paper cites write newline.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:23.770147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:23.770147Z digest=sha256:b25b6c305a78863b5070766c1ab97907557b312151698650a1a416d1e8c43ed2

Observation 5de36d5c-a3cd-4aa9-90c9-651ae928d77b · outbound

This paper cites write newline.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:23.846339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:23.846339Z digest=sha256:7e9ae58a6feb96eb489a9ff2d569afae85a05e6250351ebd1f0089e0409960b4

Observation d22a7033-1274-48b7-8a45-d975a962a426 · outbound

This paper cites @esa (Ref.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training @esa (Ref

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:23.965479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:23.965479Z digest=sha256:21ba800c69996ef21c89a22fdb9162e590a19f26238028bb3e273a949d945630

Observation db16a862-1b80-41a8-9618-463fc606670b · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:24.134180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:24.134180Z digest=sha256:d3c74b513d50566b6cfd3dcce4973c637c8b556eb108f1518aa3a368ce84ed3d

Observation ead16030-8ce3-407b-bd40-d766dac12bad · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:35.730570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T12:45:24.271441Z digest=sha256:6940ffe1db89d33fd4533ac8866d6d003fe1245052ea45f085e40caf1802a6e8

Observation 45ff0517-193d-4dfa-b1a2-c1db99266a2e · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:24.334354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:24.334354Z digest=sha256:6d41720813ec5d7e354db03fa911c7af74b9ede3b5d7823d43b76e02a819099b

Observation f59b1836-af27-4c24-b324-e07d94e6bd63 · outbound

This paper cites Claude: Training helpful and harmless ai assistants.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Claude: Training helpful and harmless ai assistants

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:35.556691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T12:45:24.414307Z digest=sha256:fd72725c02cf642c42a833fd03fd19a0c4af0cb6ae445b59c8c18b47e8071232

Observation 989c91bd-78b9-4040-b10a-bb580ba78bd7 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:24.561905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:24.561905Z digest=sha256:ff1b390e51112449defbbc90a1e03967d7d0550462058ae585414dd0420af9a4

Observation a392faf4-81bc-4a94-be1f-07e9a06b1d61 · outbound

This paper cites Dota 2 with Large Scale Deep Reinforcement Learning.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Dota 2 with Large Scale Deep Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:24.616677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:24.616677Z digest=sha256:7ef3998329a91b162cf5a6801e7e47605c360d104518c0c4ef067ff1507afaf4

Observation 0217a153-b800-41d8-b2dc-6a03a2008b5f · outbound

This paper cites D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:24.670104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:24.670104Z digest=sha256:47a63d9b7c189216a82af58f4ebe6cba11ccd701af1465c4f37a6942425f12ea

Observation 9eca5e5f-1329-4d86-8666-7011d054086a · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:35.436611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T12:45:24.737131Z digest=sha256:20129a093f2864ea58a7f483bfaa3dc39d5725f4ac9d32189a25b721b26962ee

Observation e5c9a32b-14a1-4940-bb45-0c8e9d9623c9 · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:35.249352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T12:45:24.830347Z digest=sha256:408df4595fd95741e29e6a5538b3a882660d2942f0441a46a0bc800b2a1a733b

Observation b94aa1bd-1675-437b-b240-bc9492c67b0d · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Training Verifiers to Solve Math Word Problems

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:24.928868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:24.928868Z digest=sha256:ecf975525e30a24378a061939053392b968b42bdaff098a252ccdc6942d2f931

Observation 11693588-6706-486d-be96-698d8782eec0 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:25.046427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:25.046427Z digest=sha256:6d9f3e90e54f0d7cefd0dc85aec7674e5079b67b7ec7e515a91ffb5cc7a8d243

Observation f5e39292-d015-48f7-834a-6d8d9c80605f · outbound

This paper cites RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:25.186917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:25.186917Z digest=sha256:72f8dd4fad71d25923808fa6dcfd3b2ca32b17951a23aa8fad2e49f14448ddfb

Observation e708a75c-3b4d-428f-943b-9da9ad7ccb01 · outbound

This paper cites Competitive Programming with Large Reasoning Models.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Competitive Programming with Large Reasoning Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:25.443666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:25.443666Z digest=sha256:1777dc49aad9ead3c6f438bb8ef3b0685506b2ef5a3d4adb0bc85d9a600d2fe6

Observation b256365c-8f1c-482a-b6ee-372103d1a48e · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:35.084914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T12:45:25.543781Z digest=sha256:7316e174332ba94fbf7387bf1b40aecac142e39491c2acee47f23389921e1256

Observation 1aff602b-b69d-462f-8792-7d66865a4e5d · outbound

This paper cites RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:25.658504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:25.658504Z digest=sha256:831713da40822f49edc364a00df4edff800f2773b0f29790505cd88c1cce7c29

Observation 41064d76-6cec-4431-99ca-ac9028e20ae3 · outbound

This paper cites Bard: Conversational ai by google.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Bard: Conversational ai by google

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:34.891145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T12:45:25.746481Z digest=sha256:10ed17d8cafe506540b20a5a94620e746581bd25f1f9f8fed4e479595c371f2b

Observation 73fccd74-5a99-4ea2-acf7-9c161fd76513 · outbound

This paper cites Deliberative Alignment: Reasoning Enables Safer Language Models.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:25.860418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:25.860418Z digest=sha256:933852a81047d242d7edebacb3c356bb732f255669583f1129649af7fdf660f5

Observation 5f7c0837-28d7-4e25-9eba-81eca24608ef · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Measuring Mathematical Problem Solving With the MATH Dataset

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:25.995232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:25.995232Z digest=sha256:1f4b6baff30cc3b001827b74cb5ecd46725f1a2da933622112d2a5934a500e7c

Observation 2128733f-0124-4d36-8e3d-4cdc5e9d507b · outbound

This paper cites Deep Learning Scaling is Predictable, Empirically.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Deep Learning Scaling is Predictable, Empirically

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:26.107786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:26.107786Z digest=sha256:233e8a94040d452f6acca828c2a56272a02729fa3383e64c7dd139ba4616b2dc

Observation e19f7d38-bb19-49ea-a33f-163005cb279f · outbound

This paper cites Training Compute-Optimal Large Language Models.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Training Compute-Optimal Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:26.228541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:26.228541Z digest=sha256:b9058bce3c6c3730f4f1e51a60de31dc5c4c99dd5e4af7703cce7f8716a5bd84

Observation 59cc2812-bd55-49a7-89d5-b1bd41a21eeb · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:34.707529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T12:45:26.309987Z digest=sha256:38e9da90b4c96ba881caa15170bc4127c46722dc291dca8d48fd818d4060d31e

Observation c25dc515-d9e1-49fa-a8cb-a386905f912e · outbound

This paper cites OpenAI o1 System Card.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training OpenAI o1 System Card

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:26.395468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:26.395468Z digest=sha256:ab2484c55a95de8b2e04bf818706f5b6a6d6b3b7809fcb3ca00ce6d95c9db439

Observation 2d545a37-25e9-4084-93c2-11c1aa9bb837 · outbound

This paper cites and Abbeel, P.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training and Abbeel, P

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:34.582388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T12:45:26.446359Z digest=sha256:4c5fde2f3d14ccdd13a9153d636e6700c691bc9ed5d4a631dc0346efa463ff99

Observation 52141ded-4384-4fbb-bf54-6370bb16020e · outbound

This paper cites and Langford, J.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training and Langford, J

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:34.404114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T12:45:26.503995Z digest=sha256:ac34623bee749853651af1299e16d2c1bcc980b15c2276f0219dab917913efed

Observation c25ab529-3f3d-466c-8630-6f9e046841cd · outbound

This paper cites Scaling Laws for Neural Language Models.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Scaling Laws for Neural Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:26.558951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:26.558951Z digest=sha256:c1c73aafda9f3c2ce07a1e55ec4f3b8656159acb76ef5431c2c9a4c61097f48e

Observation 3000b53c-46e7-4d71-b723-93897e3a6536 · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:34.275455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T12:45:26.611239Z digest=sha256:43957d9505cfba11be7effa3ff4318b26f7dbd9fef020788a055ea651413115e

Observation 5139bbd4-11cc-4431-be00-5e305202204e · outbound

This paper cites VinePPO: Refining Credit Assignment in RL Training of LLMs.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:26.663718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:26.663718Z digest=sha256:9012c9bd8ca8521a6bc507147d93d38d7b65eea50f3ec20c1403bac5e5f1f58b

Observation 4d97c8bf-a410-47c1-abd8-07a60b8739b7 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:26.712582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:26.712582Z digest=sha256:237b2469f2fe5d533221a55df4d6ef5833e539737f54ed2410e9e20c3f58bd0d

Observation 5ed95391-f60b-442b-a38c-da2cd65d4524 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Adam: A Method for Stochastic Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:26.768253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:26.768253Z digest=sha256:9af7dd7c77b9231c4e3b1b95764a0aa07ed429e08a4ea7cec38f21c7851c63ab

Observation 85c1a2e6-39ea-4fae-a9e7-0fe0007ea9c1 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:26.809054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:26.809054Z digest=sha256:752680cecc48be729ab8b418392135408903c4b35b60db17088718523b597958

Observation baf38e29-692f-47ad-a007-0a6dfa798140 · outbound

This paper cites G., Park, J.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training G., Park, J

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:34.179070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T12:45:26.854355Z digest=sha256:d9d739222b7533bcaeb78bfcdfb3621790ff422ff9d8fdd6884afe706b7a67ec

Observation 8d3eb950-d117-4390-b03a-00a18af5e0d2 · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:34.060963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T12:45:26.901233Z digest=sha256:123845517492198b70c54d7ce07b0c3f3c441052c8ecf8ae5b94ed26cb695bba

Observation 629e7fe6-a2ae-4098-9814-c7fc615ab6e3 · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:33.950050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T12:45:26.959962Z digest=sha256:eb061a5290bb77030002d860d6c1413669f66353e9bdd937ba31d019ace6a2b5

Observation b2b022ba-afae-4eea-b518-3aae00c6c6f1 · outbound

This paper cites Let's Verify Step by Step.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Let's Verify Step by Step

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:27.040318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:27.040318Z digest=sha256:345d0df68386a3920d6f08b25d0b790f80ea013c93d1857c765f161a7c90b057

Observation 759d7f3d-bc02-4068-8885-51aa739c7f6e · outbound

This paper cites The Llama 3 Herd of Models.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training The Llama 3 Herd of Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:27.098290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:27.098290Z digest=sha256:9c07f004eab2bb625918d8d2ec2ce5967939092bfbe8553edd6e3c91c13974a0

Observation 2ac4ad63-b16f-4bfc-9372-3147a3f164c1 · outbound

This paper cites Decoupled Weight Decay Regularization.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Decoupled Weight Decay Regularization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:27.144979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:27.144979Z digest=sha256:414a95fb9ddf6dd56291c0cce554d84de8cd403e0dab33709607933bd60b5e5e

Observation 753d3c29-85d2-4bb2-8f2b-8768ebdde0c3 · outbound

This paper cites P., Paprocki, M., C ert\' i k, O., Kirpichev, S.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training P., Paprocki, M., C ert\' i k, O., Kirpichev, S

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:27.201064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:27.201064Z digest=sha256:47415e1f6d5dabe6b1fde45d79139b09bc094b9ef396a6d7a5674dde03332b54

Observation 1b625d55-d532-4228-b7c6-1d331bebadd3 · outbound

This paper cites P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:33.815726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T12:45:27.243910Z digest=sha256:550faf6e2c79363b14265596e77a90f8cf714e9188a954290fc1c2a9cfb55e84

Observation 2ae4da01-07d6-4862-b8f6-c9fd40a544b0 · outbound

This paper cites I., and Stoica, I.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training I., and Stoica, I

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:33.696311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T12:45:27.297791Z digest=sha256:4053e195b98aa2fee960059e3d5a1dc0c3c821f73fec532fd92dbb43ae4e2f39

Observation eca310e8-d82c-44cf-aaf6-aff43d74a587 · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:33.572033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T12:45:27.353704Z digest=sha256:427436d8b7cdddde5501ddeb987f467892632775e9189f0c47d406767e9465c9

Observation e0f34c75-af43-4eb7-afbc-39fa7ab6c496 · outbound

This paper cites B., Singh, V., Lin, M., Gimelshein, N., Desmaison, A., and Yang, E.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training B., Singh, V., Lin, M., Gimelshein, N., Desmaison, A., and Yang, E

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:33.440164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T12:45:27.404939Z digest=sha256:20d9c8d9589471cbac2281da5d914162442e6b83fe9c5c5b8b0c88af6a8f60cc

Observation 9c42c152-03f6-416f-aa9e-44c0921b7bfc · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:33.302036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T12:45:27.448474Z digest=sha256:83eaed687b13632bc0c752ff7054fc868305166992e6ec96fc5fbb2d33060835

Observation e224529b-dd60-40be-972f-44b209a252f9 · outbound

This paper cites GPT-4 Technical Report.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training GPT-4 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:27.526241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:27.526241Z digest=sha256:e138033411dc51a408fde3f553eea116a14dbbb70a33055c9a37ebab3f2d3f6e

Observation 49944a65-5e2f-4e27-9d42-68d438bdb65e · outbound

This paper cites Training language models to follow instructions with human feedback.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Training language models to follow instructions with human feedback

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:27.595282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:27.595282Z digest=sha256:8bd6a6c92d60743bec27dc1d01af224b138d49c2f5cad8c7e1d8ef66dd9ce385

Observation 2169cf50-ffc1-47c5-9777-6ed001b2056f · outbound

This paper cites and et al.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training and et al

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:33.167654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T12:45:27.679104Z digest=sha256:9ba19f06d5ca1aed4b7323b52110e162a4ab77fe990632a620156df32a37876f

Observation e2749790-728d-4965-9cf0-94d6d3fc1d7c · outbound

This paper cites S., and Singh, S.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training S., and Singh, S

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:33.056767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T12:45:27.725997Z digest=sha256:82bdf891c8e41ab68b52198f016679ca5bf091ff5bf149dfea25dcace5ca2571

Observation b2f0289c-d1e7-4be0-a362-40bab67b28cf · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:32.959344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T12:45:27.813391Z digest=sha256:e130a543c25d046ba0f7beeb44c7b5f5165b05e53c2d2d444d9353270d410784

Observation 085e1425-1e98-4dbf-82c4-9f6528c057f5 · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:32.826777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T12:45:27.880640Z digest=sha256:2064758a84be00a78ec9afa9e1fa6af5e5948a481f3dc21c965be68f115b44f3

Observation 497bfe96-31bf-49da-b6e3-5dea5b5e6a9c · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:32.650731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T12:45:27.990465Z digest=sha256:bf0c3716ca6a0104324a8d0e8e54703053db891c07379eb30d43177f3af978d4

Observation 2dabefd9-2f45-4bb8-86fe-7f6d9ef369d8 · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:32.433370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T12:45:28.113659Z digest=sha256:ee939665133b27cad92217b7f891943eeb70433e04ee6b2471486474923fd151

Observation 0c3b2671-921c-4580-ba77-6cf25f206ba7 · outbound

This paper cites Proximal Policy Optimization Algorithms.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Proximal Policy Optimization Algorithms

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:28.227786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:28.227786Z digest=sha256:3fada40fec2f6f058e502c8b9a4e3aad566d18d0c15b9003ed62914e88d10f7d

Observation b69a40a9-1ab4-438c-9e68-a9c91c601b4f · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:28.335605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:28.335605Z digest=sha256:52b28f599df164710c174f64ea53089652b0972d53595e9c5d117ad9b8c419bf

Observation 84f342ad-16dc-4c18-a48c-25f7d2a1a718 · outbound

This paper cites S., Aithal, A., and Kuchaiev, O.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training S., Aithal, A., and Kuchaiev, O

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:32.250872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T12:45:28.486086Z digest=sha256:9269d234ece48212152e1709c367e5658510d26463e6ea889c6db74f75164c1f

Observation a682b063-6ddc-4529-8a0f-2a2c1fb2c606 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training HybridFlow: A Flexible and Efficient RLHF Framework

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:28.605509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:28.605509Z digest=sha256:93f2e07520eb0aa28566051ee285c333f86beb9fde06123e4e67064b978b80db

Observation 0d72c3d2-2a49-4dfb-8db5-d2060873a8e9 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:28.702500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:28.702500Z digest=sha256:a196c70c55853f43eb6baf353795c490cf3217b08626392bd0d7140479b8f0a7

Observation bd313a9f-5794-4dbb-9f22-8b1be12cc3de · outbound

This paper cites Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:28.836184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:28.836184Z digest=sha256:222c8fce81cc8a8bf143e27aca8abea613e95a40243c76a88b19e3ee741d8b83

Observation 29d07c9a-94ee-4274-a160-df653b7a835d · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:31.985189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T12:45:28.918413Z digest=sha256:5bf7a3ce07c5b9285c4c4200b07e2af8a4ebf31eaef03d7802e62a74c72d19a8

Observation 2591cb05-32c7-4872-831e-d37cb3368112 · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:31.792103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T12:45:28.973764Z digest=sha256:cd73c9c544fd99e4478acc21efdaaa8fb7b5ef2a96aca9559be1c8e1e07d4245

Observation 40392732-00a9-4837-8ca2-61d1d0092168 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Gemini: A Family of Highly Capable Multimodal Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:29.052265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:29.052265Z digest=sha256:e4fe38f7aaaa75f05bf4b6b96793ef988b3ca4e361510d2a6df82296bead6798

Observation a060fd00-2d23-4b38-929f-5030ff16084f · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:31.620568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T12:45:29.206789Z digest=sha256:32ed7dd170ed082c56de3bedcb27e4fbb7cd396e2093cca901d1e00c46342417

Observation 7239647a-83cb-4ef0-bb04-9746ced809ca · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:29.254014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:29.254014Z digest=sha256:1796d6ab57e03a0e6d70b1883d8a039ba6b3c661f44db415c96e392d07176785

Observation 0e9e40be-f8b0-41bc-9414-010bce4df230 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Solving math word problems with process- and outcome-based feedback

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:29.295485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:29.295485Z digest=sha256:f17965a3b6e04fbe202c65ba530b3443a31ddc5d0e03100c6ab2dca4d37f8c36

Observation a39cd7e0-210c-44cd-9ec5-768f76a8034e · outbound

This paper cites M., Dudzik, A., Huang, A., Georgiev, P., Powell, R., et al.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training M., Dudzik, A., Huang, A., Georgiev, P., Powell, R., et al

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:31.458469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T12:45:29.338170Z digest=sha256:498ee634bf41a5891f371dc916f714f3e078bb77bb9df32b6c56c1160f5a21a7

Observation 81515745-7875-4b0b-8e5c-5ced99f458a4 · outbound

This paper cites Sample Efficient Actor-Critic with Experience Replay.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Sample Efficient Actor-Critic with Experience Replay

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:29.387646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:29.387646Z digest=sha256:39476d4fd04f976c8460136fdac2141f08f3af366b6db388bdbe071705541574

Observation 8f639335-f82c-48d5-a3be-fdd1be02f552 · outbound

This paper cites SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:29.562365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:29.562365Z digest=sha256:03e36d1e13c4ccc8134618bf687e4bdf3f556e98ee408de150bfcdda87bb0028

Observation 9bb2f068-38ab-4814-b125-249b1d0199b6 · outbound

This paper cites An Adaptive Placement and Parallelism Framework for Accelerating RLHF Training.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training An Adaptive Placement and Parallelism Framework for Accelerating RLHF Training

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:29.683650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:29.683650Z digest=sha256:9d82a6bb39cc72fc112be1dcf433641a3aa1c1313f5533bda38e0b2735aa895a

Observation 0486456e-f3ee-40a1-9652-95381310eb25 · outbound

This paper cites A., Jin, D., Peng, K., Han, E., Nie, S., Zhu, C., Zhang, H., Zhou, W., Zeng, Z., He, Y., Mandyam, K., Talabzadeh, A., Khabsa, M., Cohen, G., Tian, Y., Ma, H., Wang, S., and Fang, H.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training A., Jin, D., Peng, K., Han, E., Nie, S., Zhu, C., Zhang, H., Zhou, W., Zeng, Z., He, Y., Mandyam, K., Talabzadeh, A., Khabsa, M., Cohen, G., Tian, Y., Ma, H., Wang, S., and Fang, H

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:31.278750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T12:45:29.826054Z digest=sha256:bb39c18bec6a9a97c39d679a7e2ef7bc8e2eb656ac4dadfcb1103bf2530aae0d

Observation 46341aad-f537-46d2-8202-9159ac88dfac · outbound

This paper cites Y., Ruwase, O., Rajbhandari, S., Wu, X., Awan, A.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Y., Ruwase, O., Rajbhandari, S., Wu, X., Awan, A

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:31.128954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T12:45:30.031519Z digest=sha256:85d0b10a2eb48f60d2eeaee0a5c99e6eb29da14d60806b22157989b0540c75b3

Observation 7010431b-99de-49de-b0f4-33e42bccc211 · outbound

This paper cites When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:30.128496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:30.128496Z digest=sha256:4e7137934422507bf45cb1471109105c44047eff7b55e098bc24f0601243491a

Observation 8648f7e9-c63d-4282-9317-11333bb0cad1 · outbound

This paper cites an unresolved cited work.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:30.901240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T12:45:30.191698Z digest=sha256:3845b37b766d6692c35e2cd82c953803b68681fcff353730b832b4f2f15bb1ec

Observation 4944aca5-0cd3-4705-aef9-659817159aaf · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Fine-Tuning Language Models from Human Preferences

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:30.237588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:30.237588Z digest=sha256:abc38f5d7b3d08e9749efe9c6d42e75717d804a3688fa411cdd25ca2cffee09d

Pith citing papers

Observation e74338b5-eb36-4ba3-884a-ddd1729123f4 · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

Reference 133

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:32:01.297713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:223d2135758b9b2d599621827ed0dbc435ed12e347d216a7a1dacdcfaa0fc18e

Observation 62aa2849-4bc7-4bce-953f-b36661f1c2ce · inbound

Magistral cites this paper.

Magistral LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:26.241394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:26.241394Z digest=sha256:6ebd9b8f505d691c4f77ba349d8b6ae7bd4eb7bc3d2e50f621e1b68e689125f6

Observation 77243c07-76f2-453d-878e-2288b1fa8aa7 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

Reference 200

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:43.979849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:43.979849Z digest=sha256:35a2442e674fcfe940b2585b5d6655d79b6b69143f416cb93e12252b707aa48c

Observation 317ba9ef-ed90-4c32-8c98-33dd0c7d27ea · inbound

HetRL: Efficient Reinforcement Learning for LLMs in Heterogeneous Environments cites this paper.

HetRL: Efficient Reinforcement Learning for LLMs in Heterogeneous Environments LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:23:36.641772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-16T22:21:26.271796Z digest=sha256:68e550b873d4297b3a970af0f0b9e07cabc4e25c5e496f2a0a7ab48d496784eb

Observation 4e20ecf2-a14a-4eec-9edc-ab8728858d75 · inbound

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training cites this paper.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:54.516468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:54.516468Z digest=sha256:142da08baa8540ef432dbb28a7de3b1070070741f46a0e33d758900f78203377

Observation ee4ba978-e8d1-489f-b615-3c4137c3991b · inbound

TensorHub: Scalable and Elastic Weight Transfer for LLM RL Training cites this paper.

TensorHub: Scalable and Elastic Weight Transfer for LLM RL Training LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:11:00.631266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T17:16:33.020610Z digest=sha256:cd5fb9d4d51f78ae8d8f484a23a1fd2a2d012a733e80b31d35a5a6aa8536a6d2

Observation e6ba9bdf-05e5-4cc8-8ccf-873d988d8145 · inbound

DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training cites this paper.

DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:46:26.425646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-07T13:49:26.560459Z digest=sha256:7674a8e16fc6d69b724b474d73b2bcace7353410397dd70befd3720460d10e45

Observation 3521626d-9bfb-47d7-924e-d54365baccd4 · inbound

When to Stop Reusing: Dynamic Gradient Gating for Sample-Efficient RLVR cites this paper.

When to Stop Reusing: Dynamic Gradient Gating for Sample-Efficient RLVR LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:33:07.364766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T07:31:23.194325Z digest=sha256:fa827601590e2d5b869e8e5c48f4086a2263b652f7fffd5bc34f28abe8228e8c

Observation 3ca7aa93-bd09-4601-9abd-9e0159acd6aa · inbound

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training cites this paper.

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:53:55.792604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-29T19:47:09.817243Z digest=sha256:c6d0ed97571d35812ba83a6ec910db9a9569d8cc9f7a78376e686d6dc98fab0a

Observation cda6d2b5-4d4e-4fcb-9c45-e253884a7721 · inbound

Libra: Efficient Resource Management for Agentic RL Post-Training cites this paper.

Libra: Efficient Resource Management for Agentic RL Post-Training LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:27.151232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T11:02:00.385932Z digest=sha256:6d38c59949bbd29cbad03b1d416a04e147dcf0c304afab1a57013e11089df144

Observation 1a77efd9-f915-4fbb-b87c-7876560eb8d5 · inbound

Rollout-Level Advantage-Prioritized Experience Replay for GRPO cites this paper.

Rollout-Level Advantage-Prioritized Experience Replay for GRPO LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-07-02T06:06:41.077555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-28T07:43:21.284574Z digest=sha256:d39c88be4627c941d70525a89e6508a5727966af8685f150b574bb5c86bc5c28

Observation 9394f3ea-be4d-41d8-a70f-3022c8445dcc · inbound

AsyncWebRL: Efficient Asynchronous Reinforcement Learning for Multi-Step Visual Web Agents cites this paper.

AsyncWebRL: Efficient Asynchronous Reinforcement Learning for Multi-Step Visual Web Agents LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T11:56:55.543867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-28T02:45:25.694378Z digest=sha256:b15137e496626032c11bf80b79d07f977e5b2c16b3e90d9ba2afb14f05eaf7e5

Observation 2f99b87d-9796-49cc-ab6b-9a492c5824bc · inbound

Sparrow: Sparse Rollout for Stable and Efficient Long-context RL of Large Language Models cites this paper.

Sparrow: Sparse Rollout for Stable and Efficient Long-context RL of Large Language Models LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:07:26.505309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T19:10:53.882876Z digest=sha256:bb630407be8e1d1c4ed504c9ce4e30f7f86db44666de58e3480026a93eda1275

Observation a9cd52f0-5201-422d-9aed-8d6667281aaa · inbound

Harnessing Routing Foresight for Micro-step-level MoE load balancing in RL Post-training cites this paper.

Harnessing Routing Foresight for Micro-step-level MoE load balancing in RL Post-training LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-07-03T12:58:08.711158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T08:35:16.435272Z digest=sha256:3aeebebb19bbf2a27d3d08f7b32a6b9088e22d8a8467bb554c9099c3ecf22bdc