Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:45:30.237588Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 14 inbound Pith citation observations for arXiv:2505.24034.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:45:30.237588Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:19:26.241394Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
74 of 74 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 07a64a7c-f311-4425-bb11-eeedc2081e68 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5de36d5c-a3cd-4aa9-90c9-651ae928d77b · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d22a7033-1274-48b7-8a45-d975a962a426 · outbound
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db16a862-1b80-41a8-9618-463fc606670b · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ead16030-8ce3-407b-bd40-d766dac12bad · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 45ff0517-193d-4dfa-b1a2-c1db99266a2e · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f59b1836-af27-4c24-b324-e07d94e6bd63 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Claude: Training helpful and harmless ai assistants
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 989c91bd-78b9-4040-b10a-bb580ba78bd7 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a392faf4-81bc-4a94-be1f-07e9a06b1d61 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Dota 2 with Large Scale Deep Reinforcement Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0217a153-b800-41d8-b2dc-6a03a2008b5f · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9eca5e5f-1329-4d86-8666-7011d054086a · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e5c9a32b-14a1-4940-bb45-0c8e9d9623c9 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b94aa1bd-1675-437b-b240-bc9492c67b0d · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Training Verifiers to Solve Math Word Problems
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11693588-6706-486d-be96-698d8782eec0 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5e39292-d015-48f7-834a-6d8d9c80605f · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e708a75c-3b4d-428f-943b-9da9ad7ccb01 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Competitive Programming with Large Reasoning Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b256365c-8f1c-482a-b6ee-372103d1a48e · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1aff602b-b69d-462f-8792-7d66865a4e5d · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41064d76-6cec-4431-99ca-ac9028e20ae3 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Bard: Conversational ai by google
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 73fccd74-5a99-4ea2-acf7-9c161fd76513 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f7c0837-28d7-4e25-9eba-81eca24608ef · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Measuring Mathematical Problem Solving With the MATH Dataset
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2128733f-0124-4d36-8e3d-4cdc5e9d507b · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Deep Learning Scaling is Predictable, Empirically
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e19f7d38-bb19-49ea-a33f-163005cb279f · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Training Compute-Optimal Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59cc2812-bd55-49a7-89d5-b1bd41a21eeb · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c25dc515-d9e1-49fa-a8cb-a386905f912e · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training OpenAI o1 System Card
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d545a37-25e9-4084-93c2-11c1aa9bb837 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training and Abbeel, P
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 52141ded-4384-4fbb-bf54-6370bb16020e · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training and Langford, J
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c25ab529-3f3d-466c-8630-6f9e046841cd · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Scaling Laws for Neural Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3000b53c-46e7-4d71-b723-93897e3a6536 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5139bbd4-11cc-4431-be00-5e305202204e · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d97c8bf-a410-47c1-abd8-07a60b8739b7 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ed95391-f60b-442b-a38c-da2cd65d4524 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Adam: A Method for Stochastic Optimization
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85c1a2e6-39ea-4fae-a9e7-0fe0007ea9c1 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation baf38e29-692f-47ad-a007-0a6dfa798140 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training G., Park, J
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8d3eb950-d117-4390-b03a-00a18af5e0d2 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 629e7fe6-a2ae-4098-9814-c7fc615ab6e3 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b2b022ba-afae-4eea-b518-3aae00c6c6f1 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Let's Verify Step by Step
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 759d7f3d-bc02-4068-8885-51aa739c7f6e · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training The Llama 3 Herd of Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ac4ad63-b16f-4bfc-9372-3147a3f164c1 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Decoupled Weight Decay Regularization
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 753d3c29-85d2-4bb2-8f2b-8768ebdde0c3 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training P., Paprocki, M., C ert\' i k, O., Kirpichev, S
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b625d55-d532-4228-b7c6-1d331bebadd3 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2ae4da01-07d6-4862-b8f6-c9fd40a544b0 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training I., and Stoica, I
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation eca310e8-d82c-44cf-aaf6-aff43d74a587 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e0f34c75-af43-4eb7-afbc-39fa7ab6c496 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training B., Singh, V., Lin, M., Gimelshein, N., Desmaison, A., and Yang, E
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9c42c152-03f6-416f-aa9e-44c0921b7bfc · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e224529b-dd60-40be-972f-44b209a252f9 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training GPT-4 Technical Report
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49944a65-5e2f-4e27-9d42-68d438bdb65e · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Training language models to follow instructions with human feedback
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2169cf50-ffc1-47c5-9777-6ed001b2056f · outbound
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e2749790-728d-4965-9cf0-94d6d3fc1d7c · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training S., and Singh, S
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b2f0289c-d1e7-4be0-a362-40bab67b28cf · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 085e1425-1e98-4dbf-82c4-9f6528c057f5 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 497bfe96-31bf-49da-b6e3-5dea5b5e6a9c · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2dabefd9-2f45-4bb8-86fe-7f6d9ef369d8 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0c3b2671-921c-4580-ba77-6cf25f206ba7 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Proximal Policy Optimization Algorithms
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b69a40a9-1ab4-438c-9e68-a9c91c601b4f · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84f342ad-16dc-4c18-a48c-25f7d2a1a718 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training S., Aithal, A., and Kuchaiev, O
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a682b063-6ddc-4529-8a0f-2a2c1fb2c606 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training HybridFlow: A Flexible and Efficient RLHF Framework
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d72c3d2-2a49-4dfb-8db5-d2060873a8e9 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd313a9f-5794-4dbb-9f22-8b1be12cc3de · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29d07c9a-94ee-4274-a160-df653b7a835d · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2591cb05-32c7-4872-831e-d37cb3368112 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 40392732-00a9-4837-8ca2-61d1d0092168 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Gemini: A Family of Highly Capable Multimodal Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a060fd00-2d23-4b38-929f-5030ff16084f · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7239647a-83cb-4ef0-bb04-9746ced809ca · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e9e40be-f8b0-41bc-9414-010bce4df230 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Solving math word problems with process- and outcome-based feedback
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a39cd7e0-210c-44cd-9ec5-768f76a8034e · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training M., Dudzik, A., Huang, A., Georgiev, P., Powell, R., et al
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 81515745-7875-4b0b-8e5c-5ced99f458a4 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Sample Efficient Actor-Critic with Experience Replay
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f639335-f82c-48d5-a3be-fdd1be02f552 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bb2f068-38ab-4814-b125-249b1d0199b6 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training An Adaptive Placement and Parallelism Framework for Accelerating RLHF Training
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0486456e-f3ee-40a1-9652-95381310eb25 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training A., Jin, D., Peng, K., Han, E., Nie, S., Zhu, C., Zhang, H., Zhou, W., Zeng, Z., He, Y., Mandyam, K., Talabzadeh, A., Khabsa, M., Cohen, G., Tian, Y., Ma, H., Wang, S., and Fang, H
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 46341aad-f537-46d2-8202-9159ac88dfac · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Y., Ruwase, O., Rajbhandari, S., Wu, X., Awan, A
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7010431b-99de-49de-b0f4-33e42bccc211 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8648f7e9-c63d-4282-9317-11333bb0cad1 · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Unresolved cited work
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4944aca5-0cd3-4705-aef9-659817159aaf · outbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Fine-Tuning Language Models from Human Preferences
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e74338b5-eb36-4ba3-884a-ddd1729123f4 · inbound
Reinforcement Learning from Human Feedback LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training
Reference 133
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 62aa2849-4bc7-4bce-953f-b36661f1c2ce · inbound
Magistral LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77243c07-76f2-453d-878e-2288b1fa8aa7 · inbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training
Reference 200
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 317ba9ef-ed90-4c32-8c98-33dd0c7d27ea · inbound
HetRL: Efficient Reinforcement Learning for LLMs in Heterogeneous Environments LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4e20ecf2-a14a-4eec-9edc-ab8728858d75 · inbound
StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee4ba978-e8d1-489f-b615-3c4137c3991b · inbound
TensorHub: Scalable and Elastic Weight Transfer for LLM RL Training LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e6ba9bdf-05e5-4cc8-8ccf-873d988d8145 · inbound
DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3521626d-9bfb-47d7-924e-d54365baccd4 · inbound
When to Stop Reusing: Dynamic Gradient Gating for Sample-Efficient RLVR LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3ca7aa93-bd09-4601-9abd-9e0159acd6aa · inbound
Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cda6d2b5-4d4e-4fcb-9c45-e253884a7721 · inbound
Libra: Efficient Resource Management for Agentic RL Post-Training LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1a77efd9-f915-4fbb-b87c-7876560eb8d5 · inbound
Rollout-Level Advantage-Prioritized Experience Replay for GRPO LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9394f3ea-be4d-41d8-a70f-3022c8445dcc · inbound
AsyncWebRL: Efficient Asynchronous Reinforcement Learning for Multi-Step Visual Web Agents LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2f99b87d-9796-49cc-ab6b-9a492c5824bc · inbound
Sparrow: Sparse Rollout for Stable and Efficient Long-context RL of Large Language Models LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a9cd52f0-5201-422d-9aed-8d6667281aaa · inbound
Harnessing Routing Foresight for Micro-step-level MoE load balancing in RL Post-training LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.