Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:01:18.613783Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 19 inbound Pith citation observations for arXiv:2508.18588.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:01:18.613783Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T06:25:47.766365Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T23:59:07.264671Z
64 of 64 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 09b5210d-2431-4060-884d-662f8a443ee6 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 962fb6bb-b756-4e97-9fb0-671857785666 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11054d9b-a0c9-4516-940f-1452bc232d41 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL The Llama 4 herd: The beginning of a new era of natively multi- modal AI innovation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 73efe1df-856c-4a05-a50f-4d380d5cd11e · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Qwen3 Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f000dd27-7ef6-4e91-a3ac-13503b49ba18 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Introducing Claude 4
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ee676022-ac45-4145-915d-cf752048e605 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7e528cb-879c-46a8-a1fc-0d1125c83759 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acafe9ae-355b-41a2-afab-49b0e901551d · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b31c13fb-7e25-4bed-a014-f516bd5f5501 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL OmniScience: A Domain-Specialized LLM for Scientific Reasoning and Discovery
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c738fb90-9627-439f-a3ee-726a5cca853c · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3756319d-1e52-411c-a6bf-d56b8ba7f9f5 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL s1: Simple test-time scaling
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14130940-3e90-4133-9b88-5ce88ae0932b · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80bb81a3-1e60-4dfb-bfd4-00af83cf5d66 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b776967d-8ba7-4e65-9848-4991b5bf8127 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f41e1d7-81c2-45b0-86d8-93f1976f96a8 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aac7fa0f-cd44-4eba-a046-58145567403b · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4786741e-81d0-4753-bc35-b021e4cad92c · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3cf7ff51-5593-496d-9db0-c5d9daf3abdd · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Group Sequence Policy Optimization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 030be016-4478-4e41-a069-fb8e1bfdc2c7 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e11b492d-cb3b-42b1-bf47-e78e88b309f2 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 247edac1-dbc8-42fa-a1dc-69802efbb0a4 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Yang and S.S
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4c57d2be-43ad-4385-b370-de869db141d8 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL PPO-Clip Attains Global Optimality: Towards Deeper Understandings of Clipping
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0138876a-2fb8-4d45-9769-51914f4b9534 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Introducing OpenAI o3 and o4-mini
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 24a89868-a17f-41d6-914f-3927e251c735 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Proximal Policy Optimization Algorithms
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee0388c9-eaae-4ec1-936b-54f95ae897e5 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b99ed55c-507d-4907-a713-8fb4e01f8194 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL OpenAI o1 System Card
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8db97eec-7c87-48ba-960d-a16195fde901 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3ed52654-9418-4ecf-ab9c-981798faf910 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Accelerating Large Language Model Decoding with Speculative Sampling
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3cb22e8-5286-4ac7-8098-d836ac7b3805 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Reward-Guided Speculative Decoding for Efficient LLM Reasoning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59646347-f248-406f-9b85-7d2230e9ab08 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 994ae41b-b2c4-49ab-9fd9-4e23d69028e4 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76ff464d-45b4-4fc6-9e17-465f801219b7 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d1f88e7-fa6a-4d3d-b185-58845e5ca3fd · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc09ac30-e460-4744-8903-82567dc2a135 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Reinforcement Speculative Decoding for Fast Ranking
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2e055d5-7e57-4212-a186-cb742e2f172e · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f1a8af8-3006-4a18-aa60-5d2206b65592 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7715227a-f736-460a-a545-57690ef20e3a · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Inference with Reference: Lossless Acceleration of Large Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6376decf-b9cd-4938-a136-83dc36767ddd · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 464dfc39-9d4f-4cd4-8c27-e1b833bf47eb · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2595c532-af99-4371-a7c1-853590ddaea4 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 472291bc-bbee-4d96-9c07-368ed5ed0133 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb239698-3976-4af1-b306-56be931776ed · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL (vLLM official blog) How Speculative Decoding Boosts vLLM Performance by up to 2.8x
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c117057e-0673-4c3c-9bd1-4a8301674a69 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5e66320c-69e9-423a-8c8a-fb8a9b1eb240 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL NeMo: A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c4c4a0f8-aea0-47bf-88f6-90bbfa621a69 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 582ee1c2-835a-4bb7-8281-dfa8677d1212 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL DistFlow: A Fully Distributed RL Framework for Scalable and Efficient LLM Post-Training
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6e4cc92a-6724-4872-b700-9a21e2766e6c · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL (veRL) One Step Off Policy Async Trainer
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation beab0c3e-fd0d-47c6-b39d-4e3b15a74a52 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a23c2ab-44f8-4297-90f6-cfee82c50245 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Gonzalez, Clark Barrett, and Ying Sheng
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 92bf3a87-cb62-4e37-a6f4-be4d75aed8db · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c2c199d-dde7-4001-af01-3a9817678fb5 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e470b97c-70e9-4180-b2f0-79c04a16952b · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 680bb46c-3e39-46b3-a02d-e1cff1f2118f · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL cgroups - Linux control groups
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 795c19e9-8381-4710-ab5f-913542030f75 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL (vLLM documentation) vLLm Speculative Decoding
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e905edad-db00-483d-99ab-38248151b648 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Qwen2.5 Technical Report
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a125e698-9a3e-428d-9355-b2bb24d4bcd9 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL The Llama 3 Herd of Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7412a422-ffa9-4bc6-9282-1440f7c572bf · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c9ad1c60-2e67-44aa-8047-c8f27f2c60d4 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL DeepSeek-V3 Technical Report
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d7b3aaa-1ced-4db7-981f-dded3ed14b1a · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL (vLLM issue) vLLM Eagle performance is worse than expected
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 55dfeec6-392b-465e-bc23-354fba9680f4 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Scaling Laws for Neural Language Models
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d89ec082-cda3-49d3-bc1d-74e6ee95cfdd · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work
Reference 198
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4742bee0-995a-424f-a194-e9c396e56fa4 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL In Computer Architecture (ISCA), 2017 ACM/IEEE 44th Annual International Symposium on
Reference 2017
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 273d52c8-3e9f-48e0-991c-79c04ac01104 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbbd30f6-11cc-451c-9354-6784853f4804 · outbound
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Kimi K2: Open Agentic Intelligence
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3800836e-76ec-4e92-a022-d6015b1b7204 · inbound
RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8c2c3fe3-f660-4953-838a-d78764c03ecb · inbound
Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a46dc5c9-5719-4b7b-b5d7-0787bff3f378 · inbound
StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96d1c965-bc1d-46f9-bcf6-a69040c52b5e · inbound
ECHO-2: A Large-Scale Distributed Rollout Framework for Cost-Efficient Reinforcement Learning History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11ab2c75-bcae-42df-a493-f22c02e03898 · inbound
TensorHub: Scalable and Elastic Weight Transfer for LLM RL Training History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ff08c0c9-878f-4727-9471-72957cf6f468 · inbound
DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 104cac29-2996-4b38-bdab-8afba1fa7e39 · inbound
ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2cbc7f6b-7da4-453e-aaa2-1a48b797e015 · inbound
ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 83ad0f94-1742-4f74-8d6c-1ba0c29dea2f · inbound
Gradient Extrapolation-Based Policy Optimization History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d1022d61-df46-4f6c-ae73-1317e9940e9f · inbound
Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ea386b91-1233-496f-a36f-0bf3911ab6ef · inbound
BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 77dc928c-00a6-4bbd-84a0-7c3ee9de34c7 · inbound
AIS: Adaptive Importance Sampling for Quantized RL History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 69499fa6-ac48-494e-8c0b-ac469511e626 · inbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 073c2ca5-8078-4af1-94ba-be90e8adc967 · inbound
How Off-Policy Can GRPO Be? Mu-GRPO for Efficient LLM Reinforcement Learning History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e0c15d7d-b9b1-4f87-aee3-ca0310a88838 · inbound
Libra: Efficient Resource Management for Agentic RL Post-Training History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e157f5db-02dd-4ea5-8329-034ddbea5d8c · inbound
Harnessing Routing Foresight for Micro-step-level MoE load balancing in RL Post-training History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6ef0ae08-8da4-46f6-abff-172ecf02e0b6 · inbound
EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4dd1f559-4df5-4ea8-9935-c44e2eaa6a7e · inbound
Bidirectional Resource Scheduling for Disaggregated and Asynchronous RL Post-Training History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f176268-ebbb-498f-a953-8ba9718194a7 · inbound
QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.