Pith. sign in

Paper Citation Record · LEDGER

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL

As of 18 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 19 inbound Pith citation observations for arXiv:2508.18588.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.18588 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:01:18.613783Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T06:25:47.766365Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T23:59:07.264671Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 09b5210d-2431-4060-884d-662f8a443ee6 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.312585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.312585Z digest=sha256:2b4fef99495a31ce6fdd5d7280b4eee07fbdbe175e325e42f1820c0e6220fab7

Observation 962fb6bb-b756-4e97-9fb0-671857785666 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.319025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.319025Z digest=sha256:57e5a895f6df638d027509d37b9a41d73a5b634200f02a5f6f76b75cf3986bfa

Observation 11054d9b-a0c9-4516-940f-1452bc232d41 · outbound

This paper cites The Llama 4 herd: The beginning of a new era of natively multi- modal AI innovation.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL The Llama 4 herd: The beginning of a new era of natively multi- modal AI innovation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:01:20.085860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:01:18.324179Z digest=sha256:ebb264a9bfea896b906ce98276431923706b3bbf0dc3ece0783a08e1f90abe51

Observation 73efe1df-856c-4a05-a50f-4d380d5cd11e · outbound

This paper cites Qwen3 Technical Report.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Qwen3 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.329164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.329164Z digest=sha256:35f2df1504824c0353e77f54071acea2d7c3d476f8ffe4a9d5a683f5a6dcd661

Observation f000dd27-7ef6-4e91-a3ac-13503b49ba18 · outbound

This paper cites Introducing Claude 4.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Introducing Claude 4

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:01:20.068728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:01:18.334164Z digest=sha256:8e489325cc385c2fb800b9c312e34077b02a528384e86fc2b4f9d6f106780ed6

Observation ee676022-ac45-4145-915d-cf752048e605 · outbound

This paper cites an unresolved cited work.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.338827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.338827Z digest=sha256:9c7fdb7e35090bd58c07724acba2e4bf861cfbd9098d7ff45da2eb82e989733e

Observation e7e528cb-879c-46a8-a1fc-0d1125c83759 · outbound

This paper cites DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.347926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.347926Z digest=sha256:22043bcf026f00b4be58b3b0e9e08b8ee4a03dbedf28d9baeb867cf1e0e93dbc

Observation acafe9ae-355b-41a2-afab-49b0e901551d · outbound

This paper cites Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.353364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.353364Z digest=sha256:c422ffc90b75294b33cf60a119e7c0599ba206f74a95fa721a6e34d46076fde5

Observation b31c13fb-7e25-4bed-a014-f516bd5f5501 · outbound

This paper cites OmniScience: A Domain-Specialized LLM for Scientific Reasoning and Discovery.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL OmniScience: A Domain-Specialized LLM for Scientific Reasoning and Discovery

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.358564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.358564Z digest=sha256:9471ab4c76bfe73f80a344dd6ad3e09c1cbdc5a32b6b33b57de599c7b6e22931

Observation c738fb90-9627-439f-a3ee-726a5cca853c · outbound

This paper cites SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.363452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.363452Z digest=sha256:dad144bcc1863e6151c7ae9dab38188225b15a832533bea2b4c1c259d1bcf4bf

Observation 3756319d-1e52-411c-a6bf-d56b8ba7f9f5 · outbound

This paper cites s1: Simple test-time scaling.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL s1: Simple test-time scaling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.368084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.368084Z digest=sha256:7e886a32f84b410bdc98f7319ef91030161cfeb47003641316469bb2ce7368aa

Observation 14130940-3e90-4133-9b88-5ce88ae0932b · outbound

This paper cites an unresolved cited work.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.372700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.372700Z digest=sha256:e8bd6a0412279d19fa1c0a622141af84ce99a66154c82cd4c5e01322054c9151

Observation 80bb81a3-1e60-4dfb-bfd4-00af83cf5d66 · outbound

This paper cites an unresolved cited work.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:01:19.923104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:01:18.377372Z digest=sha256:e15f81313d545e3385e4e7fdbdc3d94b0b59303cccb5dd334167c476f92dfdd2

Observation b776967d-8ba7-4e65-9848-4991b5bf8127 · outbound

This paper cites StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.387139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.387139Z digest=sha256:2f8d6238e0e289935822a35edc36a64ed7a8d4709a303b746e6885653cde154d

Observation 6f41e1d7-81c2-45b0-86d8-93f1976f96a8 · outbound

This paper cites AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.391862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.391862Z digest=sha256:9e4848acc211c1e9e6db6f99d6074aa8c56c9526e94d0594c8465d307d4b20b9

Observation aac7fa0f-cd44-4eba-a046-58145567403b · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.396807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.396807Z digest=sha256:5debb4bbfec2c761184edf372cfaf835910ef3fbb1fea30af71e47fd3df22a29

Observation 4786741e-81d0-4753-bc35-b021e4cad92c · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:01:19.906568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:01:18.403335Z digest=sha256:79f62be70c00c90f0c12f48729fc2239e0791cf494d7502f8934426dfeb9c5a2

Observation 3cf7ff51-5593-496d-9db0-c5d9daf3abdd · outbound

This paper cites Group Sequence Policy Optimization.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Group Sequence Policy Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.407912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.407912Z digest=sha256:1f4859c29756f3efe39d72ed9bde36aa11b78365cf82b3892d2ee41b8af9959c

Observation 030be016-4478-4e41-a069-fb8e1bfdc2c7 · outbound

This paper cites an unresolved cited work.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:01:19.887336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:01:18.412621Z digest=sha256:0923a67b081923d65eecf40756d840bf2d4fc20e5b6a81cfa39c3dc5f6539018

Observation e11b492d-cb3b-42b1-bf47-e78e88b309f2 · outbound

This paper cites an unresolved cited work.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:01:19.868700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:01:18.422340Z digest=sha256:fde5d34247d9401d9d2d28a869502b694fcf71db553a2a0b0a580afbd9b85631

Observation 247edac1-dbc8-42fa-a1dc-69802efbb0a4 · outbound

This paper cites Yang and S.S.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Yang and S.S

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:01:19.850999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:01:18.427035Z digest=sha256:24a724be61611c3949eeacdc0ea6f2efb81f63a1aee994cd24dcbff8c92d941a

Observation 4c57d2be-43ad-4385-b370-de869db141d8 · outbound

This paper cites PPO-Clip Attains Global Optimality: Towards Deeper Understandings of Clipping.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL PPO-Clip Attains Global Optimality: Towards Deeper Understandings of Clipping

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T17:01:19.373396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:01:18.417145Z digest=sha256:7db74063cc3b7efcc49610f6f78a5bbd007cf292792f878994b4cf4cd2f5b060

Observation 0138876a-2fb8-4d45-9769-51914f4b9534 · outbound

This paper cites Introducing OpenAI o3 and o4-mini.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Introducing OpenAI o3 and o4-mini

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:01:19.834556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:01:18.441077Z digest=sha256:04060b458e471bdcb0d8b35e9b59f939e31f3551bb2c65430ebe7d3bea3637b8

Observation 24a89868-a17f-41d6-914f-3927e251c735 · outbound

This paper cites Proximal Policy Optimization Algorithms.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Proximal Policy Optimization Algorithms

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.445579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.445579Z digest=sha256:04a234317356682d8e9493d24053d01ab364b0f9a2249f06409ef1f45040e64f

Observation ee0388c9-eaae-4ec1-936b-54f95ae897e5 · outbound

This paper cites an unresolved cited work.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:01:19.818596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:01:18.450347Z digest=sha256:82c17b7f16dc635b4c6eab0d42c67e2e86316d0ffe9ae892ba82593108b240fd

Observation b99ed55c-507d-4907-a713-8fb4e01f8194 · outbound

This paper cites OpenAI o1 System Card.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL OpenAI o1 System Card

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.436172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.436172Z digest=sha256:acee3dd37622ebe5aa4c6b6dc40702fe5dd6f26c55e48aae9119119e40025da8

Observation 8db97eec-7c87-48ba-960d-a16195fde901 · outbound

This paper cites an unresolved cited work.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:01:19.802498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:01:18.459458Z digest=sha256:56c2a4f66c83cf1a4f503fa3abdc8bbab6d441f31fb3681ac39670bac6b8f9bc

Observation 3ed52654-9418-4ecf-ab9c-981798faf910 · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Accelerating Large Language Model Decoding with Speculative Sampling

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.463990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.463990Z digest=sha256:f5670b55d9f9133a396ca97b12248d720d125ee91a8a22866f4814b760918415

Observation d3cb22e8-5286-4ac7-8098-d836ac7b3805 · outbound

This paper cites Reward-Guided Speculative Decoding for Efficient LLM Reasoning.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Reward-Guided Speculative Decoding for Efficient LLM Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.468677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.468677Z digest=sha256:75b099f23d0be5dfb702418beacaf790fde83a111f87ff77a4bc3e92585b767f

Observation 59646347-f248-406f-9b85-7d2230e9ab08 · outbound

This paper cites an unresolved cited work.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.454865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.454865Z digest=sha256:0375ead73e7b7e5596811bb165b4bd25ba10a4396ebcce7807dc1f85db4c690d

Observation 994ae41b-b2c4-49ab-9fd9-4e23d69028e4 · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.478489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.478489Z digest=sha256:0ac64de033b15cfe282a719073608b51d67178e1727b5ed65b97df5acffb45ad

Observation 76ff464d-45b4-4fc6-9e17-465f801219b7 · outbound

This paper cites EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.483193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.483193Z digest=sha256:cd1b6ecff7feabe78f8d3d6bd71aa00845dab5e740b323e8ac6ccf0079449aa1

Observation 8d1f88e7-fa6a-4d3d-b185-58845e5ca3fd · outbound

This paper cites EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.487879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.487879Z digest=sha256:baf467f249a93c8d11a69ca7cee2037096359c139e60edc8092ff9b2f0a9fbae

Observation fc09ac30-e460-4744-8903-82567dc2a135 · outbound

This paper cites Reinforcement Speculative Decoding for Fast Ranking.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Reinforcement Speculative Decoding for Fast Ranking

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.473636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.473636Z digest=sha256:550ae966d7df51a76ea0abb9cf7274d9cbf5884d2e555234eece0e4beb60448a

Observation d2e055d5-7e57-4212-a186-cb742e2f172e · outbound

This paper cites an unresolved cited work.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.497503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.497503Z digest=sha256:9bd84cfd40bc3cc3cb87c1f67a97a430d49f0ed5186fd0bec31e921ba3a9c527

Observation 1f1a8af8-3006-4a18-aa60-5d2206b65592 · outbound

This paper cites an unresolved cited work.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.502202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.502202Z digest=sha256:ed6eda39ca983736a6d12bccf45efa4e4a18b50ffae189acacd8354a9b575291

Observation 7715227a-f736-460a-a545-57690ef20e3a · outbound

This paper cites Inference with Reference: Lossless Acceleration of Large Language Models.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Inference with Reference: Lossless Acceleration of Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.506626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.506626Z digest=sha256:9f781302b349b7bd71738f04d2518974b3bfef33e0dd3b72cc68679c94d204d1

Observation 6376decf-b9cd-4938-a136-83dc36767ddd · outbound

This paper cites EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.492656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.492656Z digest=sha256:f238605c797851d25d71cd81dc54ddddbee5c9f647fd6eb8a6347ea5259bf4f9

Observation 464dfc39-9d4f-4cd4-8c27-e1b833bf47eb · outbound

This paper cites an unresolved cited work.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.516099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.516099Z digest=sha256:1dbe07e292bec1b3e0dbe8ea240ebeeaf86f857741d4767b8f30a2b24d3c2b85

Observation 2595c532-af99-4371-a7c1-853590ddaea4 · outbound

This paper cites an unresolved cited work.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:01:19.760002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:01:18.520618Z digest=sha256:ab114b5ce7548e7a1d75d904235dd4d48f00d58796d24509e056c194e16567bc

Observation 472291bc-bbee-4d96-9c07-368ed5ed0133 · outbound

This paper cites AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.525104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.525104Z digest=sha256:7bcbd61dccbebbddc11403d0e57cf71a522a5c832677bb2d81ad907ea9ddaad2

Observation eb239698-3976-4af1-b306-56be931776ed · outbound

This paper cites (vLLM official blog) How Speculative Decoding Boosts vLLM Performance by up to 2.8x.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL (vLLM official blog) How Speculative Decoding Boosts vLLM Performance by up to 2.8x

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:01:19.775713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:01:18.511680Z digest=sha256:ee5e2128742dcae70823751e1585e7771cbc15248e52b11c3fb3c4ca46c8c046

Observation c117057e-0673-4c3c-9bd1-4a8301674a69 · outbound

This paper cites an unresolved cited work.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:01:19.743931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:01:18.534523Z digest=sha256:2363f289efdbfe51ef9aa7fd3c2d29f8c1c794fe3a304a1941c13953faff0df9

Observation 5e66320c-69e9-423a-8c8a-fb8a9b1eb240 · outbound

This paper cites NeMo: A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL NeMo: A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:01:19.728625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:01:18.538879Z digest=sha256:b3ba5691cb13d0f96de8ebb3bd5e2835924091cb3cc7dae1db4eee01b384c656

Observation c4c4a0f8-aea0-47bf-88f6-90bbfa621a69 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.543506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.543506Z digest=sha256:226d37e2c377190c36bcfd6e92ac2421e043be0aa5a770ee563d1fcfef4704ac

Observation 582ee1c2-835a-4bb7-8281-dfa8677d1212 · outbound

This paper cites DistFlow: A Fully Distributed RL Framework for Scalable and Efficient LLM Post-Training.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL DistFlow: A Fully Distributed RL Framework for Scalable and Efficient LLM Post-Training

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-15T17:01:18.902703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:01:18.529888Z digest=sha256:909b9ae9ee18a8b7f47df4fa6ef14e2be25b2b18719d8ceefdfb9d4d6b2431bb

Observation 6e4cc92a-6724-4872-b700-9a21e2766e6c · outbound

This paper cites (veRL) One Step Off Policy Async Trainer.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL (veRL) One Step Off Policy Async Trainer

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:01:19.713044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:01:18.553428Z digest=sha256:9bb4da86a019956e6ca21dc5951df7a906e15b6e70016e03864398b1595b8db7

Observation beab0c3e-fd0d-47c6-b39d-4e3b15a74a52 · outbound

This paper cites an unresolved cited work.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.558899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.558899Z digest=sha256:3d7561ef47e4d8f9995cab222a5beacf448b1cdc623a29a1d532369384f77a76

Observation 7a23c2ab-44f8-4297-90f6-cfee82c50245 · outbound

This paper cites Gonzalez, Clark Barrett, and Ying Sheng.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Gonzalez, Clark Barrett, and Ying Sheng

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:01:19.696011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:01:18.563549Z digest=sha256:1aee9a7bfbe40cf938450aecff9817a3c76dbf50b83ab79639f3d03497b271fb

Observation 92bf3a87-cb62-4e37-a6f4-be4d75aed8db · outbound

This paper cites Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.548778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.548778Z digest=sha256:9ee30d42afd4b42b043345e35ecf6dc605ef75900b21b7656ace2153c41b3def

Observation 0c2c199d-dde7-4001-af01-3a9817678fb5 · outbound

This paper cites an unresolved cited work.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.572899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.572899Z digest=sha256:6ffdd2b3077b56aab106e6e37547e4c73f272bb4816fa5541b2dafab1b4ffeb3

Observation e470b97c-70e9-4180-b2f0-79c04a16952b · outbound

This paper cites an unresolved cited work.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.577877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.577877Z digest=sha256:8b0e2732387a4de2231e7ed66d3aa8e97b007300d033fcab22b844f568a278b4

Observation 680bb46c-3e39-46b3-a02d-e1cff1f2118f · outbound

This paper cites cgroups - Linux control groups.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL cgroups - Linux control groups

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:01:19.663794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:01:18.582278Z digest=sha256:8eeb89f014fdd0880e3eb6004124d08ff13ee83b035a3af582ebc96ef71f0ffc

Observation 795c19e9-8381-4710-ab5f-913542030f75 · outbound

This paper cites (vLLM documentation) vLLm Speculative Decoding.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL (vLLM documentation) vLLm Speculative Decoding

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:01:19.679324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:01:18.568388Z digest=sha256:e11efe80eeb3dce6a7af4dcfc6dee5b749a907907e907af6bc86883e60c6053d

Observation e905edad-db00-483d-99ab-38248151b648 · outbound

This paper cites Qwen2.5 Technical Report.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Qwen2.5 Technical Report

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.591066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.591066Z digest=sha256:10fbb16e4e818fd2e45691ce54a720435b72faf57405e90905f398c4791eb188

Observation a125e698-9a3e-428d-9355-b2bb24d4bcd9 · outbound

This paper cites The Llama 3 Herd of Models.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL The Llama 3 Herd of Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.595473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.595473Z digest=sha256:3239aacf2379003b3bbe4d0209c80b75523bf05e0cf7f65d5248112408b873d8

Observation 7412a422-ffa9-4bc6-9282-1440f7c572bf · outbound

This paper cites an unresolved cited work.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:01:19.648516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:01:18.600019Z digest=sha256:2ded1ba3babccb1c01796a804108af31557df1c0c0b4a565f51fe0bd650a3176

Observation c9ad1c60-2e67-44aa-8047-c8f27f2c60d4 · outbound

This paper cites DeepSeek-V3 Technical Report.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL DeepSeek-V3 Technical Report

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.586778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.586778Z digest=sha256:40dac321592782fef9e5a896aca12a41911034866ce2cd52e4b92881ace9d008

Observation 4d7b3aaa-1ced-4db7-981f-dded3ed14b1a · outbound

This paper cites (vLLM issue) vLLM Eagle performance is worse than expected.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL (vLLM issue) vLLM Eagle performance is worse than expected

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:01:19.617101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:01:18.613783Z digest=sha256:9d0f3438a697445d8271d5e5a013e5b626bfbb0bc106a93e16f14a34ca832781

Observation 55dfeec6-392b-465e-bc23-354fba9680f4 · outbound

This paper cites Scaling Laws for Neural Language Models.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Scaling Laws for Neural Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.608944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.608944Z digest=sha256:a573b92472f0df8d61089c04fd7440083bfe4cade4015ac0a6977d270e12803d

Observation d89ec082-cda3-49d3-bc1d-74e6ee95cfdd · outbound

This paper cites an unresolved cited work.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Unresolved cited work

Reference 198

Resolution
metadata mismatch
raw_fallback, observed 2026-08-15T17:01:19.350667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:01:18.431560Z digest=sha256:0de35b2b51ca1d98d3b68639ec0478cdc2ad8bfb442e23e3498f4196f749db9c

Observation 4742bee0-995a-424f-a194-e9c396e56fa4 · outbound

This paper cites In Computer Architecture (ISCA), 2017 ACM/IEEE 44th Annual International Symposium on.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL In Computer Architecture (ISCA), 2017 ACM/IEEE 44th Annual International Symposium on

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:01:19.632932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:01:18.604438Z digest=sha256:b7513724f6301321af9ac217ccd7af03e17d943e3873f1e808239d5c6fbd980d

Observation 273d52c8-3e9f-48e0-991c-79c04ac01104 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.343360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.343360Z digest=sha256:024b2f0c6dd8fab1662ea5bb6d6b80e11bf91ab56851f9b2719ace8fd664ce15

Observation bbbd30f6-11cc-451c-9354-6784853f4804 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL Kimi K2: Open Agentic Intelligence

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T17:01:18.382175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:01:18.382175Z digest=sha256:f45276ab574d78685ab4ccebd081f4e5bf4782c48658a2f57c775993c8662383

Pith citing papers

Observation 3800836e-76ec-4e92-a022-d6015b1b7204 · inbound

RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs cites this paper.

RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:30:55.188214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T05:29:46.136115Z digest=sha256:4ceb58c0df27721b161ae7acd3a7db20e8755d9a3b9bca3c0e5c0ac27a3b6a4f

Observation 8c2c3fe3-f660-4953-838a-d78764c03ecb · inbound

Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning cites this paper.

Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:40:14.503471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T20:38:30.169363Z digest=sha256:2d59e9439e77e5604dcc09f9f1e30218870062206648c7fa06d1bf07e6392f9f

Observation a46dc5c9-5719-4b7b-b5d7-0787bff3f378 · inbound

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training cites this paper.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:47.766365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:47.766365Z digest=sha256:9a9d066bc88fb9e21f31e3dcdebd1b65d50953f9c223aa2b9f8f7b5af05d2fde

Observation 96d1c965-bc1d-46f9-bcf6-a69040c52b5e · inbound

ECHO-2: A Large-Scale Distributed Rollout Framework for Cost-Efficient Reinforcement Learning cites this paper.

ECHO-2: A Large-Scale Distributed Rollout Framework for Cost-Efficient Reinforcement Learning History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T05:32:40.894045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:32:40.894045Z digest=sha256:b7bc3e1a7e165130e7f13984b74fd18b27e4ab2203f92907e6ffbed90f2a4943

Observation 11ab2c75-bcae-42df-a493-f22c02e03898 · inbound

TensorHub: Scalable and Elastic Weight Transfer for LLM RL Training cites this paper.

TensorHub: Scalable and Elastic Weight Transfer for LLM RL Training History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:11:00.608168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T17:16:33.020610Z digest=sha256:6393cddac32cbd7919f65e04e86b96a1b68a077639a4034fcb5596546a5ded05

Observation ff08c0c9-878f-4727-9471-72957cf6f468 · inbound

DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training cites this paper.

DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:46:26.419119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-07T13:49:26.560459Z digest=sha256:542e6fd40e6e82b4ec86b33fa1a9c48b7416c050de280a257ee9eec829479b3c

Observation 104cac29-2996-4b38-bdab-8afba1fa7e39 · inbound

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL cites this paper.

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:31:15.958221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T05:14:14.168753Z digest=sha256:13727a5c781c8ef63a2e447296091d4c0799b2708aa0d944912887adb13e6e01

Observation 2cbc7f6b-7da4-453e-aaa2-1a48b797e015 · inbound

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL cites this paper.

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:39:53.249603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T08:39:31.911497Z digest=sha256:70135122555b9c49834fbe12785ab0ee41069b2e7e290ff294f7ed98900f4ca7

Observation 83ad0f94-1742-4f74-8d6c-1ba0c29dea2f · inbound

Gradient Extrapolation-Based Policy Optimization cites this paper.

Gradient Extrapolation-Based Policy Optimization History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:56.750695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-11T02:07:08.792030Z digest=sha256:3f099a651fdc0f158705d1e8691f0b5dc41ce2527a8fb21a3bbaf572726356a1

Observation d1022d61-df46-4f6c-ae73-1317e9940e9f · inbound

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models cites this paper.

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:00:55.193120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-11T02:02:41.411795Z digest=sha256:b2dd779c3ab340071abc2c165a4cd4eeeba535a08a141520839d8285c0f463bf

Observation ea386b91-1233-496f-a36f-0bf3911ab6ef · inbound

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning cites this paper.

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:01:30.843384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T01:23:13.071107Z digest=sha256:28799e73b69a071ab2a7d9413b7844043a704c89e81a30ce87a4726233be2ab8

Observation 77dc928c-00a6-4bbd-84a0-7c3ee9de34c7 · inbound

AIS: Adaptive Importance Sampling for Quantized RL cites this paper.

AIS: Adaptive Importance Sampling for Quantized RL History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:14:52.805415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T03:13:14.384567Z digest=sha256:a283e6988a81af8586c4cfe13725a29ce3b91bcd0b26c26ee41addfc9e99f4f5

Observation 69499fa6-ac48-494e-8c0b-ac469511e626 · inbound

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs cites this paper.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:13:43.860135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:21689368a0f765e8b6afbdd9af5f69e4b1507f0b9dc6fdc0808db00aa8d9462e

Observation 073c2ca5-8078-4af1-94ba-be90e8adc967 · inbound

How Off-Policy Can GRPO Be? Mu-GRPO for Efficient LLM Reinforcement Learning cites this paper.

How Off-Policy Can GRPO Be? Mu-GRPO for Efficient LLM Reinforcement Learning History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:43:19.700305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T13:38:45.835754Z digest=sha256:e700d08488cb139f9b0da5a5ab16f91f12e22e39e31b0601d0440e68cfdbe4b6

Observation e0c15d7d-b9b1-4f87-aee3-ca0310a88838 · inbound

Libra: Efficient Resource Management for Agentic RL Post-Training cites this paper.

Libra: Efficient Resource Management for Agentic RL Post-Training History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:27.228306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T11:02:00.385932Z digest=sha256:fadcfe682f1fb12d14eb2cd511e67dc0d5c0b90f067cef3a8400d5e330247b82

Observation e157f5db-02dd-4ea5-8329-034ddbea5d8c · inbound

Harnessing Routing Foresight for Micro-step-level MoE load balancing in RL Post-training cites this paper.

Harnessing Routing Foresight for Micro-step-level MoE load balancing in RL Post-training History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T12:58:08.759839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T08:35:16.435272Z digest=sha256:5a4df9cc4a1207a59bc4d77e338935f4672867129dbcbba5b9397a95cfc8459c

Observation 6ef0ae08-8da4-46f6-abff-172ecf02e0b6 · inbound

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts cites this paper.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:59:07.266551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:576bcc8def46ca8b08d6a650e28095abe7963898baf5f89312d90fd6c30bba28

Observation 4dd1f559-4df5-4ea8-9935-c44e2eaa6a7e · inbound

Bidirectional Resource Scheduling for Disaggregated and Asynchronous RL Post-Training cites this paper.

Bidirectional Resource Scheduling for Disaggregated and Asynchronous RL Post-Training History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T04:42:35.589143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T04:42:35.589143Z digest=sha256:6d5b6116399e173193298b39af9ca26bbd7c2cf2788d3c336e0b6cc9084200c7

Observation 3f176268-ebbb-498f-a953-8ba9718194a7 · inbound

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization cites this paper.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:18.067189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:18.067189Z digest=sha256:8b44bdbf6435f5aa3707236b642a98019769f4e5916a489ade9567c712ee5c88