Pith. sign in

Paper Citation Record · LEDGER

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

As of 10 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 15 inbound Pith citation observations for arXiv:2507.07451.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07451 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:47:53.026897Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T08:15:59.069553Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:59:40.165966Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b088bdc2-b2a5-4392-a83f-8feaf9ade318 · outbound

This paper cites an unresolved cited work.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:47:53.580139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:47:52.741698Z digest=sha256:461313166d0a074e3ee0485805093616013d52c1e70435f9e9ded88bd413e1e6

Observation 73a3f962-c41a-4bd2-91c6-7e2c8e929205 · outbound

This paper cites an unresolved cited work.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:47:52.863138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:47:52.863138Z digest=sha256:775f06a000564c654bb6405ae6b0e9b7cd0176f80da805e1dded7975cbf168d6

Observation e44144e7-bbef-4dc3-abaf-78aa41bf1c3f · outbound

This paper cites an unresolved cited work.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:47:53.565275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:47:52.963492Z digest=sha256:7c067101dbd5febda32534cab7a9f1d3f20e86382978451b6b3cbc2b18fbb98f

Observation b559eaa4-2bff-45fc-9c63-92603b92741b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:47:52.967896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:47:52.967896Z digest=sha256:f643421695db8b6085a503d0a8f730721f4a93a062c79f464e177f4fee2f373d

Observation 52e57824-a06d-481f-b04a-5e8790af495b · outbound

This paper cites an unresolved cited work.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:47:52.972497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:47:52.972497Z digest=sha256:222cf34432b198d481d5175af05190ff670593370ed20dc70eede1b713d5c378

Observation 474fbb0d-2761-4a12-a0ba-c3bdcad7ada2 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:47:52.976774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:47:52.976774Z digest=sha256:03663ee53835ff1486782681378e854b20a051a38d2790f3dc8b7309700784ef

Observation 9507fb39-46cf-407e-afec-818af53a1e7e · outbound

This paper cites an unresolved cited work.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:47:53.550737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:47:52.981900Z digest=sha256:6cb9d37816f8601dfaa50979bf0c6fcb5a8fb105f293c3c95e143c2759620822

Observation 9477c727-1f68-498f-820c-9995f14a3279 · outbound

This paper cites Openai o1 system card.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Openai o1 system card

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:47:53.534839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:47:52.986066Z digest=sha256:f1b0587c6c2a38dc14ae269c26c686a745a0cdd0714a200e3d6ac08231c702c8

Observation 4ae26289-3377-4e46-8f47-369694799228 · outbound

This paper cites Prioritized Experience Replay.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Prioritized Experience Replay

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:47:52.990150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:47:52.990150Z digest=sha256:e200f5b057133afaf9b6ad51cbae42b0f610b3a1cbb9f77533e74e3329f83a80

Observation 73bca3c6-127e-4ed5-8307-793c812424b6 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:47:52.994487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:47:52.994487Z digest=sha256:fec3eaf09488a932f5c457d5d6bc767b4844805fa693fa5e42865254d03af910

Observation 953f5c33-f686-48ba-b15d-14589505cdf6 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:47:52.998807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:47:52.998807Z digest=sha256:28eba317e2bf495e7635e815fdccf051c9849b01651eeb03f029d35c1a7c6816

Observation 89149ea1-74f9-4a6e-b5a6-fc3827d7b7b3 · outbound

This paper cites an unresolved cited work.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:47:53.003183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:47:53.003183Z digest=sha256:0d51d7f8843437431ece06ddf849a048a7a9d35a0159bf4d5e68b64a6d3cf992

Observation 70b477f0-1efe-40e1-acb5-a0f6cbcc92b2 · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:47:53.007802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:47:53.007802Z digest=sha256:0e7ddf21de9f1c56793cf5b9cd8a5075671669711f0c5d57e307094a5866d0e6

Observation d01f2037-0923-4a59-94f8-19a5ee13f10d · outbound

This paper cites Learning to Reason under Off-Policy Guidance.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Learning to Reason under Off-Policy Guidance

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:47:53.012363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:47:53.012363Z digest=sha256:dd9e60d92c45a3355b2cfe33858a766ef07e79560770fff25adc6d695968e549

Observation a1c3d17e-85a6-4121-b93c-b096f2e75002 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:47:53.017126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:47:53.017126Z digest=sha256:7283dadd5f1977307b4bd38673453dffd2e213e1f1d088efd45d129fb1acf4f9

Observation 99272396-9262-43cf-b742-4022c013c4f1 · outbound

This paper cites Qwen3 Technical Report.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning Qwen3 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:47:53.022014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:47:53.022014Z digest=sha256:8e5fd092c4f65036c06a72cadeee3889cfc6795d81731f92b15122eae6fb182e

Observation 753b1d4b-3e9c-4d5f-b4c2-96ad20b0e56f · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:47:53.026897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:47:53.026897Z digest=sha256:80d7d2912845f986150a275c934be579b13bfff7928c30c26b169f5d34b5279a

Pith citing papers

Observation dac9cfc3-8c42-440f-ae91-ff827c37b32c · inbound

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models cites this paper.

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Reference 112

Resolution
unresolved
no resolver link, observed 2026-08-04T08:15:59.069553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:15:59.069553Z digest=sha256:7b57c1185442ef4e229c5a35f70aafd193ba606b99a1fc828518fd699dd77de0

Observation d7f21a67-f54a-4b20-ac92-4d3c06536a5d · inbound

EasyVideoR1: Easier RL for Video Understanding cites this paper.

EasyVideoR1: Easier RL for Video Understanding RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:42:06.748293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T07:41:27.231098Z digest=sha256:69635bfdc04cf6ac9e696a312d25056e4cd25064d8934b9f3d422212fc338b7e

Observation 2092eeef-c21b-4e07-9145-e23e6f9a9eaa · inbound

Near-Future Policy Optimization cites this paper.

Near-Future Policy Optimization RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:54:48.603133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T00:51:36.580600Z digest=sha256:c9909f71c3a86b7b44e9acf11d354bc3c4837e9339f31514d3a443a83f513411

Observation a8b9d20c-620a-4815-9db1-9c01ea163064 · inbound

Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective cites this paper.

Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:30:55.227343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T01:19:58.448344Z digest=sha256:114d2b0fa2ef589020bc19ec3da09f50046b3e31441a44dada8df8c575f7da20

Observation ac6c1eb1-842d-4db9-8521-9f10ca992b96 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Reference 275

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.655426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:7a4410dbf68fd9842d27ef10f1ee1027f01af2c8842170e91c8f49af2eb949a2

Observation c4d71b5c-3edd-49e8-8ea1-e94e8cbd8310 · inbound

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning cites this paper.

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:16:13.343127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T17:25:50.758630Z digest=sha256:1af6258ba51466ab7fc66d20b108ede62e7dfd6e4f0dd6379d9825a238b6bb3b

Observation 7e06d7af-d662-47fa-95f4-a09215f3a59b · inbound

Learning to Solve, Forgetting to Retain: Correct-Set Turnover in RLVR cites this paper.

Learning to Solve, Forgetting to Retain: Correct-Set Turnover in RLVR RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Reference 143

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:36:25.002327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T11:50:00.954670Z digest=sha256:99970319606e1ba29c0272a4cc75c6a6db797bc4996fb6dfbb92b16943e2d0ac

Observation ddd15fc1-9f0d-44a3-8ea3-1575dfdde3e1 · inbound

Rollout-Level Advantage-Prioritized Experience Replay for GRPO cites this paper.

Rollout-Level Advantage-Prioritized Experience Replay for GRPO RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-07-02T06:06:41.007463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T07:43:21.284574Z digest=sha256:6310f985143f1f121e1bf8b8e263c932a7f8a2261eca4f8f958cfbe887ea77a9

Observation ceee3d62-14e9-492e-b822-b47524bb934e · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:56.024146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:45b97a5f9e8d4524264c9fa59f286fca62d2515a2c18e5ef50303c0759ecf28e

Observation 08963af0-1e52-4b44-8a39-67a91b27948b · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Reference 261

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.167556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:77180dc438ed554da66e9130608d16c30ae0262c97e4b2fd106f7227f75e99fb

Observation cd42a4e4-5c79-42c5-b36b-22bfd527b87d · inbound

DRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training cites this paper.

DRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:24:21.451888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T07:19:00.012926Z digest=sha256:87eb7db189923067ff326235715558da904bd8f0c19e58c3d939584cc0f8e68f

Observation 6e9c4b3e-42fc-4ebd-b38b-3d63fb707d2e · inbound

Experience Augmented Policy Optimization for LLM Reasoning cites this paper.

Experience Augmented Policy Optimization for LLM Reasoning RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:54:20.299729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T06:52:55.368187Z digest=sha256:1a06d35b481b6ef664405ac708d19a19b938d38843ac6d9c4ab3de296ec47d8b

Observation e9e37322-2c65-4c7c-9231-098e61e842ed · inbound

Experience Augmented Policy Optimization for LLM Reasoning cites this paper.

Experience Augmented Policy Optimization for LLM Reasoning RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T09:32:16.230452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:32:16.230452Z digest=sha256:9600828ff88f567356efc965cf28d377233197a79640682478daee6e87f26e44

Observation 501b1519-0b51-4cc8-9d05-ace24da922c5 · inbound

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples cites this paper.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-14T11:23:30.063231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T11:23:30.063231Z digest=sha256:873d12478772ead23499859c1189d4be4f5e948a1788c8a0d863ea771bc03f09

Observation e769e426-7214-4d05-bd84-4105242bdd2e · inbound

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples cites this paper.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:29.013207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:29.013207Z digest=sha256:f0e0665d6ac8a0f61fbc7599f90a4bd10a148c14838ff03f8ebe32ff4b3acf8f