Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:53:29.001062Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 6 inbound Pith citation observations for arXiv:2504.14363.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:53:29.001062Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T18:30:38.055667Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T07:59:40.189087Z
56 of 56 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 888524d1-e700-4d13-b272-6e7915a7e389 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay IEEE Signal Processing Magazine 34(6), 26–38 (2017)
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dae19bd0-8887-42d9-9b17-6e1071ebc1e6 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c90a7971-3e79-4e58-bc1b-850676693ab8 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c39bf08-9c78-4d61-9ebe-bcffa6d48bb5 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d20b8933-9a52-46b1-bfc6-e781c9329f0a · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay DocFusion: A Unified Framework for Document Parsing Tasks
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42bc350b-17e5-423b-92a7-b384c0860eac · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay PanGu-Coder: Program Synthesis with Function-Level Language Modeling
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad4659b5-0974-46d4-a975-452382807278 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay Training Verifiers to Solve Math Word Problems
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65229ffc-8d53-46e3-b205-549ceff1b21c · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay In: 2008 Interna- tional Conference on Computational Intelligence for Modelling Control Automa- tion
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8b9976a4-4f52-4bd9-aa9d-cf85a28bd283 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay Mitigating Tail Narrowing in LLM Self-Improvement via Socratic-Guided Sampling
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f00bda77-aa71-440a-8c60-b776ea5265cf · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay arXiv preprint arXiv:2407.06153 (2024)
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae96fd29-0008-44bd-beab-b6cc578ff7bf · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ce400777-20f3-462d-ac3f-b7f3fcd0fb7a · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay MetaRM: Shifted Distributions Alignment via Meta-Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03c2a41e-5bc4-4e4c-a7df-49adf5076bd3 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay Towards Understanding the Capability of Large Language Models on Code Clone Detection: A Survey
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c00eeada-e0ed-49f8-86fe-a6ec1ce69bcc · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay arXiv preprint arXiv:2506.02672 (2025)
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ded2841-43b0-4cf3-9a7e-c520a900f59e · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay In: Proceedings of the 62nd Annual Meeting of the AssociationforComputationalLinguistics(Volume1:LongPapers).pp.1932–1945 (2024)
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a4f13fcd-9947-40fd-92d0-73e41fcf66b6 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay Nature590, 580 – 586 (2020),https://api.semanticscholar.org/ CorpusID:216552951
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 265801ab-340f-4118-a497-9cccdf90e407 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay In: Proceedings of the 41st International Conference on Machine Learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b5a32ec2-5821-47b7-bbf0-a158e84ef373 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eefee861-05aa-4a28-b84a-77bf10127813 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 48a7c01a-d34b-4642-a506-bddcf65291cd · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay In: Vanschoren, J., Yeung, S
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7b9dcbfb-6158-4163-9e86-7b9800b51638 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay Measuring Mathematical Problem Solving With the MATH Dataset
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9055ac67-6aae-419c-95a0-8be5bed2262c · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay Towards Reasoning in Large Language Models: A Survey
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a2d2248-5edb-4f01-a135-7b470cc0818d · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay RATIONALYST: Mining Implicit Rationales for Process Supervision of Reasoning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2e2a4f09-b6db-4598-92c3-e9807baa9d1a · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay Understanding the Effects of RLHF on LLM Generalisation and Diversity
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25a5fd15-d033-42ca-bac4-4b1d4b32e08d · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay LLM Post-Training: A Deep Dive into Reasoning Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9c4d114-ae70-4143-b99d-e46c18016fa9 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay Information Fusion85, 1–22 (2022)
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1896f1dd-c6fb-48f3-b105-f6967430cbcd · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay StarCoder: may the source be with you!
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f60984d5-12cf-40a0-889c-d2d0327c1bd5 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay Deep Reinforcement Learning: An Overview
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d67fe6a5-165f-4abe-981a-cdc3bc643bc2 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay RLTF: Reinforcement Learning from Unit Test Feedback
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b834e9ad-e8c2-4d10-8381-da9a0ba90d6f · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay WizardCoder: Empowering Code Large Language Models with Evol-Instruct
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4196ed8a-44ca-45b0-8c1b-b0e3c47716de · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay Brief analysis of DeepSeek R1 and its implications for Generative AI
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec987063-8bc8-4f1c-8f08-1107e78c04e4 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay In: 2018 IEEE Inter- national Conference on Robotics and Automation (ICRA)
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b74eb59b-54e3-4ada-ae5b-0b7b012fce14 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay arXiv preprint arXiv:2407.11511 (2024)
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de719de5-add6-4a7a-aa2d-16575cc7cd27 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay Code Llama: Open Foundation Models for Code
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98c32d55-ab36-466a-a87f-bc3f7c533e1b · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay Prioritized Experience Replay
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 667eed7b-7bc1-4c79-9e4e-437d47d7f5e1 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52cc8e12-af0e-4e6a-b088-e953c55ea7b2 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay Proximal Policy Optimization Algorithms
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9db5262c-478d-4508-abde-356ab74bef36 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay 5-thinking: Advancing superb reasoning models with reinforcement learning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4bc0226-3566-42ed-aba4-52742dfa2fda · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f232fc04-9978-43c8-8f7e-0226eb9bcb16 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay In: The 2023 Conference on Empirical Methods in Natural Language Processing
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation dfe5a085-9bf6-49bb-bcf9-4e0876a6106a · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay Execution-based Code Generation using Deep Reinforcement Learning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85e0fc10-7300-46e5-98b4-581cc48cf3a4 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay Advances in neural informa- tion processing systems12 (1999)
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 562cef5c-f270-45ac-8046-207313ef9b73 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay Reverse Forward Curriculum Learning for Extreme Sample and Demonstration Efficiency in Reinforcement Learning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32e0b9d7-d87f-4c01-975a-4981ca8b7e02 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 32f7067e-076a-4510-b437-f9e204aa0d67 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a8e414d-0b82-4ddc-a9a5-3eaa050741c2 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay Solving math word problems with process- and outcome-based feedback
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a10ab00-4e0a-4284-89eb-e8ac65cf2a98 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61e3575d-57e1-43c4-ac9a-d246f320caf0 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28577e6f-3437-4add-90e9-300765c13701 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay Machine learning8, 229–256 (1992)
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4fcf402-90c7-4a45-a758-3bc594580228 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay Progressive Mastery: Customized Curriculum Learning with Guided Prompting for Mathematical Reasoning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a571a8ea-e785-4453-ad81-2414c8c6a3d3 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay In: International Conference on Machine Learning
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a016837c-4a05-42e1-b30e-e793c6fd24c5 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay Exploration in Deep Reinforcement Learning: From Single-Agent to Multiagent Domain
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ede8af13-74e9-4136-b940-c3d979966938 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay arXiv preprint arXiv:2505.17793 (2025)
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 455a8c2b-08a9-48e3-9529-729c736bc6d6 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay A Deeper Look at Experience Replay
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bff13c8d-45e2-4636-a8c6-262d0a021569 · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6421c666-6154-4714-ba74-d565060ac54b · outbound
Improving RL Exploration for LLM Reasoning through Retrospective Replay In: NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following (2023)
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a57d2fe3-68d0-4906-9adc-01e65938bc09 · inbound
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning Improving RL Exploration for LLM Reasoning through Retrospective Replay
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 331e8f51-bedd-44b1-8469-cd6695202c97 · inbound
A Survey of Reinforcement Learning for Large Reasoning Models Improving RL Exploration for LLM Reasoning through Retrospective Replay
Reference 118
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8d570112-a7fc-4c0c-84ae-491828bceca0 · inbound
OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning Improving RL Exploration for LLM Reasoning through Retrospective Replay
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0d92ba2d-2def-4264-a3d6-9daafaa5fb0a · inbound
Trust Region On-Policy Distillation Improving RL Exploration for LLM Reasoning through Retrospective Replay
Reference 266
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f71fc0eb-23d5-43eb-abf0-9f60c6da2b44 · inbound
Learning to Solve, Forgetting to Retain: Correct-Set Turnover in RLVR Improving RL Exploration for LLM Reasoning through Retrospective Replay
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f3695488-aca2-48e0-b879-1b07bbbd95a5 · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Improving RL Exploration for LLM Reasoning through Retrospective Replay
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.