Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2504.14286.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:40:19.810322Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T04:27:36.773999Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 8fa0b0cb-40ff-44f7-9096-c22dd012a6ec · inbound
rStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified Dataset SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0550c118-947f-459e-bfbe-4a6c265c01e0 · inbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 824023c0-625e-448b-a72b-e5d5cb7d259b · inbound
FinLMM-R1: Enhancing Financial Reasoning in LMM through Scalable Data and Reward Design SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfda8c5e-4bc2-4b25-bbef-9a520947345d · inbound
Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef670f0a-6fbe-4d06-9322-eeac51fcc7e5 · inbound
AdapThink: Adaptive Thinking Preferences for Reasoning Language Model SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ee80aa1-e4b6-4d90-ba9b-0c678caca20c · inbound
From Data-Centric to Sample-Centric: Enhancing LLM Reasoning via Progressive Optimization SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 588b069e-cf15-4971-b6ae-7fdd8a2ffc51 · inbound
First Return, Entropy-Eliciting Explore SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e16bb927-3e76-40a4-91cc-875e83af6f8a · inbound
KAT-V1: Kwai-AutoThink Technical Report SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3ded519-52c8-419b-8937-bbaaa9dccd8a · inbound
The Challenge of Teaching Reasoning to LLMs Without RL or Distillation SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01cf387a-dc15-4d58-9191-cbe4f1ed3b14 · inbound
CodeReasoner: Enhancing the Code Reasoning Ability with Reinforcement Learning SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ff6a99d-6525-439b-bee4-924540413cb4 · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 34d0a9c6-8d97-4206-9615-6b63681e28f7 · inbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM
Reference 240
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fcae94b-67bf-4505-ae59-0626b50c976e · inbound
Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86c664fe-99db-4d12-bcd6-9eed609058f3 · inbound
Mitigating Lost in Multi-turn Conversation via Curriculum RL with Verifiable Accuracy and Abstention Rewards SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bbee64ab-ad32-429f-a653-8048f268b749 · inbound
KG-Hopper: Empowering Compact Open LLMs with Knowledge Graph Reasoning via Reinforcement Learning SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 040d6b68-72c5-4668-9004-f56b942cb283 · inbound
Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM
Reference 113
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f0f34b89-99ca-49ff-8421-e1ce57ea0b24 · inbound
StepCodeReasoner: Aligning Code Reasoning with Stepwise Execution Traces via Reinforcement Learning SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8e9b1bbe-6e87-460d-b37f-80140c6a958c · inbound
DISA: Offline Importance Sampling for Distribution-Matching LLM-RL SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6929a437-184e-42fc-852b-7815e614f977 · inbound
Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d2ff5de9-1909-4aff-841b-382f44db42a0 · inbound
RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 33ec4857-2c52-4929-896a-e3658182a151 · inbound
AdaGRPO: A Capability-Aware Adaptive Enhancement for Flow-based GRPO SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f659790b-0b17-460e-a7b4-f359adc99aca · inbound
TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0721e090-647b-4ff0-8756-616091f58f48 · inbound
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed222e3f-3531-41b9-ac31-46ac4fb7ef93 · inbound
ReCo: Reweighting GRPO Against Distributional Concentration SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.