Pith. sign in

Paper Citation Record · LEDGER

SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2504.14286.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.14286 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:40:13.081592Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:27:36.773999Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e58237ea-71b5-4bb8-b946-dbe383eb6477 · inbound

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization cites this paper.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.081592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.081592Z digest=sha256:2c94afeaf6d547da0722b32b99140aa035fd2f4676af8949a741a3b48c57220d

Observation 8fa0b0cb-40ff-44f7-9096-c22dd012a6ec · inbound

rStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified Dataset cites this paper.

rStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified Dataset SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 35

Resolution
malformed identifier
no resolver link, observed 2026-08-07T13:40:19.810322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:40:19.810322Z digest=sha256:319c173fa7326c743c3dd9c9b55ded5448e2ceec8a38ec787dff679050489f48

Observation 0550c118-947f-459e-bfbe-4a6c265c01e0 · inbound

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts cites this paper.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:57.154091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:57.154091Z digest=sha256:45c3cee7f13b1f67c22fbc438389509acd4c3537059648c229c9ed4cbe845a73

Observation 824023c0-625e-448b-a72b-e5d5cb7d259b · inbound

FinLMM-R1: Enhancing Financial Reasoning in LMM through Scalable Data and Reward Design cites this paper.

FinLMM-R1: Enhancing Financial Reasoning in LMM through Scalable Data and Reward Design SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:40.069218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:40.069218Z digest=sha256:e2bd6191058185fabd1fbbceca6ea6cf56c3b721d25ae4bb44306b3c6a209075

Observation cfda8c5e-4bc2-4b25-bbef-9a520947345d · inbound

Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs cites this paper.

Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:23.492554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:23.492554Z digest=sha256:97772cbbfa9abca0c11e737fff32ec5a11264da5d13a9cc219f8653aed681f9c

Observation ef670f0a-6fbe-4d06-9322-eeac51fcc7e5 · inbound

AdapThink: Adaptive Thinking Preferences for Reasoning Language Model cites this paper.

AdapThink: Adaptive Thinking Preferences for Reasoning Language Model SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:41.334745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:28:41.334745Z digest=sha256:234e6af4f358de3501740ab9b59d9182a3baa0b32cbf64e383f1075435ad714d

Observation 3ee80aa1-e4b6-4d90-ba9b-0c678caca20c · inbound

From Data-Centric to Sample-Centric: Enhancing LLM Reasoning via Progressive Optimization cites this paper.

From Data-Centric to Sample-Centric: Enhancing LLM Reasoning via Progressive Optimization SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:07:36.624737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:07:36.624737Z digest=sha256:57f38d9f6eafb537c975314d0fd9f5793d0638f544cfbb56b78b19ca3dd46cc2

Observation 588b069e-cf15-4971-b6ae-7fdd8a2ffc51 · inbound

First Return, Entropy-Eliciting Explore cites this paper.

First Return, Entropy-Eliciting Explore SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:15.010459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:15.010459Z digest=sha256:aa750a407d28c40aa45332fd8b5044cdead0afd3208658cefe4a83a09eb49ba3

Observation e16bb927-3e76-40a4-91cc-875e83af6f8a · inbound

KAT-V1: Kwai-AutoThink Technical Report cites this paper.

KAT-V1: Kwai-AutoThink Technical Report SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:28:30.436815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:28:30.436815Z digest=sha256:cd2701d33b4aaf8971efcae484469da3e0033cf28a17b57d04e91a7478efaf05

Observation b3ded519-52c8-419b-8937-bbaaa9dccd8a · inbound

The Challenge of Teaching Reasoning to LLMs Without RL or Distillation cites this paper.

The Challenge of Teaching Reasoning to LLMs Without RL or Distillation SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:18.054670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:51:18.054670Z digest=sha256:8dd608adce1a75495c34266dc75f143427e7bca619aa5cbdec025f5598b60a31

Observation 01cf387a-dc15-4d58-9191-cbe4f1ed3b14 · inbound

CodeReasoner: Enhancing the Code Reasoning Ability with Reinforcement Learning cites this paper.

CodeReasoner: Enhancing the Code Reasoning Ability with Reinforcement Learning SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T14:50:28.067043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:50:28.067043Z digest=sha256:42d6c1e80ba39a0a55898a2cd4cff19449a71a8db986b1330f6fe34e7c063c1f

Observation 2b99a34e-ab6b-47cb-bdbd-f3d009cdce61 · inbound

CURE: Critical-Token-Guided Re-Concatenation for Entropy-Collapse Prevention cites this paper.

CURE: Critical-Token-Guided Re-Concatenation for Entropy-Collapse Prevention SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T17:36:21.490968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:36:21.490968Z digest=sha256:500bb0d81126265ff5081830b785c38b1a35b673b703be4e4a74ce961e634437

Observation 5ff6a99d-6525-439b-bee4-924540413cb4 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.803557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:fc0076f73e499cf026bac7967c191fa402bda5b2a03ff33da2fd70a0f5ad3444

Observation 34d0a9c6-8d97-4206-9615-6b63681e28f7 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 240

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:47.154757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:47.154757Z digest=sha256:ae9343dd9897afd7ee9bbaaafd6be8371a73b162d4e04fec3bad79d8ed07743d

Observation 7fcae94b-67bf-4505-ae59-0626b50c976e · inbound

Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning cites this paper.

Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T12:52:28.407488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:52:28.407488Z digest=sha256:19291b82d7fa100d3a24ca79bb295cd3bf8a77f0b198be29b4abeb776dfc0d59

Observation 86c664fe-99db-4d12-bcd6-9eed609058f3 · inbound

Mitigating Lost in Multi-turn Conversation via Curriculum RL with Verifiable Accuracy and Abstention Rewards cites this paper.

Mitigating Lost in Multi-turn Conversation via Curriculum RL with Verifiable Accuracy and Abstention Rewards SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T05:05:54.784655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T05:04:10.226166Z digest=sha256:4c76308acf366d91dff06fc904862aa9aa7480758203f3cd4b2e89a4229faacf

Observation bbee64ab-ad32-429f-a653-8048f268b749 · inbound

KG-Hopper: Empowering Compact Open LLMs with Knowledge Graph Reasoning via Reinforcement Learning cites this paper.

KG-Hopper: Empowering Compact Open LLMs with Knowledge Graph Reasoning via Reinforcement Learning SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:25:07.567438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T06:24:23.965754Z digest=sha256:10b1004a760b43518ab9a29d412be41cb2269ef5282d07e0b47953e26a5ef835

Observation 040d6b68-72c5-4668-9004-f56b942cb283 · inbound

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models cites this paper.

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:00:55.027452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-11T02:02:41.411795Z digest=sha256:2cff5285cea9351f1c22c2dbb2465bdb3a549840c1956199e23a7576b2e3e849

Observation f0f34b89-99ca-49ff-8421-e1ce57ea0b24 · inbound

StepCodeReasoner: Aligning Code Reasoning with Stepwise Execution Traces via Reinforcement Learning cites this paper.

StepCodeReasoner: Aligning Code Reasoning with Stepwise Execution Traces via Reinforcement Learning SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:32:20.256262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-13T05:27:37.521421Z digest=sha256:37aaf35fe0d800a7a301283e22416815b3c798dbf155bcc9ea1790b7c0f21aca

Observation 8e9b1bbe-6e87-460d-b37f-80140c6a958c · inbound

DISA: Offline Importance Sampling for Distribution-Matching LLM-RL cites this paper.

DISA: Offline Importance Sampling for Distribution-Matching LLM-RL SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T15:13:24.931997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-20T15:11:43.235574Z digest=sha256:c351a87facaf4b192c901aea9aae4733c206c4d9975e8ab3abe1cfca29180268

Observation 6929a437-184e-42fc-852b-7815e614f977 · inbound

Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards cites this paper.

Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:59:41.075776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-21T05:55:45.654673Z digest=sha256:9475e9e90bb4f84d96e86383fc164c52ab406afa8fbe546cb910e4d8f7fc1dce

Observation d2ff5de9-1909-4aff-841b-382f44db42a0 · inbound

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning cites this paper.

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:16:13.400279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T17:25:50.758630Z digest=sha256:190ae8ea9fdda311578c75f237c011ce4a56976b53c02436eb01823a93b7ebdb

Observation 33ec4857-2c52-4929-896a-e3658182a151 · inbound

AdaGRPO: A Capability-Aware Adaptive Enhancement for Flow-based GRPO cites this paper.

AdaGRPO: A Capability-Aware Adaptive Enhancement for Flow-based GRPO SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:17:08.889756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:54:21.740892Z digest=sha256:386722f2e039713e0c2c8ae56e5723050295c803aeed4ba49f912d5589bc9abb

Observation f659790b-0b17-460e-a7b4-f359adc99aca · inbound

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning cites this paper.

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:27:36.775325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-27T13:55:35.363377Z digest=sha256:9b6a9896b5296fcbd85505b08de4cfbc053eac24ada2aa31a566dcf729b4ef65

Observation 0721e090-647b-4ff0-8756-616091f58f48 · inbound

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR cites this paper.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T00:29:00.841014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:29:00.841014Z digest=sha256:a9b34c431d9359c92b30284879257bf143d092d0c968e610edfcb17a7c148b83

Observation ed222e3f-3531-41b9-ac31-46ac4fb7ef93 · inbound

ReCo: Reweighting GRPO Against Distributional Concentration cites this paper.

ReCo: Reweighting GRPO Against Distributional Concentration SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-30T18:55:22.314325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T18:55:22.314325Z digest=sha256:119a1b218b8cd3fabd424eb266b5e957a6cb7c6a912f69977c5ae25b00884431

Observation 3e71b636-a63a-4298-b12c-07c50937e5aa · inbound

Evidence-Grounded Forensic Reasoning for Detecting and Grounding Multi-Modal Media Manipulation cites this paper.

Evidence-Grounded Forensic Reasoning for Detecting and Grounding Multi-Modal Media Manipulation SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T00:38:22.040895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:38:22.040895Z digest=sha256:cd4dce164cda6e3db9e0b46f4165cf2c8bce29fd1e482ddc3e98bf92368182f5