Pith. sign in

Paper Citation Record · LEDGER

SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2504.14286.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.14286 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:40:19.810322Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:27:36.773999Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8fa0b0cb-40ff-44f7-9096-c22dd012a6ec · inbound

rStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified Dataset cites this paper.

rStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified Dataset SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 35

Resolution
malformed identifier
no resolver link, observed 2026-08-07T13:40:19.810322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:40:19.810322Z digest=sha256:165a6447c5cd96e3cbe268923115ff66a45c569a554175af9fe1076cf04461e8

Observation 0550c118-947f-459e-bfbe-4a6c265c01e0 · inbound

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts cites this paper.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:57.154091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:57.154091Z digest=sha256:28b203b69a81865c4bfdbd4048f0a456557beb35afcda8a942896c0e301c32bd

Observation 824023c0-625e-448b-a72b-e5d5cb7d259b · inbound

FinLMM-R1: Enhancing Financial Reasoning in LMM through Scalable Data and Reward Design cites this paper.

FinLMM-R1: Enhancing Financial Reasoning in LMM through Scalable Data and Reward Design SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:40.069218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:40.069218Z digest=sha256:2cf9bff85aa82554dfe34d94bd912aa97dd9e0cd46d6c7d4e96fb2e8c632a57d

Observation cfda8c5e-4bc2-4b25-bbef-9a520947345d · inbound

Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs cites this paper.

Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:23.492554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:23.492554Z digest=sha256:72d942e41eec4268d7a280bd2092f16b088d3d2ac03a0d0bbbadff849ba21990

Observation ef670f0a-6fbe-4d06-9322-eeac51fcc7e5 · inbound

AdapThink: Adaptive Thinking Preferences for Reasoning Language Model cites this paper.

AdapThink: Adaptive Thinking Preferences for Reasoning Language Model SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:41.334745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:28:41.334745Z digest=sha256:5597a0886e0bc152d296e104cde63f461c794cc6e242b60fb61d80cdad0dc5cd

Observation 3ee80aa1-e4b6-4d90-ba9b-0c678caca20c · inbound

From Data-Centric to Sample-Centric: Enhancing LLM Reasoning via Progressive Optimization cites this paper.

From Data-Centric to Sample-Centric: Enhancing LLM Reasoning via Progressive Optimization SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:07:36.624737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:07:36.624737Z digest=sha256:e7c9610a04281fd8df4ad30763be2a4ae56d113fba5ea972a5646e29a885e2e8

Observation 588b069e-cf15-4971-b6ae-7fdd8a2ffc51 · inbound

First Return, Entropy-Eliciting Explore cites this paper.

First Return, Entropy-Eliciting Explore SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:15.010459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:15.010459Z digest=sha256:043099b1e1ed97c40f796178ba61e12e351105be92dd0ba5fcaec476a76b0600

Observation e16bb927-3e76-40a4-91cc-875e83af6f8a · inbound

KAT-V1: Kwai-AutoThink Technical Report cites this paper.

KAT-V1: Kwai-AutoThink Technical Report SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:28:30.436815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:28:30.436815Z digest=sha256:560265413d1eb45aa59c5ea834224faeba027dca12e7bb11a8aac58bd8c77096

Observation b3ded519-52c8-419b-8937-bbaaa9dccd8a · inbound

The Challenge of Teaching Reasoning to LLMs Without RL or Distillation cites this paper.

The Challenge of Teaching Reasoning to LLMs Without RL or Distillation SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:18.054670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:51:18.054670Z digest=sha256:4ba7113171174c68d1200648b1c772317e0f2ba468e575a8545a0218a86c03bb

Observation 01cf387a-dc15-4d58-9191-cbe4f1ed3b14 · inbound

CodeReasoner: Enhancing the Code Reasoning Ability with Reinforcement Learning cites this paper.

CodeReasoner: Enhancing the Code Reasoning Ability with Reinforcement Learning SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T14:50:28.067043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:50:28.067043Z digest=sha256:a0e7ed164af363a03309ae29afe37d8d52322101c45ec4ea86457097bfc9e645

Observation 5ff6a99d-6525-439b-bee4-924540413cb4 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.803557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:7cbb00b8f2d477c1d2c146ab3039150741182ad7ec3c6e3e8e21752df59432ca

Observation 34d0a9c6-8d97-4206-9615-6b63681e28f7 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 240

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:47.154757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:47.154757Z digest=sha256:3f25391583c89d1387240c73f1995e1d65fb30e2ee741e31d5e43ecd536b630f

Observation 7fcae94b-67bf-4505-ae59-0626b50c976e · inbound

Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning cites this paper.

Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T12:52:28.407488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:52:28.407488Z digest=sha256:32745aaffb36e08774323da8440ca6fc8bb025aa28f0639928cf377bbf867116

Observation 86c664fe-99db-4d12-bcd6-9eed609058f3 · inbound

Mitigating Lost in Multi-turn Conversation via Curriculum RL with Verifiable Accuracy and Abstention Rewards cites this paper.

Mitigating Lost in Multi-turn Conversation via Curriculum RL with Verifiable Accuracy and Abstention Rewards SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T05:05:54.784655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T05:04:10.226166Z digest=sha256:1ac72cd5fdcab5c754959e169e83b40d6f8d516195c6e445c62d888c378d4e02

Observation bbee64ab-ad32-429f-a653-8048f268b749 · inbound

KG-Hopper: Empowering Compact Open LLMs with Knowledge Graph Reasoning via Reinforcement Learning cites this paper.

KG-Hopper: Empowering Compact Open LLMs with Knowledge Graph Reasoning via Reinforcement Learning SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:25:07.567438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T06:24:23.965754Z digest=sha256:00b208f20e24bfc8356534aaf2f788df3571a36b6de5e8a58ba5787e58d05c86

Observation 040d6b68-72c5-4668-9004-f56b942cb283 · inbound

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models cites this paper.

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:00:55.027452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T02:02:41.411795Z digest=sha256:51aa35c6a97307b1df538259602e3ad0be5f0ab271f59cc01fb6627b61420241

Observation f0f34b89-99ca-49ff-8421-e1ce57ea0b24 · inbound

StepCodeReasoner: Aligning Code Reasoning with Stepwise Execution Traces via Reinforcement Learning cites this paper.

StepCodeReasoner: Aligning Code Reasoning with Stepwise Execution Traces via Reinforcement Learning SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:32:20.256262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T05:27:37.521421Z digest=sha256:091a7851ed7aba1bdc737e34a66fccd929861e0f6bf0b5685a9c80f624e2f9a9

Observation 8e9b1bbe-6e87-460d-b37f-80140c6a958c · inbound

DISA: Offline Importance Sampling for Distribution-Matching LLM-RL cites this paper.

DISA: Offline Importance Sampling for Distribution-Matching LLM-RL SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T15:13:24.931997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T15:11:43.235574Z digest=sha256:223283600da280eb328f3e8a6cb92d9ecd51c3ecfeb25991c2fefef9545ee796

Observation 6929a437-184e-42fc-852b-7815e614f977 · inbound

Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards cites this paper.

Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:59:41.075776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T05:55:45.654673Z digest=sha256:b16a983a59ef8e575199c601e3950dc7c704950443d47099bd06133619896389

Observation d2ff5de9-1909-4aff-841b-382f44db42a0 · inbound

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning cites this paper.

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:16:13.400279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T17:25:50.758630Z digest=sha256:02a4142bbb8493812f95314c7329536e6128a7a491dfefbc1d5cd3f9318bb2fe

Observation 33ec4857-2c52-4929-896a-e3658182a151 · inbound

AdaGRPO: A Capability-Aware Adaptive Enhancement for Flow-based GRPO cites this paper.

AdaGRPO: A Capability-Aware Adaptive Enhancement for Flow-based GRPO SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:17:08.889756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T22:54:21.740892Z digest=sha256:b6896d24ea438336edb1fa37e384cb4bbe0b058db8b2fc268baddf508d0863f2

Observation f659790b-0b17-460e-a7b4-f359adc99aca · inbound

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning cites this paper.

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:27:36.775325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T13:55:35.363377Z digest=sha256:b98978d653e8425490a800c008f6016e7fdb61edd770bf759e4c7186560998ac

Observation 0721e090-647b-4ff0-8756-616091f58f48 · inbound

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR cites this paper.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T00:29:00.841014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:29:00.841014Z digest=sha256:b833e2a4d84841031882fff269c3bf212b8e839d13a93f5270c345570d9b58b0

Observation ed222e3f-3531-41b9-ac31-46ac4fb7ef93 · inbound

ReCo: Reweighting GRPO Against Distributional Concentration cites this paper.

ReCo: Reweighting GRPO Against Distributional Concentration SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-30T18:55:22.314325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T18:55:22.314325Z digest=sha256:d65462187633895838eb583ad0530a682fac5f40f31e3d0918797821aa26987a