Pith. sign in

Paper Citation Record · LEDGER

Self-rewarding correction for mathematical reasoning

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2502.19613.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.19613 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:59:21.879268Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bac2e894-6082-431f-ac62-b315f956605a · inbound

Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL cites this paper.

Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL Self-rewarding correction for mathematical reasoning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T00:59:21.879268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:59:21.879268Z digest=sha256:56cafda66c3ad1724dba185b986bc687e6f6f4e59396510b9092f91c75ae95b9

Observation 526399d3-1612-40f6-95c7-11537eb8752b · inbound

A Survey of Slow Thinking-based Reasoning LLMs using Reinforced Learning and Inference-time Scaling Law cites this paper.

A Survey of Slow Thinking-based Reasoning LLMs using Reinforced Learning and Inference-time Scaling Law Self-rewarding correction for mathematical reasoning

Reference 126

Resolution
unresolved
no resolver link, observed 2026-08-16T00:48:19.274791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:48:19.274791Z digest=sha256:9ff5a5acb7693e92387789a4c53786b8474fddabc5c2901228165f78ddd71014

Observation 7840a340-42c9-4ada-a926-8534ee4d2064 · inbound

Scalable Chain of Thoughts via Elastic Reasoning cites this paper.

Scalable Chain of Thoughts via Elastic Reasoning Self-rewarding correction for mathematical reasoning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T23:13:33.639132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:13:33.639132Z digest=sha256:3e6eb1e640974676aa8622409e5ff271d81d084d522ad2d7b4f623f44e693c64

Observation 968ddbac-06f6-4cee-adb7-a6b99431d84a · inbound

Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards cites this paper.

Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards Self-rewarding correction for mathematical reasoning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:14.112520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:14.112520Z digest=sha256:a8066f8302d54f6986d2a32325c267d04e140c778fc6e68b21d52ecf4cae3537

Observation 33f319fc-9290-4edf-ac9a-6f53d1eb7b31 · inbound

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning cites this paper.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Self-rewarding correction for mathematical reasoning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:33.115481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:33.115481Z digest=sha256:5a21225bc362c52a5e06eb1ccff91841374e1af7b55cbd947de8208e0a968482

Observation a00c80d3-824f-4ab2-922a-582cb76fbd30 · inbound

Boosting LLM Reasoning via Spontaneous Self-Correction cites this paper.

Boosting LLM Reasoning via Spontaneous Self-Correction Self-rewarding correction for mathematical reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:30.740802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:30.740802Z digest=sha256:882287e632b92c29156a8274808dc3fee7399653103bfa30860aba3e15d7a844

Observation 3d8ec851-bcbe-4610-825e-f5be039ea645 · inbound

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning cites this paper.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Self-rewarding correction for mathematical reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:45.657899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:45.657899Z digest=sha256:5e56093759b30eb5ec53a1dc31f4b009a5d6b610c14e758cd3672b9857289343

Observation 128b7cd0-6b9c-4786-9296-6e6a7a9fc255 · inbound

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier cites this paper.

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier Self-rewarding correction for mathematical reasoning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:39:37.601699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:39:37.601699Z digest=sha256:34e94a3289f1e643682a9b7c383ca869808a8251f2885b421b6df187cb91f6f8

Observation 97f62b28-d998-499f-aa58-9eb0923c1a72 · inbound

Scaling Test-time Compute for LLM Agents cites this paper.

Scaling Test-time Compute for LLM Agents Self-rewarding correction for mathematical reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:48.404875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:48.404875Z digest=sha256:ac6137f5815729ae8663d31a85a8b2fa5ae559353d196349332bc0a38ea29aec

Observation cc1bd704-6917-4eda-a7cf-32a282e1a413 · inbound

Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training cites this paper.

Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training Self-rewarding correction for mathematical reasoning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:40:43.171530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-21T22:38:57.833414Z digest=sha256:388293c0caefd393ba9e362546cfd1f652783cfe0118b72813088b8d87bb8e79

Observation 06743f5c-b56f-4ab0-a5ae-fcede9ccce86 · inbound

LightReasoner: Can Small Language Models Teach Large Language Models Reasoning? cites this paper.

LightReasoner: Can Small Language Models Teach Large Language Models Reasoning? Self-rewarding correction for mathematical reasoning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:01:34.019544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T12:59:31.011283Z digest=sha256:9f39194ad0e08700d33f15ca273a959fe5e67f176fe7f10f829fb31a8aa521c5

Observation 5830c2c4-5a28-44d0-9fea-33295906121a · inbound

Breaking the Self-Confirming Loop: Diagnosing and Mitigating Systemic Reward Bias in Self-Rewarding RL cites this paper.

Breaking the Self-Confirming Loop: Diagnosing and Mitigating Systemic Reward Bias in Self-Rewarding RL Self-rewarding correction for mathematical reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T10:44:31.460549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:44:31.460549Z digest=sha256:9403736b988dd42f9e2b726bfcf3585dcaaf2f6e3566e0d874ea913e9238fbc3

Observation b896a909-8106-4d06-ac64-1436ef0624b9 · inbound

CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning cites this paper.

CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning Self-rewarding correction for mathematical reasoning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T05:14:21.533431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:14:21.533431Z digest=sha256:954566f20203be9089dc32527db514f40b8f5cd11a04a07e0bd8b0dde5729b2b

Observation 81f621f4-3604-428e-8cc2-e0bf2c72abc3 · inbound

rePIRL: Learn PRM with Inverse RL for LLM Reasoning cites this paper.

rePIRL: Learn PRM with Inverse RL for LLM Reasoning Self-rewarding correction for mathematical reasoning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:14:10.976923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T13:13:13.293921Z digest=sha256:8ee2174f82e676559ef49001e5eaaf480440a140f76a69965cb5a06dc9aef0db

Observation dbfc3b72-9643-43e4-85a6-9b411c13cfd7 · inbound

rePIRL: Learn PRM with Inverse RL for LLM Reasoning cites this paper.

rePIRL: Learn PRM with Inverse RL for LLM Reasoning Self-rewarding correction for mathematical reasoning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:45.150810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:45.150810Z digest=sha256:750379dc072def825b570815a900828f957fc6c334c19cbc7488dddedca1be85

Observation 45e07ccf-0df1-4722-b143-7e48f81aada3 · inbound

Can LLMs Learn to Reason Robustly under Noisy Supervision? cites this paper.

Can LLMs Learn to Reason Robustly under Noisy Supervision? Self-rewarding correction for mathematical reasoning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:08:01.264453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T16:58:42.129870Z digest=sha256:aa7e9d656d1137c0f83dd6ece497e202aeedbf3b184f07fd7ba20d36ddba0950

Observation 6c7cb4b3-0047-45bd-a6ed-f07c5cd35174 · inbound

Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning cites this paper.

Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning Self-rewarding correction for mathematical reasoning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:26:27.829461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T06:13:09.898530Z digest=sha256:f662494f220159dfbd5ec7b5e377ace4df7ec497c7a6144cbc322ae98e80fbd7

Observation cf61d227-58ea-4435-9385-6261c2420a7c · inbound

ACE: Self-Evolving LLM Coding Framework via Adversarial Unit Test Generation and Preference Optimization cites this paper.

ACE: Self-Evolving LLM Coding Framework via Adversarial Unit Test Generation and Preference Optimization Self-rewarding correction for mathematical reasoning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-21T00:59:19.336113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T00:55:40.784295Z digest=sha256:3a8624f47e883da004885cb3d88c4b518972347c46edfef11b6a27f8c23298ec

Observation f414539e-896d-4e51-bb05-6e65dc61007e · inbound

ACE: Self-Evolving LLM Coding Framework via Adversarial Unit Test Generation and Preference Optimization cites this paper.

ACE: Self-Evolving LLM Coding Framework via Adversarial Unit Test Generation and Preference Optimization Self-rewarding correction for mathematical reasoning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:14:47.165809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T10:14:00.478000Z digest=sha256:3c233e3b9f47ee786648b36637c27286d999b5e0956610f13a2a00627cbd50d2

Observation 0dd7aacd-e04f-4881-a417-eea4ab73c190 · inbound

Guarded Repair for Harm-Aware Post-hoc Replacement of LLM Mathematical Reasoning cites this paper.

Guarded Repair for Harm-Aware Post-hoc Replacement of LLM Mathematical Reasoning Self-rewarding correction for mathematical reasoning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:44:40.809593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-30T13:38:48.617204Z digest=sha256:0b2773b9dabdfaeda7010ac77278f572ebfe56fee0b2d233d53ec7e2430309f0

Observation 03fe794f-2c38-44d0-961d-2ae2f83409a2 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Self-rewarding correction for mathematical reasoning

Reference 162

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.435749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:6f695bfda767ada65102651d8240f7f3a0576b0a50e8ac305d9f4ae38c1f0de2

Observation c9f7a4fa-0632-4437-b449-b34408d22a14 · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes Self-rewarding correction for mathematical reasoning

Reference 274

Resolution
verified exact
arxiv_id, observed 2026-06-27T13:00:55.970436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:a7da285e3fab4d444e4271b0f2c8c9b5c99e1451dfbae134a3aa5e9b5bdbcccf

Observation 86fff995-f14c-4448-b088-6cff9d7586e5 · inbound

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning cites this paper.

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning Self-rewarding correction for mathematical reasoning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:08:33.388470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T06:39:34.199607Z digest=sha256:8768af6404da1450bd73e303ce284c362154771f9ffd4fedbb6db9016b7df523

Observation deeadaf6-7404-490b-a02c-e0291d1527f5 · inbound

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning cites this paper.

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning Self-rewarding correction for mathematical reasoning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-03T02:12:24.010055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:12:24.010055Z digest=sha256:346ea5903cbb40e627ce7e34aa5e08c0815ef4ff21c0029e9bc1753752bd10ab

Observation 2cf7f342-4619-40d3-8df7-e5cc461b49d7 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Self-rewarding correction for mathematical reasoning

Reference 238

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.132829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:e0cb21c456c84a960bf504b7823f0d0968f4dc832cf732bbb2919f2d11cdb079

Observation af7e3537-efcf-4951-ba7a-d7583bc4b8f7 · inbound

Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction cites this paper.

Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction Self-rewarding correction for mathematical reasoning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-01T17:15:51.588940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T03:58:32.372896Z digest=sha256:3d8195f593f125baaa7b91125a97af9c38bd77bd5483bb9485ed2ca38c48aef1

Observation 8428ec32-f496-4353-9265-0eb5a8b61781 · inbound

Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction cites this paper.

Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction Self-rewarding correction for mathematical reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T17:12:24.565155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:12:24.565155Z digest=sha256:60f2b3e1a065e51d19f80538b8da7891c9fb170a745d589a078de7f244306a3d