Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2502.19613.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:59:21.879268Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation bac2e894-6082-431f-ac62-b315f956605a · inbound
Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL Self-rewarding correction for mathematical reasoning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 526399d3-1612-40f6-95c7-11537eb8752b · inbound
A Survey of Slow Thinking-based Reasoning LLMs using Reinforced Learning and Inference-time Scaling Law Self-rewarding correction for mathematical reasoning
Reference 126
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7840a340-42c9-4ada-a926-8534ee4d2064 · inbound
Scalable Chain of Thoughts via Elastic Reasoning Self-rewarding correction for mathematical reasoning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 968ddbac-06f6-4cee-adb7-a6b99431d84a · inbound
Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards Self-rewarding correction for mathematical reasoning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33f319fc-9290-4edf-ac9a-6f53d1eb7b31 · inbound
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Self-rewarding correction for mathematical reasoning
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a00c80d3-824f-4ab2-922a-582cb76fbd30 · inbound
Boosting LLM Reasoning via Spontaneous Self-Correction Self-rewarding correction for mathematical reasoning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d8ec851-bcbe-4610-825e-f5be039ea645 · inbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Self-rewarding correction for mathematical reasoning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 128b7cd0-6b9c-4786-9296-6e6a7a9fc255 · inbound
PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier Self-rewarding correction for mathematical reasoning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97f62b28-d998-499f-aa58-9eb0923c1a72 · inbound
Scaling Test-time Compute for LLM Agents Self-rewarding correction for mathematical reasoning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc1bd704-6917-4eda-a7cf-32a282e1a413 · inbound
Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training Self-rewarding correction for mathematical reasoning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 06743f5c-b56f-4ab0-a5ae-fcede9ccce86 · inbound
LightReasoner: Can Small Language Models Teach Large Language Models Reasoning? Self-rewarding correction for mathematical reasoning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5830c2c4-5a28-44d0-9fea-33295906121a · inbound
Breaking the Self-Confirming Loop: Diagnosing and Mitigating Systemic Reward Bias in Self-Rewarding RL Self-rewarding correction for mathematical reasoning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b896a909-8106-4d06-ac64-1436ef0624b9 · inbound
CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning Self-rewarding correction for mathematical reasoning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81f621f4-3604-428e-8cc2-e0bf2c72abc3 · inbound
rePIRL: Learn PRM with Inverse RL for LLM Reasoning Self-rewarding correction for mathematical reasoning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation dbfc3b72-9643-43e4-85a6-9b411c13cfd7 · inbound
rePIRL: Learn PRM with Inverse RL for LLM Reasoning Self-rewarding correction for mathematical reasoning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45e07ccf-0df1-4722-b143-7e48f81aada3 · inbound
Can LLMs Learn to Reason Robustly under Noisy Supervision? Self-rewarding correction for mathematical reasoning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6c7cb4b3-0047-45bd-a6ed-f07c5cd35174 · inbound
Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning Self-rewarding correction for mathematical reasoning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cf61d227-58ea-4435-9385-6261c2420a7c · inbound
ACE: Self-Evolving LLM Coding Framework via Adversarial Unit Test Generation and Preference Optimization Self-rewarding correction for mathematical reasoning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f414539e-896d-4e51-bb05-6e65dc61007e · inbound
ACE: Self-Evolving LLM Coding Framework via Adversarial Unit Test Generation and Preference Optimization Self-rewarding correction for mathematical reasoning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0dd7aacd-e04f-4881-a417-eea4ab73c190 · inbound
Guarded Repair for Harm-Aware Post-hoc Replacement of LLM Mathematical Reasoning Self-rewarding correction for mathematical reasoning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 03fe794f-2c38-44d0-961d-2ae2f83409a2 · inbound
Trust Region On-Policy Distillation Self-rewarding correction for mathematical reasoning
Reference 162
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c9f7a4fa-0632-4437-b449-b34408d22a14 · inbound
The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes Self-rewarding correction for mathematical reasoning
Reference 274
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 86fff995-f14c-4448-b088-6cff9d7586e5 · inbound
ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning Self-rewarding correction for mathematical reasoning
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation deeadaf6-7404-490b-a02c-e0291d1527f5 · inbound
ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning Self-rewarding correction for mathematical reasoning
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cf7f342-4619-40d3-8df7-e5cc461b49d7 · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Self-rewarding correction for mathematical reasoning
Reference 238
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation af7e3537-efcf-4951-ba7a-d7583bc4b8f7 · inbound
Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction Self-rewarding correction for mathematical reasoning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8428ec32-f496-4353-9265-0eb5a8b61781 · inbound
Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction Self-rewarding correction for mathematical reasoning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.