Pith. sign in

Paper Citation Record · LEDGER

S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2502.12853.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.12853 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:13:51.984499Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T06:54:01.056295Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3c58cff3-de1e-4c63-b657-f717887b1988 · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 200

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:36:24.296155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:c3a66fb9e81f0f62bf9c6230642a6031db04fac38253b3b9b9081038089b14db

Observation 93e86eb5-d40a-4b50-8c8c-b65249a2b758 · inbound

Training-Free Reasoning and Reflection in MLLMs cites this paper.

Training-Free Reasoning and Reflection in MLLMs S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:51.984499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:51.984499Z digest=sha256:57e35365bef583beeb07c464e5eeca26af15e748b0ab516fb6b7a0aedf192b97

Observation 3c54ff20-26cc-4bae-93be-5128e3dd5c00 · inbound

LeTS: Learning to Think-and-Search via Process-and-Outcome Reward Hybridization cites this paper.

LeTS: Learning to Think-and-Search via Process-and-Outcome Reward Hybridization S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:51:05.696234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:51:05.696234Z digest=sha256:b48cda81ad83982707de81848bc0e5d16a3fab4a2e3ae1fb9b79915b913fef87

Observation 2d5e7b6f-4e1c-4faf-9f00-b351615363dd · inbound

Boosting LLM Reasoning via Spontaneous Self-Correction cites this paper.

Boosting LLM Reasoning via Spontaneous Self-Correction S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:30.691200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:30.691200Z digest=sha256:85b8c2155d23297236130002070316a1b1d8893cd44207755e83862fa4deb663

Observation e812ee5e-535c-40df-9a12-ddc265680545 · inbound

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier cites this paper.

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:39:34.398292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:39:34.398292Z digest=sha256:5bc1ecc4fec974179b2993ea33609de1170a0645e16e52377f6d8d3a4fad060d

Observation 31143d57-8403-4307-9567-3d65736b6bca · inbound

CogDual: Enhancing Dual Cognition of LLMs via Reinforcement Learning with Implicit Rule-Based Rewards cites this paper.

CogDual: Enhancing Dual Cognition of LLMs via Reinforcement Learning with Implicit Rule-Based Rewards S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:26.932292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:02:26.932292Z digest=sha256:706535e7d032f7a02b868abf4f975486b0e89936333afb2e27553c620c216e83

Observation 686afce3-f261-4af1-b3ef-9204c8fdb512 · inbound

Self-Reflective Generation at Test Time cites this paper.

Self-Reflective Generation at Test Time S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T12:41:42.687257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:41:42.687257Z digest=sha256:eee6d03bf163c9bdc7f1f63165a3d48e52819da6b2fadd50127232fdd634ce8e

Observation 391e202b-2abc-490a-af25-b77342bd120f · inbound

Towards Sparse Video Understanding and Reasoning cites this paper.

Towards Sparse Video Understanding and Reasoning S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T23:33:12.494748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:33:12.494748Z digest=sha256:e02633561f9b0d49f23ca34d079547ffb5dd5370e835c0b2c23a3cd638cee75b

Observation aa851a64-26fc-49e8-a29b-ea8e5102c6e2 · inbound

SVSR: A Self-Verification and Self-Rectification Paradigm for Multimodal Reasoning cites this paper.

SVSR: A Self-Verification and Self-Rectification Paradigm for Multimodal Reasoning S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:31:03.849486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:59:04.802241Z digest=sha256:7fbcdf7ca88c6d46a021f2e592703bb1b9e93de60e85a742813175ee7b12972d

Observation f2f9c1d9-9e73-4d04-9ba7-3a2762563404 · inbound

SVSR: A Self-Verification and Self-Rectification Paradigm for Multimodal Reasoning cites this paper.

SVSR: A Self-Verification and Self-Rectification Paradigm for Multimodal Reasoning S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T22:47:33.917420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:47:33.917420Z digest=sha256:43f86cc2bada812be8b5cd2bdf80e8413acbde997cf5b67aad208ef2e987d077

Observation e91f14a7-9f89-43fd-94ad-3a06bde8597e · inbound

Multi-modal Reasoning with LLMs for Visual Semantic Arithmetic cites this paper.

Multi-modal Reasoning with LLMs for Visual Semantic Arithmetic S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:56:10.337177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T02:35:17.215879Z digest=sha256:8d31fc141854c6e1c08203eafd4652b87cf125dafb02c59a2e65253369350ca5

Observation 90a1b760-73b2-47f8-8e9b-649ec3c0717e · inbound

Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning cites this paper.

Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:31:31.255048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T06:13:09.898530Z digest=sha256:626b12236c4bc9ce63b787846f0327ff37acaef915092591b33fcf14de8e8c54

Observation 68b6e2e8-4366-47fe-95da-e734b1f24516 · inbound

The Hidden Signal of Verifier Strictness: Controlling and Improving Step-Wise Verification via Selective Latent Steering cites this paper.

The Hidden Signal of Verifier Strictness: Controlling and Improving Step-Wise Verification via Selective Latent Steering S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:54:01.058137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T06:51:56.556213Z digest=sha256:8c0e62b55803686a28f32074031093756c9e75ba6aca16318ef998e450dcf23c

Observation b5942e9d-b718-4427-8708-efc52af0e119 · inbound

Spectral Origins of the Self-Correction Blind Spot in Autoregressive Generation cites this paper.

Spectral Origins of the Self-Correction Blind Spot in Autoregressive Generation S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T15:30:23.485228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T15:30:23.485228Z digest=sha256:4f7a95b7d609a5a903b6187507cb0d77d909a30098682184df91e8ca0a0005f1