Pith. sign in

Paper Citation Record · LEDGER

Teaching Large Language Models to Reason with Reinforcement Learning

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 43 inbound Pith citation observations for arXiv:2403.04642.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.04642 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 43 of 43 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:11:08.137792Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T02:42:26.083324Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e620d219-e4d2-4d99-b7cc-e5b38a53b4cf · inbound

Training Language Models to Self-Correct via Reinforcement Learning cites this paper.

Training Language Models to Self-Correct via Reinforcement Learning Teaching Large Language Models to Reason with Reinforcement Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:04:10.444859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T12:04:10.210508Z digest=sha256:f230a8b8e8d6090d6907bd3e85e0605419b0bbba23bd10ea0f83a45f7242b43a

Observation 5e742e69-9d5d-4e1b-ab04-905424c5eb95 · inbound

Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning cites this paper.

Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning Teaching Large Language Models to Reason with Reinforcement Learning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:42:19.179140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T01:42:19.004468Z digest=sha256:5f4190d964caa11628a670e6248f7aa7ff0e0a25f64de0885253b63600c3c85f

Observation 81073105-9561-4640-9bdf-431d133e25bb · inbound

Training Large Language Models to Reason in a Continuous Latent Space cites this paper.

Training Large Language Models to Reason in a Continuous Latent Space Teaching Large Language Models to Reason with Reinforcement Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:29:05.706544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T10:29:05.384381Z digest=sha256:87864f862b0b26b7d3dac4a725ad94d649b18778510533539ac7c626f2579a01

Observation 0d0b062e-1036-4c1f-b7fe-403edeb2d6d2 · inbound

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning cites this paper.

Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning Teaching Large Language Models to Reason with Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T15:11:08.137792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:11:08.137792Z digest=sha256:82ff676c8f1e813cf1784bbe3e4687bdd74c428939c2c4240338c79e76ce0d73

Observation 943e4545-6612-4952-a7c6-53c23682a60a · inbound

On the Emergence of Thinking in LLMs I: Searching for the Right Intuition cites this paper.

On the Emergence of Thinking in LLMs I: Searching for the Right Intuition Teaching Large Language Models to Reason with Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:53.357179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:25:53.357179Z digest=sha256:1c28efb2457799941118efe1bc52bf5822b537a2972d6ea18401719116bc2c60

Observation c98a33f2-13cf-4e37-974c-51909915dc0d · inbound

Process Reward Models for LLM Agents: Practical Framework and Directions cites this paper.

Process Reward Models for LLM Agents: Practical Framework and Directions Teaching Large Language Models to Reason with Reinforcement Learning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T18:39:38.006409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:39:38.006409Z digest=sha256:a681970423c8b60104497f7eb73bcf8ffe9b7643cccea3aa7ade8a399044ce55

Observation 23390e2f-76d4-4c61-b3fb-2f047a4552c5 · inbound

Learning to Reason at the Frontier of Learnability cites this paper.

Learning to Reason at the Frontier of Learnability Teaching Large Language Models to Reason with Reinforcement Learning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.085841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:05fb05a7faa8696529f5c1192db94b7d006a21bfe7fafd4cd83148e52b82eebc

Observation 1a9aa297-dd75-445b-9d9f-0739788963a9 · inbound

Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space cites this paper.

Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space Teaching Large Language Models to Reason with Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:43.520186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:43.520186Z digest=sha256:0020f0ce946679ade0f5eb35ec74827a33a5f1cf4275e6802888594ca9e026f1

Observation 49fc1608-b320-4489-bbfa-ae4d409b03ac · inbound

HCRMP: A LLM-Hinted Contextual Reinforcement Learning Framework for Autonomous Driving cites this paper.

HCRMP: A LLM-Hinted Contextual Reinforcement Learning Framework for Autonomous Driving Teaching Large Language Models to Reason with Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:43.610133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:43.610133Z digest=sha256:a15f9da49bc0f6a4142b00b86f3b3ec23ad33c979a28268377575dbd236f0cec

Observation d19f54ef-db2a-4b5b-867e-f353c9648dd4 · inbound

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models cites this paper.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Teaching Large Language Models to Reason with Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:53.173407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:53.173407Z digest=sha256:38637c02cc28e53bdb71e148235d34e9e58d0fd70c2ae18958e7ee0951296a89

Observation 2120d7eb-4477-465d-8dea-891dc5dd9fb2 · inbound

Learning to Select In-Context Demonstration Preferred by Large Language Model cites this paper.

Learning to Select In-Context Demonstration Preferred by Large Language Model Teaching Large Language Models to Reason with Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:22.127164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:09:22.127164Z digest=sha256:597afd8db371c852ef2d160c8480b03f63a6e89c64ebd103f79f13ccba506722

Observation b280998b-7aeb-4866-9259-74fd9d347658 · inbound

LoVeC: Reinforcement Learning for Better Verbalized Confidence in Long-Form Generations cites this paper.

LoVeC: Reinforcement Learning for Better Verbalized Confidence in Long-Form Generations Teaching Large Language Models to Reason with Reinforcement Learning

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:47:17.959953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T12:43:51.019983Z digest=sha256:e7b1cb026b11c6d424bf2e420b3b8ff2448c3816e4fe03b9f11e787bdf70ea3e

Observation 7e35f750-3716-42a7-88b1-4dbb4526b9d0 · inbound

SCOUT: Teaching Pre-trained Language Models to Enhance Reasoning via Flow Chain-of-Thought cites this paper.

SCOUT: Teaching Pre-trained Language Models to Enhance Reasoning via Flow Chain-of-Thought Teaching Large Language Models to Reason with Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:59.544322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:59.544322Z digest=sha256:2accc889a709ba5224c319d47aa03111e4e083f4e18f5fee2ca692c2349fff33

Observation e135594d-1c46-4d77-8a61-9b2fe77ab4d0 · inbound

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning cites this paper.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Teaching Large Language Models to Reason with Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:28.455709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:28.455709Z digest=sha256:a6e73e70e9360149cf4c216d20457ebdb908ae139109d38407df05ce9b64bc95

Observation a53a291c-66f3-4219-8f0f-dc8fd58cf89f · inbound

Truly Self-Improving Agents Require Intrinsic Metacognitive Learning cites this paper.

Truly Self-Improving Agents Require Intrinsic Metacognitive Learning Teaching Large Language Models to Reason with Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:19.442762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:28:19.442762Z digest=sha256:ee46386ab7896d011566ec34a75b61aa79c4fddf75764c81834601e1db1fa3d9

Observation 28fa47da-3602-46cb-bef4-1b7022a8e596 · inbound

RePO: Replay-Enhanced Policy Optimization cites this paper.

RePO: Replay-Enhanced Policy Optimization Teaching Large Language Models to Reason with Reinforcement Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T04:53:56.348495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:53:56.348495Z digest=sha256:16b48bb3e629f33315a8ec93186b4facf1661886525644ebb72268ef55c528fb

Observation 1c3bf2f5-bbaa-4cbe-aa0d-b4f53a75aaf8 · inbound

Intent Factored Generation: Unleashing the Diversity in Your Language Model cites this paper.

Intent Factored Generation: Unleashing the Diversity in Your Language Model Teaching Large Language Models to Reason with Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:47:42.354458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:47:42.354458Z digest=sha256:1e249ce58d909ad33586123d1eedae469097238e1aedd93ea72b6a44e21d932a

Observation 31b4fc56-00b6-434c-9a07-5a53afa3bab2 · inbound

RAST: Reasoning Activation in LLMs via Small-model Transfer cites this paper.

RAST: Reasoning Activation in LLMs via Small-model Transfer Teaching Large Language Models to Reason with Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:53.893423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:53.893423Z digest=sha256:f36f690f643ba675c9c7bade686c3f3b74a93769e53872784bc100e25ae3a5a8

Observation 30cebf1b-9edf-4c93-a4c6-0d06a81f86d1 · inbound

Learning Efficient Robotic Garment Manipulation with Standardization cites this paper.

Learning Efficient Robotic Garment Manipulation with Standardization Teaching Large Language Models to Reason with Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:00.026680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:00.026680Z digest=sha256:34ec11b787415a97b09d94d0933cd9e8078cedfd3c84713c5a2c0f08a16391d4

Observation 1e441468-f270-49bd-8275-064d1d46a9b2 · inbound

Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models cites this paper.

Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models Teaching Large Language Models to Reason with Reinforcement Learning

Reference 139

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:04.082222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:04.082222Z digest=sha256:ae3948297e3629767b1bdaadd636e9d928128017ba9ce53d2ad89442db44c534

Observation 32d4999e-f949-4bcf-ba50-b84d03b796ca · inbound

When LLMs Copy to Think: Uncovering Copy-Guided Attacks in Reasoning LLMs cites this paper.

When LLMs Copy to Think: Uncovering Copy-Guided Attacks in Reasoning LLMs Teaching Large Language Models to Reason with Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:07:27.843871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:07:27.843871Z digest=sha256:28f62385fb648167e786946ee9d7f0afc92dc5da3c790c2ac18e501bb746c117

Observation 2ee5efad-7c3b-49cb-b87e-062e32774c0e · inbound

Med-R$^3$: Enhancing Medical Retrieval-Augmented Reasoning of LLMs via Progressive Reinforcement Learning cites this paper.

Med-R$^3$: Enhancing Medical Retrieval-Augmented Reasoning of LLMs via Progressive Reinforcement Learning Teaching Large Language Models to Reason with Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T10:44:26.838066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:44:26.838066Z digest=sha256:7b887e20799adbc559877d69b0cfe39ed88bd0b82f46436d5a6d6ac777b394fe

Observation 4cccdf65-b86f-4eaa-8584-d24c5534aadb · inbound

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization cites this paper.

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization Teaching Large Language Models to Reason with Reinforcement Learning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-19T01:16:57.097300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T01:15:49.658412Z digest=sha256:5fba1ecd8bb35f8ab2a3628a8827d51cfe8892d6c9c879fe8aec12688f09ffd2

Observation d1ae46cd-7879-4233-b5ff-171a85413323 · inbound

Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling cites this paper.

Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling Teaching Large Language Models to Reason with Reinforcement Learning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:51:50.809262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T20:49:26.966293Z digest=sha256:14af581ee86578aa06b717905be1e222425a1913d92ac09d17dae6342cfa7560

Observation 7c481f30-7042-4931-9867-16411b6c3d47 · inbound

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance cites this paper.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Teaching Large Language Models to Reason with Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:08.022212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:08.022212Z digest=sha256:85ac5931826da516f34e4b82880992a95af53f58c6b93f9ce325d4a05ae8e173

Observation 86d251eb-e19e-42fc-aca2-e7b41d675a63 · inbound

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment cites this paper.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Teaching Large Language Models to Reason with Reinforcement Learning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:20:54.606098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:6865a41f542c0e1fd7da81542ebd22c2d90ce31aea75875716f07c445b6cf7ae

Observation f7fc2164-4bbc-444f-8172-de9b9d233944 · inbound

Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning cites this paper.

Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning Teaching Large Language Models to Reason with Reinforcement Learning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:13:47.929498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:11:26.411893Z digest=sha256:f60bba84a799e1eed44945bb3d9dc67d257ba6c0c33eedd25340c0c36b527155

Observation f81f4d0a-b7c7-4490-9bbc-57456cdccdf8 · inbound

SeLaR: Selective Latent Reasoning in Large Language Models cites this paper.

SeLaR: Selective Latent Reasoning in Large Language Models Teaching Large Language Models to Reason with Reinforcement Learning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:35:49.556259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T18:27:36.132030Z digest=sha256:cb33487c76fdb6d96f6d6a6cadb01efe23a8165d96c6668616975733c16e3242

Observation 3fa2893a-9d5a-40d5-a0d3-b56879453ce5 · inbound

LiFT: Does Instruction Fine-Tuning Improve In-Context Learning for Longitudinal Modelling by Large Language Models? cites this paper.

LiFT: Does Instruction Fine-Tuning Improve In-Context Learning for Longitudinal Modelling by Large Language Models? Teaching Large Language Models to Reason with Reinforcement Learning

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:18:22.899179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T00:13:44.570359Z digest=sha256:23ee53bf937f39166380818ad72b083c5f8b7ec27cb99baa2df4107d334e12e5

Observation 6a38efa3-6279-45c4-b446-85466f863fef · inbound

Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding cites this paper.

Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding Teaching Large Language Models to Reason with Reinforcement Learning

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:16:06.034886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T23:05:05.251150Z digest=sha256:be758908b77f684cb34d6652d35717e3a4d1770d882c0af1e820a5756a5d38a4

Observation 94c2ee82-3972-41d2-9102-f160f9404b65 · inbound

Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning cites this paper.

Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning Teaching Large Language Models to Reason with Reinforcement Learning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:26:27.797962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T06:13:09.898530Z digest=sha256:4e69a12b62a1ba0a8e1850a47e5c905ef71aec272be20451a495429552490b49

Observation d02a8372-ccfc-427e-8cbc-d60d3256639e · inbound

Logic-Regularized Verifier Elicits Reasoning from LLMs cites this paper.

Logic-Regularized Verifier Elicits Reasoning from LLMs Teaching Large Language Models to Reason with Reinforcement Learning

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:51:10.941397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T10:54:01.229934Z digest=sha256:ccb5281e654b8f8c223f1a6ef520629e22a82abbbee1e28611c4161556165eb8

Observation dda6e555-6545-49a4-8474-436f8a5da3d9 · inbound

NoisyCoconut: Counterfactual Consensus via Latent Space Reasoning cites this paper.

NoisyCoconut: Counterfactual Consensus via Latent Space Reasoning Teaching Large Language Models to Reason with Reinforcement Learning

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:41:24.258184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T00:51:40.815981Z digest=sha256:716de1cad326430274d6f25987f7334d2f002511e95b05642fd173cd559c4985

Observation d6d976ea-a5e7-40f6-8fe6-1d56bae82ff5 · inbound

Epistemic Uncertainty for Test-Time Discovery cites this paper.

Epistemic Uncertainty for Test-Time Discovery Teaching Large Language Models to Reason with Reinforcement Learning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:57:06.061818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T01:52:41.192353Z digest=sha256:186d14d19d02efa45f73ffd0af18ab98dcfa09623ceaabf81916495a0efd6844

Observation c81a34d1-d755-4ac4-914f-1d137bb1437d · inbound

When Reasoning Traces Become Performative: Step-Level Evidence that Chain-of-Thought Is an Imperfect Oversight Channel cites this paper.

When Reasoning Traces Become Performative: Step-Level Evidence that Chain-of-Thought Is an Imperfect Oversight Channel Teaching Large Language Models to Reason with Reinforcement Learning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:32:24.335451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T06:30:12.558660Z digest=sha256:0d590a46380a5881ac39c480ef435ee60bc06a3f465a20987ae3660d9d3b0193

Observation 9a1e33e8-a197-49ed-ae07-e9a23d3c2147 · inbound

From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI cites this paper.

From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI Teaching Large Language Models to Reason with Reinforcement Learning

Reference 139

Resolution
unresolved
no resolver link, observed 2026-08-02T11:29:31.754741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:29:31.754741Z digest=sha256:5980419723ca72f126253d3af45021f84cd6b2597e3484b4e594e2ae77df8cb8

Observation 40d461bd-51fb-4411-a91d-074baac150f9 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Teaching Large Language Models to Reason with Reinforcement Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:1c5adb8ff9aff10503e816e7afa2ec87735fc0d929f2ceef9dc300dd1b89706a

Observation 15e17f8c-3d1d-4117-a752-b61ea7913795 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Teaching Large Language Models to Reason with Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:36.004785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:36.004785Z digest=sha256:b77584e872b7d1e9cd14f70c91d72d0acda73c96d013efabdea31efca0c19777

Observation 58f57867-c9d6-4227-84b9-37c8ca272be2 · inbound

Structured Synthetic Reasoning Data for Arithmetic Fine-Tuning of Small Language Models cites this paper.

Structured Synthetic Reasoning Data for Arithmetic Fine-Tuning of Small Language Models Teaching Large Language Models to Reason with Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T13:26:28.373498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:26:28.373498Z digest=sha256:9d14bde48f1ca199cf1acb5575125e337636b2a0e166666df572c308204c16c0

Observation d3178038-8e0d-482d-b8e8-4a088c93762d · inbound

LeAct: Learning to Reason from Expert Actions cites this paper.

LeAct: Learning to Reason from Expert Actions Teaching Large Language Models to Reason with Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T06:32:19.440448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:32:19.440448Z digest=sha256:e74f55aa59c72d0343797f0dad061cf8463eedb150c30c0adfb7ca1b35eb12f5

Observation 46f0c937-c1c1-42b7-b42e-7b9169cd431c · inbound

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning cites this paper.

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning Teaching Large Language Models to Reason with Reinforcement Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T07:34:56.790463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:34:56.790463Z digest=sha256:d7446500be9d3570e3cbd4033ed8fa2f2ea376ed64fcfb60e741f0bb81624c3c

Observation 639afc79-1014-4fc7-896c-13a2914bd6db · inbound

RepoReasoner: Evaluating Repository-Level Code Reasoning Ability of Long-Context Language Models cites this paper.

RepoReasoner: Evaluating Repository-Level Code Reasoning Ability of Long-Context Language Models Teaching Large Language Models to Reason with Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T00:57:34.320933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:57:34.320933Z digest=sha256:f975209bd8f4e891990497793131ef9d1abcce26d158ac9c42f54ca4c4278ccd

Observation 3b9c5573-8ad2-4bc5-88c1-ea212ae184dc · inbound

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details cites this paper.

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details Teaching Large Language Models to Reason with Reinforcement Learning

Reference 193

Resolution
unresolved
no resolver link, observed 2026-08-05T15:25:40.180493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:25:40.180493Z digest=sha256:cb1263d3fb56b89f9a1984146752a32519d4a67f865b71a6fe68e2f97f56be34