Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Pre-Training

As of 11 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 17 inbound Pith citation observations for arXiv:2506.08007.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08007 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:26:22.163198Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:14:24.194857Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T23:49:02.979112Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3d8f1335-deaf-46fa-80f0-f65e345af051 · outbound

This paper cites GPT-4 Technical Report.

Reinforcement Pre-Training GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.074900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.074900Z digest=sha256:e28f0f776ebe663897be34b269e1b87752eb2b4145a68f025d92f30edb050c74

Observation 6d7b5087-d78a-413f-8956-492c0623a394 · outbound

This paper cites Notably, the performance of R1-Distill-Qwen-14B is evaluated in two different manner: standard next-token prediction and reasoning-based answer prediction (indicated as ‘+ think’).

Reinforcement Pre-Training Notably, the performance of R1-Distill-Qwen-14B is evaluated in two different manner: standard next-token prediction and reasoning-based answer prediction (indicated as ‘+ think’)

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:26:22.469919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T05:26:22.163198Z digest=sha256:e8452dd0926301e332f5562615d430a0f8832cd28366cfca6a016b58829cb043

Observation 914df787-450c-4f09-adb8-3a345f3b7ae3 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Reinforcement Pre-Training Measuring Massive Multitask Language Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.097018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.097018Z digest=sha256:1f3a6213b4a2270e0fe0085aa7204ef48b6a2efb85382605f79a2e8deb820171

Observation d9211a4a-fdfe-4571-8125-a93116d8f10f · outbound

This paper cites Skywork Open Reasoner 1 Technical Report.

Reinforcement Pre-Training Skywork Open Reasoner 1 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.114660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.114660Z digest=sha256:cc506fb5878f691b73b5520301ff3f0238f6a4aae6c65bb1657a3db3df36f930

Observation 7e82d631-dda2-4351-ba9c-19ed2971b7ff · outbound

This paper cites OpenAI o1 System Card.

Reinforcement Pre-Training OpenAI o1 System Card

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.120498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.120498Z digest=sha256:317e4c861bc2de098da1498bba92d9da6b86b72737b2dff21b7cba1575f88edf

Observation c83d3403-b506-4074-8cd4-8a8e049e2775 · outbound

This paper cites Scaling Laws for Neural Language Models.

Reinforcement Pre-Training Scaling Laws for Neural Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.126155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.126155Z digest=sha256:e32467626b809bc322c7b29093882af9ae38709d56ab1748bfe0ac3f216a7994

Observation b84d26e8-1849-4665-96d1-d46135f5a881 · outbound

This paper cites General-Reasoner: Advancing LLM Reasoning Across All Domains.

Reinforcement Pre-Training General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.131730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.131730Z digest=sha256:ae551124fcb38cce0103f055672808c7243928ed656da39bad535a41f04fe3dc

Observation 331c4ee9-c514-4eaf-9c09-c942f71f0115 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Reinforcement Pre-Training HybridFlow: A Flexible and Efficient RLHF Framework

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.136750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.136750Z digest=sha256:886c3fb15c85a2ed871f758fecf548b6feebc0d8c9d1c25780c551ece88f641c

Observation 99f5089e-12c0-4f1e-b071-e2dbd9917abd · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Reinforcement Pre-Training DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.142165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.142165Z digest=sha256:31cf71e4599c07a3dacb36f5e047660e5810c3a859d65af027ff622b8d654aa1

Observation e81c777e-7081-4098-8990-1ba27f7473e8 · outbound

This paper cites Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking.

Reinforcement Pre-Training Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.147943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.147943Z digest=sha256:2370fd5953ba5de667ae59edd7042f6fb39758114fac0a4c5d36545e76fd486d

Observation 5fbd83f4-6378-40a7-ab4d-98acd0087fee · outbound

This paper cites A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?.

Reinforcement Pre-Training A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.153069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.153069Z digest=sha256:d1fe62cc9acb53fc7117063f3dc54d2baefbd71fdeab6864db12f0a1b8f0b50c

Observation 881d37f7-a003-4bc1-a8de-289095639b6c · outbound

This paper cites Reinforcing General Reasoning without Verifiers.

Reinforcement Pre-Training Reinforcing General Reasoning without Verifiers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.158083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.158083Z digest=sha256:f77a307a957df8240e2817a9a7fa760df772f9755b2672765fd1b4f4d2150473

Observation fa147932-a54b-4940-b749-83a0c62ea092 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Reinforcement Pre-Training Training Compute-Optimal Large Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.108551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.108551Z digest=sha256:017d811f5ad8bab64009dcb748b0f8751acf27439212f66d568c598b773f05f4

Observation 5389ff90-d184-48a4-92d4-a32f79566ead · outbound

This paper cites SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines.

Reinforcement Pre-Training SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.080543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.080543Z digest=sha256:53652c00ac9ac6bb2c00a8c6862125bec80e4042cc61ac063d8da53b77a2d637

Observation 1448d7ff-8ad2-4796-a615-b5970a04fb3d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Reinforcement Pre-Training DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.091745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.091745Z digest=sha256:27fc411d16497db9338c00f938e2f102228dd6d87c76c23afc3bacfe11f9c081

Observation e711829c-0ac3-47de-a042-751ba8441949 · outbound

This paper cites Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models.

Reinforcement Pre-Training Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.086147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.086147Z digest=sha256:0b5b81b99624aea1ebe5131698d53c3752100fa774362bc97ba29cbdf36505de

Pith citing papers

Observation 6367d345-5a3b-46e6-960c-b812ef785c02 · inbound

A Survey on Latent Reasoning cites this paper.

A Survey on Latent Reasoning Reinforcement Pre-Training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:14:24.194857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:14:24.194857Z digest=sha256:03d84b9a25a2d72093601bc96c0ee49699043c1a4de960bf69ea7db61af41e4f

Observation 2d97bc87-0bb9-4a59-8dba-42c4713a9f58 · inbound

Advancing Event Forecasting through Massive Training of Large Language Models: Challenges, Solutions, and Broader Impacts cites this paper.

Advancing Event Forecasting through Massive Training of Large Language Models: Challenges, Solutions, and Broader Impacts Reinforcement Pre-Training

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T14:22:31.430188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:22:31.430188Z digest=sha256:fa870495723d351006e0c099301e263250bf05d5b5671c03a511082f61594ce3

Observation df887682-effa-49cc-926f-47e4e88480b0 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Reinforcement Pre-Training

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T19:21:48.478904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:a8bb10246507b18b176174772932687d7581019acd5e8a6b06a22c9b0c8607a6

Observation c37de614-28a2-4049-bf61-f6feba96f1aa · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models Reinforcement Pre-Training

Reference 116

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:02:25.291186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:46e6571d822546c302c87bcd9362d6f2dde3baf3432f46a786614dbea1a7a912

Observation 481bd3f4-b9cc-419d-be04-f5ae99ef2ace · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Reinforcement Pre-Training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.792504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.792504Z digest=sha256:6195c9cb5e07c70fc5c8f51450bc54f8bff01becc9260debb0208cb641d5b349

Observation b9b28ba4-21ed-401a-9f34-9bec50a52788 · inbound

Refine Thought: A Test-Time Inference Method for Embedding Model Reasoning cites this paper.

Refine Thought: A Test-Time Inference Method for Embedding Model Reasoning Reinforcement Pre-Training

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T09:59:47.063555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:59:47.063555Z digest=sha256:c006cf81542786b51c9bd624b7f3081c68b518579f6abcbc21167c7b451eabba

Observation 95dfa54e-503a-4b53-bef1-2844aa2839a1 · inbound

Agentic Reasoning for Large Language Models cites this paper.

Agentic Reasoning for Large Language Models Reinforcement Pre-Training

Reference 207

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:14:26.226702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T15:14:25.558878Z digest=sha256:b08bed730311484c4903bd2566ba984855c3033f5ea1c7194203e0dbf4c81fe2

Observation 09b938cb-cfa9-409d-9fde-7fd96d8601ba · inbound

CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning cites this paper.

CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning Reinforcement Pre-Training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T05:14:18.061552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:14:18.061552Z digest=sha256:fb919aaa2eefef8a344251b39ceab0d1540a09a72ece0a692f5389026cb95f54

Observation 19deb135-22eb-4e42-a387-9f0ed445e441 · inbound

Beyond Test-Time Memory: State-Space Optimal Control for LLM Reasoning cites this paper.

Beyond Test-Time Memory: State-Space Optimal Control for LLM Reasoning Reinforcement Pre-Training

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-15T12:09:16.468055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T12:09:16.468055Z digest=sha256:10d5e295c6c27afc060a4cdae031dc4f377fb39b6315ddeda5e6e574a951f864

Observation 65083332-4263-4475-9d35-f7512bdf980f · inbound

PubSwap: Public-Data Off-Policy Coordination for Federated RLVR cites this paper.

PubSwap: Public-Data Off-Policy Coordination for Federated RLVR Reinforcement Pre-Training

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:21:00.321774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T16:05:27.319466Z digest=sha256:497a01cab6fb6d71f456c281748a029aeb7f8973bb7907d93dfff02649d1c225

Observation dddbcb56-be36-4ec5-90db-16b4ac71e80c · inbound

From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space cites this paper.

From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space Reinforcement Pre-Training

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:41:04.013130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T12:50:57.603403Z digest=sha256:1f19cf07f682bd983f22a663d9b79d2bfc2798ecb0eae73bf0b1af5ebf7bd9a1

Observation 461999dc-ad6a-4e77-a4ae-6c712e72e1bb · inbound

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities cites this paper.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Reinforcement Pre-Training

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:22:41.994710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:140affe7dadd0dd66ccbe4b3bb3660e708a02a96463ca9198a259d328eb1ebe6

Observation 014e622b-6525-4eb9-9844-da26c2ebbd97 · inbound

Detecting and Mitigating the Correct-Answer Extinction Window in Test-Time Reinforcement Learning with Majority Voting cites this paper.

Detecting and Mitigating the Correct-Answer Extinction Window in Test-Time Reinforcement Learning with Majority Voting Reinforcement Pre-Training

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:23:06.972704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-20T07:20:50.835826Z digest=sha256:b2978ef5cae025f144a890e6e64467672bad553929af8f7228f8e14887f57e7e

Observation 3534117c-d9fa-4072-af1e-56b4d2aa5b6d · inbound

Value-Gradient Hypothesis of RL for LLMs cites this paper.

Value-Gradient Hypothesis of RL for LLMs Reinforcement Pre-Training

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:54:45.940210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-22T08:52:21.361294Z digest=sha256:9a3977534f391986e1941edf2be9073378e2b8bbf7c4fff9c3f7cb6b964520b4

Observation a758bfb8-2b50-4fad-a4d5-d89675e1fabe · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Reinforcement Pre-Training

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.441534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:8e5d0f09227b96c39fdd2cc6b2c070ff297d9e6132aeae024e91d80ab0166291

Observation ba6d57a6-4722-4e50-a417-4552047a6fb5 · inbound

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training cites this paper.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Reinforcement Pre-Training

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:46:29.032828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:fb0dcc27297d0db57ac885cec5291d4c2e13f8463b187792ecf681418f032ad3

Observation 6b8fd053-aa22-480d-82f6-17d2e23ecb2a · inbound

Be Your Own Teacher: Steering Protein Language Models via Unsupervised Reward Optimization cites this paper.

Be Your Own Teacher: Steering Protein Language Models via Unsupervised Reward Optimization Reinforcement Pre-Training

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:49:02.980949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T21:40:55.946788Z digest=sha256:581349713e65b7be3d07dc7a0f4106542b9123f42affbc5f202e1374c7a9d314