Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Pre-Training

As of 8 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 17 inbound Pith citation observations for arXiv:2506.08007.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08007 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:26:22.163198Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:14:24.194857Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T23:49:02.979112Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3d8f1335-deaf-46fa-80f0-f65e345af051 · outbound

This paper cites GPT-4 Technical Report.

Reinforcement Pre-Training GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.074900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.074900Z digest=sha256:778ede36b7ab44808051f3b3f62fb04f52e689d932dee90a09b22cf90c94b069

Observation 6d7b5087-d78a-413f-8956-492c0623a394 · outbound

This paper cites Notably, the performance of R1-Distill-Qwen-14B is evaluated in two different manner: standard next-token prediction and reasoning-based answer prediction (indicated as ‘+ think’).

Reinforcement Pre-Training Notably, the performance of R1-Distill-Qwen-14B is evaluated in two different manner: standard next-token prediction and reasoning-based answer prediction (indicated as ‘+ think’)

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:26:22.469919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:26:22.163198Z digest=sha256:5cc6e39180c056da15dc1818a7beb91279c1a161c7026a4aa73fff602748c5ff

Observation 914df787-450c-4f09-adb8-3a345f3b7ae3 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Reinforcement Pre-Training Measuring Massive Multitask Language Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.097018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.097018Z digest=sha256:bce568752a794dab12aca5a2846d65f6127e150056d83992d9a498178e6561f1

Observation d9211a4a-fdfe-4571-8125-a93116d8f10f · outbound

This paper cites Skywork Open Reasoner 1 Technical Report.

Reinforcement Pre-Training Skywork Open Reasoner 1 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.114660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.114660Z digest=sha256:bacf59cdc59d668bacab3ec6ac691a6056d41233e8bb94396def02c157a2bada

Observation 7e82d631-dda2-4351-ba9c-19ed2971b7ff · outbound

This paper cites OpenAI o1 System Card.

Reinforcement Pre-Training OpenAI o1 System Card

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.120498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.120498Z digest=sha256:48ebf5408d55e4e10c30a5729caef56bf4e6a7ea8dbc278f6bf67a25dfd971bd

Observation c83d3403-b506-4074-8cd4-8a8e049e2775 · outbound

This paper cites Scaling Laws for Neural Language Models.

Reinforcement Pre-Training Scaling Laws for Neural Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.126155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.126155Z digest=sha256:c544ff79133c9dfd9b0619e7988a894a77ccc9eda39bb342e4f8c52b7fabee91

Observation b84d26e8-1849-4665-96d1-d46135f5a881 · outbound

This paper cites General-Reasoner: Advancing LLM Reasoning Across All Domains.

Reinforcement Pre-Training General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.131730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.131730Z digest=sha256:f6b3ccecd11ed6b141cd02291c3fd04d169f68b8607c67b30dfd03e572edbd9a

Observation 331c4ee9-c514-4eaf-9c09-c942f71f0115 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Reinforcement Pre-Training HybridFlow: A Flexible and Efficient RLHF Framework

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.136750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.136750Z digest=sha256:907718e26fbd22f4ad4f8120ff8b93e1c956186aa593ee68074856716000636f

Observation 99f5089e-12c0-4f1e-b071-e2dbd9917abd · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Reinforcement Pre-Training DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.142165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.142165Z digest=sha256:a96cc777eef64e75deab70ba66adb09e4fc052797d426399f39e12a1c39402e0

Observation e81c777e-7081-4098-8990-1ba27f7473e8 · outbound

This paper cites Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking.

Reinforcement Pre-Training Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.147943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.147943Z digest=sha256:ad416239e8a01400632fb5f9ecbe1806452bb869b264cc125f19121136ca4a3f

Observation 5fbd83f4-6378-40a7-ab4d-98acd0087fee · outbound

This paper cites A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?.

Reinforcement Pre-Training A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.153069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.153069Z digest=sha256:932da4a211f1618024c5748cdf5ab3f7b3728d11f805042ff6f6c6841cdeb906

Observation 881d37f7-a003-4bc1-a8de-289095639b6c · outbound

This paper cites Reinforcing General Reasoning without Verifiers.

Reinforcement Pre-Training Reinforcing General Reasoning without Verifiers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.158083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.158083Z digest=sha256:4f5bc93fc9f969ec97eb39b06eb273851c96ef631791d8f16f4c3de2bee7387b

Observation fa147932-a54b-4940-b749-83a0c62ea092 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Reinforcement Pre-Training Training Compute-Optimal Large Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.108551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.108551Z digest=sha256:3f08ddeae8585becdfa3b40da8ba8db8810345d73f34df6237232e45727a1bdc

Observation 5389ff90-d184-48a4-92d4-a32f79566ead · outbound

This paper cites SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines.

Reinforcement Pre-Training SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.080543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.080543Z digest=sha256:a6f03ac65004b09282e70dc6de6a24cb80b83a2d816b4247cdd58ad6e523b479

Observation 1448d7ff-8ad2-4796-a615-b5970a04fb3d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Reinforcement Pre-Training DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.091745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.091745Z digest=sha256:b6c4ab5f9c14014bc8ff3d7dab54f6ce872c3b1538e84358dbba393b66e98c1d

Observation e711829c-0ac3-47de-a042-751ba8441949 · outbound

This paper cites Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models.

Reinforcement Pre-Training Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:22.086147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:22.086147Z digest=sha256:78975b8ae3397b3c5979384feba55aaaa22d0f34ed3ea32381815d2631d4f9ed

Pith citing papers

Observation 6367d345-5a3b-46e6-960c-b812ef785c02 · inbound

A Survey on Latent Reasoning cites this paper.

A Survey on Latent Reasoning Reinforcement Pre-Training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:14:24.194857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:14:24.194857Z digest=sha256:2d590bba6b462c2a4f5ed99d35eb2438188747f4bde65d700ae686929d59123e

Observation 2d97bc87-0bb9-4a59-8dba-42c4713a9f58 · inbound

Advancing Event Forecasting through Massive Training of Large Language Models: Challenges, Solutions, and Broader Impacts cites this paper.

Advancing Event Forecasting through Massive Training of Large Language Models: Challenges, Solutions, and Broader Impacts Reinforcement Pre-Training

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T14:22:31.430188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:22:31.430188Z digest=sha256:5ce32de1a9a737415583d5d8f675a9a1dfa4bb8f0c12e794e2a94b4ac7bf3f01

Observation df887682-effa-49cc-926f-47e4e88480b0 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Reinforcement Pre-Training

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T19:21:48.478904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:048d1d76e9165ee8e823d0d6e512a4afe7a153eccd1db1e5c0d74c50b5f88921

Observation c37de614-28a2-4049-bf61-f6feba96f1aa · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models Reinforcement Pre-Training

Reference 116

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:02:25.291186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:61a32b22f42631392fb965c067a815dc8d4b8b829f4b4db153e066ab1bc7faa5

Observation 481bd3f4-b9cc-419d-be04-f5ae99ef2ace · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Reinforcement Pre-Training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.792504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.792504Z digest=sha256:6f029aa5747c14073125b624e38f6a859b21c4b7f570fb8a499ba16d8900ea66

Observation b9b28ba4-21ed-401a-9f34-9bec50a52788 · inbound

Refine Thought: A Test-Time Inference Method for Embedding Model Reasoning cites this paper.

Refine Thought: A Test-Time Inference Method for Embedding Model Reasoning Reinforcement Pre-Training

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T09:59:47.063555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:59:47.063555Z digest=sha256:51a3ea2f8d16c6b83f1cc0dd80040f1aa0cfda26afef3ef0f7f25643873c522b

Observation 95dfa54e-503a-4b53-bef1-2844aa2839a1 · inbound

Agentic Reasoning for Large Language Models cites this paper.

Agentic Reasoning for Large Language Models Reinforcement Pre-Training

Reference 207

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:14:26.226702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T15:14:25.558878Z digest=sha256:494e815d4ba00ec1d32a60daccc57c5b54b2566619ef5e24dc6a56ba9e2d1b2b

Observation 09b938cb-cfa9-409d-9fde-7fd96d8601ba · inbound

CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning cites this paper.

CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning Reinforcement Pre-Training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T05:14:18.061552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:14:18.061552Z digest=sha256:a8973dd74c4e61b51a6cf88f145bc06b66cfd78366aa35cae4d5bb48795405e1

Observation 19deb135-22eb-4e42-a387-9f0ed445e441 · inbound

Beyond Test-Time Memory: State-Space Optimal Control for LLM Reasoning cites this paper.

Beyond Test-Time Memory: State-Space Optimal Control for LLM Reasoning Reinforcement Pre-Training

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-15T12:09:16.468055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T12:09:16.468055Z digest=sha256:ffc4d63d9e4d7f965b8a35c5fda2c8b116ce4c8c5fd18b8bd391dccead2ea060

Observation 65083332-4263-4475-9d35-f7512bdf980f · inbound

PubSwap: Public-Data Off-Policy Coordination for Federated RLVR cites this paper.

PubSwap: Public-Data Off-Policy Coordination for Federated RLVR Reinforcement Pre-Training

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:21:00.321774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:05:27.319466Z digest=sha256:5ec6eb6f3e7b02315325895d66f537dbdb31c6728cd3aa16fe1689dbe2082a8f

Observation dddbcb56-be36-4ec5-90db-16b4ac71e80c · inbound

From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space cites this paper.

From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space Reinforcement Pre-Training

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:41:04.013130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T12:50:57.603403Z digest=sha256:c2bb61dbbf6720d881e1f21e3e5cba5724f960e4e057f9ddb3a90c8edf4b314b

Observation 461999dc-ad6a-4e77-a4ae-6c712e72e1bb · inbound

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities cites this paper.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Reinforcement Pre-Training

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:22:41.994710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:edecd15fe13529de5072f36146d62a801ce66ddf7ef3e4b952c255b2da80f483

Observation 014e622b-6525-4eb9-9844-da26c2ebbd97 · inbound

Detecting and Mitigating the Correct-Answer Extinction Window in Test-Time Reinforcement Learning with Majority Voting cites this paper.

Detecting and Mitigating the Correct-Answer Extinction Window in Test-Time Reinforcement Learning with Majority Voting Reinforcement Pre-Training

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:23:06.972704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T07:20:50.835826Z digest=sha256:6251ad08b339aa520a4f83ada630abcef14f3965fdffd3dc2ba751ee8c82944d

Observation 3534117c-d9fa-4072-af1e-56b4d2aa5b6d · inbound

Value-Gradient Hypothesis of RL for LLMs cites this paper.

Value-Gradient Hypothesis of RL for LLMs Reinforcement Pre-Training

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:54:45.940210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-22T08:52:21.361294Z digest=sha256:036b75141caeb0a507933e41f98d8d471e9e4c5a40c8a282ddabd4896b90262c

Observation a758bfb8-2b50-4fad-a4d5-d89675e1fabe · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Reinforcement Pre-Training

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.441534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:799be39a87d95c1be4ae3db4f69f91d7e48221d49cb4ad3ff519c9a168b8e4e6

Observation ba6d57a6-4722-4e50-a417-4552047a6fb5 · inbound

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training cites this paper.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Reinforcement Pre-Training

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:46:29.032828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T10:37:53.152370Z digest=sha256:b9308568fb00599e5d6ab7639f9469bfc12487a98e8a1fc704fce4a11983d281

Observation 6b8fd053-aa22-480d-82f6-17d2e23ecb2a · inbound

Be Your Own Teacher: Steering Protein Language Models via Unsupervised Reward Optimization cites this paper.

Be Your Own Teacher: Steering Protein Language Models via Unsupervised Reward Optimization Reinforcement Pre-Training

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:49:02.980949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T21:40:55.946788Z digest=sha256:fd9e789e0b932b6f795d015cc577b8955b070c3b400c7e863f79c6f6af4f70bf