Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:26:22.163198Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 17 inbound Pith citation observations for arXiv:2506.08007.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:26:22.163198Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:14:24.194857Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T23:49:02.979112Z
16 of 16 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3d8f1335-deaf-46fa-80f0-f65e345af051 · outbound
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d7b5087-d78a-413f-8956-492c0623a394 · outbound
Reinforcement Pre-Training Notably, the performance of R1-Distill-Qwen-14B is evaluated in two different manner: standard next-token prediction and reasoning-based answer prediction (indicated as ‘+ think’)
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 914df787-450c-4f09-adb8-3a345f3b7ae3 · outbound
Reinforcement Pre-Training Measuring Massive Multitask Language Understanding
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9211a4a-fdfe-4571-8125-a93116d8f10f · outbound
Reinforcement Pre-Training Skywork Open Reasoner 1 Technical Report
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e82d631-dda2-4351-ba9c-19ed2971b7ff · outbound
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c83d3403-b506-4074-8cd4-8a8e049e2775 · outbound
Reinforcement Pre-Training Scaling Laws for Neural Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b84d26e8-1849-4665-96d1-d46135f5a881 · outbound
Reinforcement Pre-Training General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 331c4ee9-c514-4eaf-9c09-c942f71f0115 · outbound
Reinforcement Pre-Training HybridFlow: A Flexible and Efficient RLHF Framework
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99f5089e-12c0-4f1e-b071-e2dbd9917abd · outbound
Reinforcement Pre-Training DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e81c777e-7081-4098-8990-1ba27f7473e8 · outbound
Reinforcement Pre-Training Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fbd83f4-6378-40a7-ab4d-98acd0087fee · outbound
Reinforcement Pre-Training A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 881d37f7-a003-4bc1-a8de-289095639b6c · outbound
Reinforcement Pre-Training Reinforcing General Reasoning without Verifiers
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa147932-a54b-4940-b749-83a0c62ea092 · outbound
Reinforcement Pre-Training Training Compute-Optimal Large Language Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5389ff90-d184-48a4-92d4-a32f79566ead · outbound
Reinforcement Pre-Training SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1448d7ff-8ad2-4796-a615-b5970a04fb3d · outbound
Reinforcement Pre-Training DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e711829c-0ac3-47de-a042-751ba8441949 · outbound
Reinforcement Pre-Training Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6367d345-5a3b-46e6-960c-b812ef785c02 · inbound
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d97bc87-0bb9-4a59-8dba-42c4713a9f58 · inbound
Advancing Event Forecasting through Massive Training of Large Language Models: Challenges, Solutions, and Broader Impacts Reinforcement Pre-Training
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df887682-effa-49cc-926f-47e4e88480b0 · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Reinforcement Pre-Training
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c37de614-28a2-4049-bf61-f6feba96f1aa · inbound
A Survey of Reinforcement Learning for Large Reasoning Models Reinforcement Pre-Training
Reference 116
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 481bd3f4-b9cc-419d-be04-f5ae99ef2ace · inbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Reinforcement Pre-Training
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9b28ba4-21ed-401a-9f34-9bec50a52788 · inbound
Refine Thought: A Test-Time Inference Method for Embedding Model Reasoning Reinforcement Pre-Training
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95dfa54e-503a-4b53-bef1-2844aa2839a1 · inbound
Agentic Reasoning for Large Language Models Reinforcement Pre-Training
Reference 207
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 09b938cb-cfa9-409d-9fde-7fd96d8601ba · inbound
CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning Reinforcement Pre-Training
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19deb135-22eb-4e42-a387-9f0ed445e441 · inbound
Beyond Test-Time Memory: State-Space Optimal Control for LLM Reasoning Reinforcement Pre-Training
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65083332-4263-4475-9d35-f7512bdf980f · inbound
PubSwap: Public-Data Off-Policy Coordination for Federated RLVR Reinforcement Pre-Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dddbcb56-be36-4ec5-90db-16b4ac71e80c · inbound
From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space Reinforcement Pre-Training
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 461999dc-ad6a-4e77-a4ae-6c712e72e1bb · inbound
Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Reinforcement Pre-Training
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 014e622b-6525-4eb9-9844-da26c2ebbd97 · inbound
Detecting and Mitigating the Correct-Answer Extinction Window in Test-Time Reinforcement Learning with Majority Voting Reinforcement Pre-Training
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3534117c-d9fa-4072-af1e-56b4d2aa5b6d · inbound
Value-Gradient Hypothesis of RL for LLMs Reinforcement Pre-Training
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a758bfb8-2b50-4fad-a4d5-d89675e1fabe · inbound
Trust Region On-Policy Distillation Reinforcement Pre-Training
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ba6d57a6-4722-4e50-a417-4552047a6fb5 · inbound
RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training Reinforcement Pre-Training
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6b8fd053-aa22-480d-82f6-17d2e23ecb2a · inbound
Be Your Own Teacher: Steering Protein Language Models via Unsupervised Reward Optimization Reinforcement Pre-Training
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.