Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:52:05.518110Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 32 inbound Pith citation observations for arXiv:2506.20512.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:52:05.518110Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-14T04:32:13.384509Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T13:59:51.967723Z
53 of 53 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 929091d6-a92c-44a4-9033-f3f2ca02a490 · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Phi-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caadb374-c96d-43b2-b30f-7b6d072ec2ad · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58ed5096-4ca2-4ee0-a4b6-47b71c506583 · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2888274e-5224-47db-bc55-5cd049667693 · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Mathqa: Towards interpretable math word problem solving with operation-based formalisms
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a51e9972-3454-49ef-bab2-f8a1ceb5ff57 · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Jiang, Jia Deng, Stella Biderman, and Sean Welleck
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9d7d25c-fadd-425f-8568-f3ad62498718 · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Puzzle: Distillation-Based NAS for Inference-Optimized LLMs
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d657da24-bb83-4cbf-a411-07bce32b6ceb · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Llama-nemotron: Efficient reasoning models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c67a416f-8d7f-4307-8a13-004bc6d008b9 · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff991753-d2ba-4dc4-8cd8-4f15f47cc95d · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Training Verifiers to Solve Math Word Problems
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1a23e6c-7bb1-43ee-92c8-2c30c0b7df0d · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Enhancing chat language models by scaling high-quality instructional conversations
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61ecd802-5387-4e14-b5a6-6073b91ed178 · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Enhancing Chat Language Models by Scaling High-quality Instructional Conversations
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3421b259-5927-4b0a-a9a3-2ddc0dec2d6b · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Sailor2: Sailing in South-East Asia with Inclusive Multilingual LLMs
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 911403fa-58dd-47a8-b391-83cee0c03df9 · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling The Llama 3 Herd of Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 412ca939-8e09-423b-8593-cb07098072d2 · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b624733-24ad-41fa-a483-0ad778281b9f · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40e64994-d9c5-4039-a752-06cfd274aa5a · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Infi MM -webmath-40b: Advancing multimodal pre-training for enhanced mathematical reasoning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation fa3d8727-6f3e-4250-ade8-c5f6b732c88f · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Olympiadbench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71197d56-ceca-403a-ae74-519e4a708844 · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Measuring mathematical problem solving with the math dataset
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 517401c8-bb18-401d-a19c-32522e9b8007 · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f383e86c-955c-4243-92b1-b1a511ae4f9c · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0773918f-4cf9-4f32-b3c3-20a092f0ad9e · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1157c3ec-3dfb-4399-a34a-854b1e7e3c5f · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Open r1: A fully open reproduction of deepseek-r1, January 2025
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 292af8cf-d9ac-4c0d-81f2-200803b4fa1d · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Mawps: A math word problem repository
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9563caa2-1a9e-426e-8c31-6e698813945f · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b8feced-0230-4c83-9391-6ec317cb6918 · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation daa28f16-e706-4fe5-a113-4efd81675bce · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman - Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur - Ari, and Vedant Misra
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 89bb13e7-368a-4004-9562-02310017278c · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Datacomp-lm: In search of the next generation of training sets for language models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65cb26dd-5b41-4fec-8575-8a00428e08bf · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Let's Verify Step by Step
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1251579e-ea1f-48d4-92bd-5b7c03c374fc · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling DeepSeek-V3 Technical Report
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da330cee-41dc-4db6-b959-64dfd496c82b · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Understanding R1-Zero-Like Training: A Critical Perspective
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fbc964f-05cd-437c-91da-87bcd0c8392a · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation dfc3b450-3ab7-4840-85aa-39c4a4ad266d · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8928c801-905e-420a-91e4-ecaed0a0b48f · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling The Llama 3 Herd of Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5f940e9-4831-4e9f-a37e-64b29f917ca5 · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling A diverse corpus for evaluating and developing english math word problem solvers
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 71955f78-bbb2-49d3-973e-4d8085020009 · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling 2 OLMo 2 Furious
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0aaa0c9a-4954-4ec5-9107-6ce4b8f7816f · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Introducing openai o3 and o4-mini | openai, April 2025
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation cc9157ba-2e02-422b-b451-c9724d4b5243 · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling OpenAI o1 System Card
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac5fe4e2-e053-4b46-88d0-fea47c453f0a · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Openwebmath: An open dataset of high-quality mathematical web text
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e15ece67-2dbf-4414-b18a-b61522ce84a9 · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4c0b8fcf-9b99-4446-8d7e-3af410306ce9 · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Spurious Rewards: Rethinking Training Signals in RLVR
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a920ac3e-ef93-482e-a0cc-1a2b5cf27fc6 · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e4dd77a-3d4d-4073-bdba-4025517f6ad1 · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling HybridFlow: A Flexible and Efficient RLHF Framework
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a5d9039-bf41-4e53-b2ad-6382f8516dc2 · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Yi-Lightning Technical Report
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db169a52-7fcd-4069-9202-58c8caba5ea3 · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51ea2f7c-d7d4-45f9-b240-7a45841c4876 · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Mathpile: A billion-token-scale pretraining corpus for math
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 13e66756-f81a-4faf-992b-6f9c79c1b74c · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Chi, Quoc V
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4dd5087-a61e-42fc-8bda-303ebd4f342a · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 988b1b55-d734-4e77-b00a-5a9d17fdda05 · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Qwen3 Technical Report
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53ecc5cd-73f7-43d3-a95f-f1d81640c141 · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Qwen2.5 Technical Report
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c65920a3-dddc-4552-83ad-b0475a4caafa · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62f44857-132a-43e2-a473-329d03f21ba2 · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2d41fb7-5123-41f5-a502-a652bd9d0f90 · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Wildchat: 1m chat GPT interaction logs in the wild
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5362a314-eefa-40b4-be62-a8ef38279548 · outbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling MegaMath: Pushing the Limits of Open Math Corpora
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18ce7cce-fbc7-45bb-96c2-2bb845d70217 · inbound
Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d8941b72-95c4-4586-a154-9027845a5560 · inbound
URPO: A Unified Reward & Policy Optimization Framework for Large Language Models OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adbfa44c-b68c-4764-a063-580da13d60af · inbound
LinkQA: Synthesizing Diverse QA from Multiple Seeds Strongly Linked by Knowledge Points OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8317bc3-54de-4590-a75c-37df5c60a5d1 · inbound
Large-Scale Diverse Synthesis for Mid-Training OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef6e6e4c-790c-4b24-8c58-4716de72ec9d · inbound
SSRL: Self-Search Reinforcement Learning OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c659ab70-9778-4f58-b8fa-30abd22cbe4e · inbound
RaC: Robot Learning for Long-Horizon Tasks by Scaling Recovery and Correction OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13915e3a-a539-4303-9c73-15d01e11b28b · inbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 192
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67b20b37-35a5-48b2-8b64-50cd07d52306 · inbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60c30425-b1ca-4821-bba3-baf13a75143b · inbound
Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 939ad9d8-6458-47d3-a519-4c336360443d · inbound
SAM 3D: 3Dfy Anything in Images OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation dc0cabda-655d-41ec-87ab-9ac6b21e9e92 · inbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d203758-1e76-44d3-9140-7655cad53b77 · inbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5877dcad-ab33-452d-90da-745304b5d0a3 · inbound
CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2346ce8e-b95e-460a-a771-8ede395b587e · inbound
The Master Key Hypothesis: Unlocking Cross-Model Capability Transfer via Linear Subspace Alignment OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d4d9d98e-bd5b-4311-80bc-c0906307defc · inbound
From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9d4b2c58-29f4-4393-b7cd-4a463c65bba8 · inbound
Characterizing Model-Native Skills OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation bdbd9c5c-a6a3-438e-91ca-75d45231740d · inbound
Query-Conditioned Test-Time Self-Training for Large Language Models OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d703ccae-ce48-4c5e-8210-0a8781f004f1 · inbound
Query-Conditioned Test-Time Self-Training for Large Language Models OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0fc009b1-53fb-4da9-ad04-80a7ab58f589 · inbound
Knowledge-to-Verification: Exploring RLVR for LLMs in Knowledge-Intensive Domains OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6b5a8e6b-a893-4e3c-8a92-8a48f313ac67 · inbound
Text-to-SPARQL Generation with Reinforcement Learning: A GRPO-based Approach on DBLP OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3b478f30-6f23-4581-a2b3-d0ad9651221d · inbound
The Future of Facts: Tracing the Factual Generation-Verification Gap OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e758cdd0-5a2c-4c81-b43c-3917ec7ce817 · inbound
Exploiting Verification-Generation Gap: Test-Time Reinforcement Learning with Confidence-Conditioned Verification OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a3e1adbb-b97d-4e67-86f4-3e734304c20b · inbound
GRAIL: Gradient-Reweighted Advantages for Reinforcement Learning with Verifiable Rewards OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 669a75df-30e2-4d7a-831d-0b8a5667faa5 · inbound
On Advantage Estimates for Max@K Policy Gradients OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 32374512-3c5e-4da4-bf6d-26d1d08dc426 · inbound
OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 576f1bcd-e701-4314-a36a-0e07764004bf · inbound
Sumi: Open Uniform Diffusion Language Model from Scratch OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9dbf31ad-e306-4508-9fc4-435d05ab4ea3 · inbound
Reinforcement Learning without Ground-Truth Solutions can Improve LLMs OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 55efe14e-984f-4d5d-a268-308f1c1c4626 · inbound
BashCoder-R1: Towards Robust and Explainable Bash Code Generation with Robustness-Aware Group Relative Policy Optimization OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ba993dc2-d7c1-4805-bf0a-76acd9fd48c9 · inbound
Addressing Over-Refusal in LLMs with Competing Rewards OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d438acb0-8fb1-4806-bda1-c06e768a7a4e · inbound
When LLMs Read Tables Carelessly: Measuring and Reducing Data Referencing Errors OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a4ace8ce-f77b-4adc-84fb-f20c7b127fc1 · inbound
DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da7b42ed-f429-4851-b35a-c8196afec172 · inbound
LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.