Pith. sign in

Paper Citation Record · LEDGER

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models

As of 11 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 1 inbound Pith citation observation for arXiv:2605.08472.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.08472 v1

Coverage vector

measured 73 of 73 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T01:11:50.343466Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T16:52:29.295566Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

73 of 73 outbound references displayed

  • verified exact36
  • verified fuzzy29
  • unresolved3
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ff72c57d-92f4-4a47-8303-02b273dab3f7 · outbound

This paper cites Direct Preference Optimization with an Offset.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Direct Preference Optimization with an Offset

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:26:23.350262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:d247d6518770443b24380791c1e6066d0f4667810006728566d8851cf7bada4b

Observation 675a0eb5-0ba6-4dd0-9d49-1842335af846 · outbound

This paper cites Matharena: Evaluating llms on uncontaminated math competitions, February 2025.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Matharena: Evaluating llms on uncontaminated math competitions, February 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:15:54.253921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:0745aee38bfdde5432b2d649cca81026b92a12f02913570aff9499fcb10917c6

Observation 4965bf03-a41e-4d63-a7ce-85511b040d3d · outbound

This paper cites an unresolved cited work.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-05-14T08:15:54.257691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:c9e1ca6ccef4a68d69815fb59f99339417b9317c2aa099e0376f45136d3555f5

Observation bd904d25-fa3e-46cd-8aa9-0523b1bbb46d · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Evaluating Large Language Models Trained on Code

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:26:23.377260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:34894aacf00c62fd5cb177e2239dd599a7ac3fce0a2117829ef4b386f0a05753

Observation 0428cd94-e17c-42ca-ad2a-d88a0d704612 · outbound

This paper cites Atomic Skills are the Prerequisite: When Reinforcement Learning Synthesizes Compositional Reasoning, and When It Only Amplifies.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Atomic Skills are the Prerequisite: When Reinforcement Learning Synthesizes Compositional Reasoning, and When It Only Amplifies

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-28T02:04:12.409186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:91e0f4d25e1936e431b7ef348b08b473a053f1dab3216d6a6b8da6d241cedf32

Observation aac59382-4b8a-456a-9ae9-9605a9bb6128 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:26:23.288590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:dc2a8b9e16ccf10321257234ce29f38f6afbf712a07833a13ec94fb075988340

Observation fd2b9ec8-0272-4fd9-b465-31a9521b468c · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T12:31:34.026635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:f26c0c1a6e302c79d7ee0b44b31c49be1c8390ef44a578af4cc9b9da138c9047

Observation 4d690ac3-7f72-4110-ad4e-b70f54a63d19 · outbound

This paper cites The Vendi Score: A Diversity Evaluation Metric for Machine Learning.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models The Vendi Score: A Diversity Evaluation Metric for Machine Learning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:26:23.299429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:95feb7ea754946e2fb7e7a0679815f127454de8d82e6edef3fde12d16faefe35

Observation 59887040-9301-48f7-b223-8ffec54f4e1f · outbound

This paper cites Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:40:33.380427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:031b8d5ef7b4180e99a546a9f20a4f6631f0eabbb866182c9e15ce32c4636206

Observation 987a00bd-b665-45f1-86e2-6d4247f4054f · outbound

This paper cites The Llama 3 Herd of Models.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models The Llama 3 Herd of Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:26:23.367825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:f713004cb90e6c21160854a291507686e626ddabfaf4a621b9014c0fec238842

Observation 2589dc55-b08f-4d7a-bb31-b9f1338913fa · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:26:23.269899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:49279cc8ade12d82ecb7d9e5e617df7f1254f22639e97e9ea6f447c263379a88

Observation ad569163-ad8a-4d7d-87f1-466ebed15892 · outbound

This paper cites TarGEN: Targeted Data Generation with Large Language Models.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models TarGEN: Targeted Data Generation with Large Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:26:23.317202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:fa1d07bfeab20fea921aa5bce1ee0e103d3b11dc630e26ab9f766e74c29c595f

Observation 7789266f-52ec-495e-8127-4b7fac36b3a0 · outbound

This paper cites Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:26:23.332890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:0096ac55cdc1199d0bef58219b3104900e5ea1cf0796f6a07246220685ebc56a

Observation ce05cce0-ee35-484e-822d-823e00d079fc · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:26:23.335936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:8a13167b517a75581a8fee5ffd5c37c5c8b388df6856598400cd2325fde89cfc

Observation b35cf74e-1a9d-4ee5-ada3-20189188e3c4 · outbound

This paper cites DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:31:04.851073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:bce7fed3db333d0515da07c6b13d108bd01187fab06ae3f476d79668d7fd1f37

Observation b9752c4f-a4f7-4230-a2e6-0a434e7047ff · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Measuring Mathematical Problem Solving With the MATH Dataset

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:26:23.306433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:b4ac4c1f060315b46be137cc654653be9d22ea0dd94e6a571aec1cc325911eea

Observation 6364a0eb-cfa0-4754-91d1-44c534e3fdc1 · outbound

This paper cites Unnatural instructions: Tuning language models with (almost) no human labor.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Unnatural instructions: Tuning language models with (almost) no human labor

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:15:54.229301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:f437c632bd200b0a1d0f382ee5e13cee1656ad7e3301541ea220093708474387

Observation 3189ae89-9f0b-4cb3-9d14-6be738866498 · outbound

This paper cites Math-verify: Robust mathematical expression evaluator.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Math-verify: Robust mathematical expression evaluator

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:15:54.179887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:d7cdb9e21516e19714c59ca7a43e04195e791fdeb459164d04d01b4f237bc1d8

Observation 2100ae73-85ef-4301-afc0-06b2ad90b92d · outbound

This paper cites GPT-4o System Card.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models GPT-4o System Card

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:26:23.259514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:6c3aafc01ae62fc1f31213a77b6d6af9ebf3c073fffef0e80cecd39b470c67ef

Observation a422bd89-057c-47b4-bfda-9a446b4402be · outbound

This paper cites OpenAI o1 System Card.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models OpenAI o1 System Card

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:26:23.262633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:1635aef4261a80408a5710dce45466fff22ed30a59aee70d7a595bc51b61c532

Observation 9e8e7edf-f295-47ff-9b1f-1ba7ae188728 · outbound

This paper cites The Art of Scaling Reinforcement Learning Compute for LLMs.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:29:14.085942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:555f5d1e1f623e5751e3316b683577f89d979a2b103ad0852997178f7c375acb

Observation 1eeba29b-2ec2-405c-85f1-63dc2c17debc · outbound

This paper cites Spoc: Search-based pseudocode to code.Advances in Neural Information Processing Systems, 32.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Spoc: Search-based pseudocode to code.Advances in Neural Information Processing Systems, 32

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:15:54.184209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:116b7f688387b9ee65a102bfc80342c2efaa3f6713f7fd3b6af5a8bfb315b30a

Observation 7788295f-2a5b-4967-9d86-56a858c25aae · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Efficient memory management for large language model serving with pagedattention

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:15:54.203126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:bb0c5006b483be231ae3694fdd3255a045381d628efd3af9645db702999b1e75

Observation e145583c-698b-40c5-8677-aaac3491cc14 · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:58:29.262663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:3a2fc11fde17f57e3b8c16420401ff5144b56b4cf7e3cb8aea6883454bc73086

Observation 5f346081-10c9-4510-91fd-29809d2213ee · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:26:23.284596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:4a72e0da37fac734f43ecf2cc2a50a86b7be57a841e3403f15108f5b1030ea78

Observation c03e7620-157c-4327-bd61-18db927190b4 · outbound

This paper cites The measurement of observer agreement for categorical data.biometrics, pages 159–174.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models The measurement of observer agreement for categorical data.biometrics, pages 159–174

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:15:54.198739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:87f411972d27f34aaaf2e3e265e516b5b37fe9f9babad5a679bda157c6f5dcd4

Observation c99a57c5-8e34-4783-a228-9d9191655941 · outbound

This paper cites Small models struggle to learn from strong reasoners.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Small models struggle to learn from strong reasoners

Reference 28

Resolution
verified exact
doi, observed 2026-05-12T01:16:15.422726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:57e91e40eab36d316f6ad9ba539fdbfa31ea7e2ee7cd93a3b24575bf30f792e2

Observation 6edb31ee-ce2f-412a-acef-be306d12ea27 · outbound

This paper cites Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:17:43.404381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:d17f9a419f9fc6daf34f6f31da6be4563530c49c456a3e0f9ede51d9cbfb96f9

Observation 7b989002-0fcf-4191-9927-ff9589f81044 · outbound

This paper cites Simpo: Simple preference optimization with a reference-free reward.Advances in Neural Information Processing Systems, 37:124198–124235.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Simpo: Simple preference optimization with a reference-free reward.Advances in Neural Information Processing Systems, 37:124198–124235

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:15:54.188956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:8a43fbd1107967c40fb9186a8c5c61f1d6f7542617c4a31630cc148ac81c2600

Observation 4592ac03-26c8-4464-8822-5e5471a9ba26 · outbound

This paper cites Cross-task gener- alization via natural language crowdsourcing instructions.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Cross-task gener- alization via natural language crowdsourcing instructions

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:15:54.165918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:f6123e66a69cb8ef69f97faf0391daab8da098cd8b9e99a02e3a31d1423e1962

Observation 01e7acb6-7832-42c2-b2f0-a548372fbe6e · outbound

This paper cites Orca 2: Teaching Small Language Models How to Reason.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Orca 2: Teaching Small Language Models How to Reason

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:26:23.266438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:69631ab203771988fd5bf8ace0cb7d5a9a446d421a8fb85de146f2a1acd252ac

Observation 29024956-17e6-4c5b-8957-bb8271e22e91 · outbound

This paper cites Mid-training of large language models: A survey.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Mid-training of large language models: A survey

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:26:23.339800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:c2200058dd8599a7a30bab2a1202eb4b1e32a15c1ca73e1cb580d785a1a4aa6b

Observation 4b71ea29-b3bc-4005-bc28-ade95fcd600d · outbound

This paper cites Orca: Progressive Learning from Complex Explanation Traces of GPT-4.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Orca: Progressive Learning from Complex Explanation Traces of GPT-4

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:42:05.373278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:914d0ad68e7c2d777a719bdab169c68b687b4d7e9152ddb10fe4db0403c54d72

Observation 44ead73a-9baf-47a0-83d3-0f7efcb710e4 · outbound

This paper cites AMC 12A (2023): Problems and Solutions.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models AMC 12A (2023): Problems and Solutions

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:15:54.242758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:1ba8350cc8d932d91ba6c5d2b7d6e9f8578226e1ca72b847ee8627ac88ba7876

Observation 499f4c98-07c5-4171-a6a8-4fd8ecc010a3 · outbound

This paper cites Olmo 3.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Olmo 3

Reference 36

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T08:26:23.370856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:86ffc32d2c693c53754267371da44f28e86cf412899e5f47edc0898610452240

Observation 38add820-83e4-43b8-a773-ecc1659b13b7 · outbound

This paper cites New embedding models and api updates.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models New embedding models and api updates

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:15:54.225268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:a3df5040003cebd5442a5575528ed022333fc702b5ced915fd742ce690f63411

Observation 5fa81818-66f8-4cf8-ba06-d13b7452840a · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:15:54.155968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:15da8bbb9309bbe0d0ffad2a462250fcc48bda1fbe778bbc7dc2312a4f7fb9d9

Observation 4f51d288-cc4b-449f-bd5f-b5ac249f82f6 · outbound

This paper cites How many data samples is an additional instruction worth? InFindings of the Association for Computational Linguistics: EACL 2023, pages 1042–1057.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models How many data samples is an additional instruction worth? InFindings of the Association for Computational Linguistics: EACL 2023, pages 1042–1057

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:15:54.208003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:186ead9d9293d7d05e931227d73232afaaf16ee7de232caa564ec8e1daf11658

Observation a6e04f74-1168-4fdf-bafe-ff639c8e4d03 · outbound

This paper cites Princeton science library.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Princeton science library

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:15:54.151240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:fb917a092046549a9d752dc581b4c03e7e1c700cb9cb89084784d7d51eca648a

Observation 7deff357-4d27-4d81-a351-ee0d43279d2d · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Direct preference optimization: Your language model is secretly a reward model

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:15:54.161016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:cb40a042b6e937e33b0355578691d137c9b3415fcbaec59f32be76a4f7122c69

Observation a82b28cd-d1b4-458f-ac9b-ae3772e00f66 · outbound

This paper cites ThinkTuning: Instilling Cognitive Reflections without Distillation.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models ThinkTuning: Instilling Cognitive Reflections without Distillation

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:26:23.380848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:f180bb66f6bac0a4f18a52f4df64e5e0cb39d612044687ad5160706fde980953

Observation 3be10a95-fe9c-4d2a-a596-1b5b5e479f9c · outbound

This paper cites Triple Preference Optimization: Achieving Better Alignment using a Single Step Optimization.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Triple Preference Optimization: Achieving Better Alignment using a Single Step Optimization

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:26:23.346238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:247ae63bb6efb5ed1cc48b4e8dbd9cbc92fc866a738745597b21fc4b54a3d347

Observation 62eb4ce1-8673-4f51-bc0c-db154e5e3fec · outbound

This paper cites Proximal Policy Optimization Algorithms.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Proximal Policy Optimization Algorithms

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:26:23.309669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:2c0f50797b8a6aaf178e087281855e2e53e113185a6ca31b49ef886e08a5dda4

Observation 6171f08d-2732-40ad-a371-15a5c4404548 · outbound

This paper cites Rl on incorrect synthetic data scales the efficiency of llm math reasoning by eight-fold.Advances in Neural Information Processing Systems, 37:43000–43031.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Rl on incorrect synthetic data scales the efficiency of llm math reasoning by eight-fold.Advances in Neural Information Processing Systems, 37:43000–43031

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:15:54.175437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:ad34e3fd006bef992f40002c658744d7d9e6a66a1235d9297b61fabd67523d30

Observation fd06ad82-9169-46d6-904a-299687785907 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:26:23.373677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:612e1981d4a1f887297e499b6594616c6ba2656ef22deed6ed8a0d85a4d57018

Observation cb78f1e9-767c-4971-b5e6-ed049cea6f10 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models HybridFlow: A Flexible and Efficient RLHF Framework

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:26:23.295282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:f5b2a1a0ac562af16531d16bafac93481159b0b3ed4ba0c3d1b8e660ee566a9c

Observation 0f896885-a606-47d4-a270-dd718fc19e4f · outbound

This paper cites MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:26:23.291935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:478814c60b5eebd5ad6b3b78e41a008d9830483a695c22db030c91693cc0e30d

Observation d8c18b49-bfd4-45a2-b077-4d74ae59e0e7 · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:15:54.142446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:46e3d89e9622ff212e40e913cfdeb95a799cced2b6f1e642c81b55ba12caa43f

Observation c1d08b41-a72c-4882-9aa3-fef12409b799 · outbound

This paper cites A survey on llm mid-training.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models A survey on llm mid-training

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:26:23.313398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:c2fd91fa5157c75a562f1d10aa0e4ccc021119ca6396bd93b4d9e803ebf42d91

Observation 91c3b67a-ed70-4aec-a4cf-f7afc2a807b7 · outbound

This paper cites Trl: Transformer reinforce- ment learning.https://github.com/huggingface/trl.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Trl: Transformer reinforce- ment learning.https://github.com/huggingface/trl

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:15:54.221282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:e7c32fd419ae1e10f2068e558b6a4188840d8d4466513ce0596cdfcde497941c

Observation 20af2015-3b07-41d5-ae74-883752cb9b33 · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-12T12:12:09.004283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:5e62d0ef49e371f82f442b5a63738d5cec3e315d39cd4a3f1eec192e1a679228

Observation 4a3322ce-3a85-4435-a2b6-27c1ac169c21 · outbound

This paper cites Self-instruct: Aligning language models with self-generated instruc- tions.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Self-instruct: Aligning language models with self-generated instruc- tions

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:15:54.216521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:6832bfff958d6deb671f790490f90bb344bcbf000d348987668c4489cbdef991

Observation 1b9c4119-405b-44cd-9442-0c73ed173a06 · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Finetuned Language Models Are Zero-Shot Learners

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:26:23.342903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:f040cd3ca96993ccf66b05f83d699aaed194f4299067ffe43679b070b1d9b4e9

Observation b2c38db9-a19f-49f2-9d9d-ce3d189c299c · outbound

This paper cites Williams.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Williams

Reference 55

Resolution
verified exact
doi, observed 2026-05-12T01:16:15.425980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:c82af155f614bfac1e26618b8ddce0443d818e3be2fae790bffb5d5d97c146c9

Observation ca29502b-3249-492b-b520-311988ac9f61 · outbound

This paper cites Wizardlm: Empowering large pre-trained language models to follow complex instructions.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Wizardlm: Empowering large pre-trained language models to follow complex instructions

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:15:54.133754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:2bdae2c0533270260e22d3ffe5dc4043e9e81fba804078eafa261a8135d661ab

Observation 3517e1bb-4afc-45e7-b370-e2d00fa5ba8c · outbound

This paper cites KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:26:23.303325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:ad27442d3dd221a48b5a9a538f43be14be0c660b78b8daf8e0025cec051a02a3

Observation b486751e-90e9-42e3-b462-095e38d3da75 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:26:23.277909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:c5450ca42c5e4fc5500e1b492d7602094a2158b52b98e1e64708a47e90cb6c5e

Observation 750955e6-33f0-481d-a306-9c33e41762f0 · outbound

This paper cites arXiv preprint arXiv:2509.25123 , year=.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models arXiv preprint arXiv:2509.25123 , year=

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:26:23.357265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:19b1600b76c04180923614d33a3dc9fc91b400805e4abf97587c4e06a29a1964

Observation 1080d409-8863-4337-be3d-ec58246afb59 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:26:23.360380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:65cb246212229f3d989cab33bf7a1eeef9a67a2423b638f3f9a5189a3eacad66

Observation 4b7d1b1c-1615-47e0-94bd-54169a493084 · outbound

This paper cites Star: Bootstrapping reasoning with reasoning.Advances in Neural Information Processing Systems, 35:15476–15488.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Star: Bootstrapping reasoning with reasoning.Advances in Neural Information Processing Systems, 35:15476–15488

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:15:54.138440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:697e746840fd5a9b90ddeee44dbac02c781b80f1bd104484250d6fb60c88f9e1

Observation c7630bfd-da44-4e11-a038-cf37e83d1801 · outbound

This paper cites On the interplay of pre-training, mid-training, and rl on reasoning language models, 2025 a.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models On the interplay of pre-training, mid-training, and rl on reasoning language models, 2025 a

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:26:23.388268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:eb4f189c85db64bd7c5af0f86f71814edcd99c3a8a7d35591bc7ebce46e85841

Observation 02dc47df-a062-4b6c-b96a-c51be6529c44 · outbound

This paper cites American invitational mathematics examination (aime) 2024.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models American invitational mathematics examination (aime) 2024

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:15:54.237942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:ca5a48ace4c2527d9881f6ea224c48a076fe895e973dbfc95005675bcdd19d94

Observation 1983d042-f39c-4416-b394-30a0e1774b7c · outbound

This paper cites American invitational mathematics examination (aime) 2025.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models American invitational mathematics examination (aime) 2025

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:15:54.125582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:59b2ad735f09e6c8d62f11357351a4afc7791ce134184e50b4af3c62a003d57f

Observation 335150b7-a233-4aee-9af1-253c8b86fb5d · outbound

This paper cites Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning

Reference 65

Resolution
malformed identifier
arxiv_id, observed 2026-08-04T02:23:17.018487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:8a35db014aa7d583b7ae2664c9f2063248014fc440f04c0ac99708a492a51ef7

Observation 9d044ff9-65c6-4755-9401-dcf9d369aca8 · outbound

This paper cites an unresolved cited work.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-05-14T08:15:54.233028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:8409f0d7962a2b2d02dea28333e03a540fe63e686b533bae45fb055193e6d207

Observation f2abb68c-3cf2-4e90-88dd-75cc25ee3e7c · outbound

This paper cites an unresolved cited work.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-05-14T08:15:54.250175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:8279bbc449fdbf36dec299f9ec05e20faaaee06e6fd25b49a4a6a0c51b94c51b

Observation 5fcea277-bb05-4a79-8c2c-50ac61738c24 · outbound

This paper cites Decision: Yes.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Decision: Yes

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:15:54.120885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:9b1e367d9ce8dc03a4fa03465b90196d8f2738a08be1c67d8d56728eae625000

Observation c6405a68-7b39-448b-877c-b50aebbc9754 · outbound

This paper cites determination hope success.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models determination hope success

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:15:54.129552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:3f800e2709bd0d07e199f4da31b5bba8365d18e3590b4a032a56cabb945374b9

Observation 609123af-37b3-42ed-901e-e3c921da2c54 · outbound

This paper cites All the given conditions are satisfied, and the context makes sense.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models All the given conditions are satisfied, and the context makes sense

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:15:54.246962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:89803bbc577c0d8e9c72c02eacd147b951a1aa37a103b2e0398b3fc23c74c9bf

Observation 38a02fda-7189-4293-a34a-e5d91845e667 · outbound

This paper cites Pappus First construct or create something that helps you explore the problem.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Pappus First construct or create something that helps you explore the problem

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:15:54.194464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:8511744bcb71f9125a9b4be890f85056bffbd9e881ae58c1968146b6f463e078

Observation 59137995-f6b9-4baf-988f-37c91dd175d5 · outbound

This paper cites First, find what seems to be the right answer or pattern by exploring ex- amples, calculating, or testing possibilities.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models First, find what seems to be the right answer or pattern by exploring ex- amples, calculating, or testing possibilities

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:15:54.212244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:11d7fde85174a0a00ba6ec1cdcafbda0a140e82212cbddc5e8a93acbb46a7b36

Observation f095898f-082b-4c6f-900d-574fdc0f1fcd · outbound

This paper cites So, this possibility is inconsistent with the problem statement.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models So, this possibility is inconsistent with the problem statement

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:15:54.147173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:cc530b920c9cef98be70e05255defefb2b0f601ac7cc699a36f380aeb73ac550

Observation 84da0b5a-0dac-461e-a3ae-39c5f09f4c8d · outbound

This paper cites Now, adding both months together, 48 plus 24 gives 72.

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Now, adding both months together, 48 plus 24 gives 72

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:15:54.170472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:11:50.343466Z digest=sha256:ea08f3f685f5d26e20513bae925c9d2dc859b6849124abb976d625e390f746b1

Pith citing papers

Observation 8bbcf149-05dc-45c5-8b39-70200dde9818 · inbound

Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution cites this paper.

Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T16:52:29.295566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:52:29.295566Z digest=sha256:502ad450e3b5128605fa4b3fe54b9c47c04d9c27c79384600b213e00edb7c29e