Pith. sign in

Paper Citation Record · LEDGER

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling

As of 10 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 3 inbound Pith citation observations for arXiv:2509.01649.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.01649 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:25:28.508547Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T03:02:07.425724Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T05:49:40.779767Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact4
  • verified fuzzy16
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bebfc12a-0e3c-4625-95a5-75834a84312f · outbound

This paper cites write newline.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:22.958457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:22.958457Z digest=sha256:6bb97f3b613b2f11c721ea5c66f4214252a3a2dc6535de74b0bb685a411f5361

Observation 170ad7b5-b0ea-44cd-8d91-ef4ff033c4fa · outbound

This paper cites write newline.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:23.019471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:23.019471Z digest=sha256:9f3947514b364458237fec6d157f3392eac9e406e785eec342167b9fe1e4ca1b

Observation f6ef6a1f-07ce-4a2a-9ca5-c55a343986d3 · outbound

This paper cites @esa (Ref.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling @esa (Ref

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:23.125793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:23.125793Z digest=sha256:89b1d74a265e0c4964bc5502cca325603833e008d6aa87717948f8dd5421aad7

Observation 73c6fc5c-8d25-41fa-a7e3-621d5c5e05e2 · outbound

This paper cites an unresolved cited work.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:23.209699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:23.209699Z digest=sha256:b0487e70a69e7057768041b8ea5e1f3a451f77456cc7cae57896a37a9a20ff06

Observation f9d525d0-7efe-4c5a-8840-c8115731ff7a · outbound

This paper cites an unresolved cited work.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-05T12:25:29.878023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T12:25:23.286873Z digest=sha256:5b49e5d3d95468ea106f6c48b0621f958f60ec404375003688f94eba18b1d6eb

Observation 13ee03da-a6ba-4052-8e00-1052b8bd690e · outbound

This paper cites On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:23.384000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:23.384000Z digest=sha256:b24882c38c2e5b39344c49f47dc98c338528e3da683199a633a150f88d5f9ffb

Observation 0f4d3823-fe0f-494b-b090-8a614bc9f5fc · outbound

This paper cites Alphaevolve: A gemini-powered coding agent for designing advanced algorithms, 2025.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Alphaevolve: A gemini-powered coding agent for designing advanced algorithms, 2025

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:25:29.864076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T12:25:23.494428Z digest=sha256:a0d0f96e8c42301af7c9fa39b031d24b6a5c2aae75a7c0fdc45b8124e4900804

Observation b877d89e-b5d5-443a-b6d2-bdf26a6d18a4 · outbound

This paper cites Program Synthesis with Large Language Models.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Program Synthesis with Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:23.571192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:23.571192Z digest=sha256:e5210aba6095786f50f0aba3a32d382b6aac7b6980c72c6b5beaeb1900631aeb

Observation 7a7d6cf2-48c2-4210-b588-804baacad62c · outbound

This paper cites Jiang, Jia Deng, Stella Biderman, and Sean Welleck.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Jiang, Jia Deng, Stella Biderman, and Sean Welleck

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:25:29.849538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T12:25:23.667642Z digest=sha256:c306af6240f23e57fdfbef594cb50f92fede23fa74a7dab950dad7a092ecb1a1

Observation d9d4f780-ffd3-4fc5-b148-a943f3ff6a63 · outbound

This paper cites Do deep nets really need to be deep? Advances in neural information processing systems, 27, 2014.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Do deep nets really need to be deep? Advances in neural information processing systems, 27, 2014

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:23.760335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:23.760335Z digest=sha256:d9cdcf91620bdc8e333d1475aaf0a1c9a66bd7d992c72ebdff7f98c9637babc3

Observation 37f6152d-fbcd-4502-9ded-cd4546b6deb2 · outbound

This paper cites Scaling test-time compute with open models, 2024.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Scaling test-time compute with open models, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:25:29.824431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T12:25:23.842264Z digest=sha256:79b3386e8729ec97511625e7e2ac8b2e159a5ea07cdf18b2553107a96814d29b

Observation e31731bc-8330-4580-aa69-80e800748b77 · outbound

This paper cites Knowledge distillation: A good teacher is patient and consistent.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Knowledge distillation: A good teacher is patient and consistent

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:23.941000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:23.941000Z digest=sha256:cc6547c04962f5d3bba0bb8a31943dfc80d73c622a8be1fa5e0a984c748eabb8

Observation d4477826-47e3-49cd-bffa-20fb3c8a19ee · outbound

This paper cites Birth of a transformer: A memory viewpoint.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Birth of a transformer: A memory viewpoint

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:25:29.808008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T12:25:24.039161Z digest=sha256:1af9590097dfcb9656fdeab30ffd55ff4246977fa8e1f090fce47d39381930fe

Observation 88af22cc-8883-4c65-8514-7f7a59104101 · outbound

This paper cites Model compression.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Model compression

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:24.111826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:24.111826Z digest=sha256:87d54cbaacc299cc732658faa4513703d95af3be91dd26e46606cf49fbdc71da

Observation d9287584-c60c-4e1f-b10f-1b5121db1396 · outbound

This paper cites Distillation Scaling Laws.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Distillation Scaling Laws

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:24.197937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:24.197937Z digest=sha256:d0a179294447111b5e6a6bf549177a20853fe25cad3ffad7898cb108473c1a6f

Observation 13688017-0268-4d3b-aba3-de9cf28eacea · outbound

This paper cites Why knowledge distillation works in generative models: A minimal working explanation.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Why knowledge distillation works in generative models: A minimal working explanation

Reference 16

Resolution
verified exact
raw_fallback, observed 2026-08-05T12:25:29.535717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T12:25:24.290047Z digest=sha256:8cbee6142cb84adba2ff92276e7dadcfc8819ff5a4409283ae66eae8c3957ad6

Observation 8a84ea49-d432-49ff-813d-59b1cf6abef4 · outbound

This paper cites Rethinking fine-tuning when scaling test-time compute: Limiting confidence improves mathematical reasoning, 2025.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Rethinking fine-tuning when scaling test-time compute: Limiting confidence improves mathematical reasoning, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:24.382110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:24.382110Z digest=sha256:3dacc9b6fb427b3340431b44e1dda958d759a3c70c841034c4db9f88b201389d

Observation 3751ea62-5e02-48cc-bd39-ce387294c4a2 · outbound

This paper cites AlphaMath Almost Zero: Process Supervision without Process.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling AlphaMath Almost Zero: Process Supervision without Process

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:24.465741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:24.465741Z digest=sha256:8d029cbb5bc17134cc6dc268c0be0b4f0011d34f5f25dbce21d3f198a46f99cd

Observation bd39482b-8f99-48b4-b476-9d495338e925 · outbound

This paper cites On the Efficacy of Knowledge Distillation.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling On the Efficacy of Knowledge Distillation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:24.562713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:24.562713Z digest=sha256:e0f8e205846f96fdc098b4d5ee59fd395ddc8d732b3fa6e5722c885889dff1d2

Observation 868e1558-4f58-4504-a633-676f96595b73 · outbound

This paper cites Inference-aware fine-tuning for best-of-n sampling in large language models, 2024.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Inference-aware fine-tuning for best-of-n sampling in large language models, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:24.661688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:24.661688Z digest=sha256:580b911d2aa9b9f0a11f79b2e22d8f39898e46ce991826965b5e391cf949dca2

Observation 8c19ebc7-857c-426c-9bde-ab3eaeeee94c · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Training Verifiers to Solve Math Word Problems

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:24.753080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:24.753080Z digest=sha256:dada86361c95cc3dd8b12b31cdf27924d526793f37f853e8fc1afcb20981ce95

Observation 40bcfe5b-aacf-4f6e-8f25-f3186fb1aa5e · outbound

This paper cites Weight ensembling improves reasoning in language models, 2025.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Weight ensembling improves reasoning in language models, 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:24.878084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:24.878084Z digest=sha256:d2a17e9e946f513c59174d99053edcb48787e04cd4e1a7e2c418b947c2e48243

Observation 2331a9c0-d643-47c0-a4e9-eeb4b0349675 · outbound

This paper cites DROP : A reading comprehension benchmark requiring discrete reasoning over paragraphs.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling DROP : A reading comprehension benchmark requiring discrete reasoning over paragraphs

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:25:29.782758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T12:25:25.003825Z digest=sha256:f0991a31f5da52894b030a2309f3bcce681d7029738ea3b9ebe0d2d4cdcf68ba

Observation 3162a577-bca9-4c0a-891f-828d349a23c3 · outbound

This paper cites The Evolution of Statistical Induction Heads: In-Context Learning Markov Chains.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling The Evolution of Statistical Induction Heads: In-Context Learning Markov Chains

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:25.085058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:25.085058Z digest=sha256:0f105e2def9dcf7bab59697a69100f49938e3f0011cf84ae3f29f4099baff7df

Observation 0b7e0d21-d7d3-4292-833c-b36f0f938a44 · outbound

This paper cites Born again neural networks.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Born again neural networks

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:25:29.768885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T12:25:25.189309Z digest=sha256:f94e0e55a975555d1be7d2f02e31b85a23d4c970b655a85e04a04ebf2f342ed9

Observation 7740d992-e655-4e5b-bcf1-92634caac227 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:25.292590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:25.292590Z digest=sha256:c93e371ff5fb69af18f804dfb9ac80a6df8ca10a6183e805af3e7506c0b6884f

Observation 63fa1951-723a-4c0d-b665-b4b1e524d4a2 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Gemma: Open Models Based on Gemini Research and Technology

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:25.407357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:25.407357Z digest=sha256:90263344bd44be128ef51485218b8e68f8d2f0365190ea36d56d712912853368

Observation 15800d26-2908-495b-b1a6-dc7e07adcffe · outbound

This paper cites Gemma 3 Technical Report.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Gemma 3 Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:25.533802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:25.533802Z digest=sha256:450be746e09c63ea8abaa925cf54d5c989b556bc6f06e6bec79a4b234a8df965

Observation efb07af1-884e-43c0-9ec7-f8dbe9355320 · outbound

This paper cites Multi-Token Prediction Needs Registers.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Multi-Token Prediction Needs Registers

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:25:29.120281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T12:25:25.612205Z digest=sha256:e21fd2a6f9707a9c7e9882c32561ed1a2a2f932eb3d581baff25fc577724a150

Observation c07f773e-84fc-4187-a908-a2a26fe6a00a · outbound

This paper cites Better & Faster Large Language Models via Multi-token Prediction.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Better & Faster Large Language Models via Multi-token Prediction

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:25.693098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:25.693098Z digest=sha256:04396feb5fb9d749dc3e99b228910e669b3e0c69630b82e123c63f5dbcb45a47

Observation 916976b7-535c-48aa-a155-adb098a62325 · outbound

This paper cites Knowledge distillation: A survey.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Knowledge distillation: A survey

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:25:29.754807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T12:25:25.826857Z digest=sha256:12141cc5934151b5a9361195fd06e08e6711771fce8178589e1abd3efca1ceb1

Observation 950d04fa-88ce-495d-9437-76ff921952f3 · outbound

This paper cites Context-Parametric Inversion: Why Instruction Finetuning Can Worsen Context Reliance.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Context-Parametric Inversion: Why Instruction Finetuning Can Worsen Context Reliance

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:25.983146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:25.983146Z digest=sha256:4ca8e50cc30fc2a092d092ec0efb9b912c75c513729c8816c4bb442c5ec1f8e1

Observation a16dd5d6-37af-4f07-bb5b-3bcafd78c977 · outbound

This paper cites MiniPLM: Knowledge Distillation for Pre-Training Language Models.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling MiniPLM: Knowledge Distillation for Pre-Training Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:26.071158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:26.071158Z digest=sha256:75e467e5a29f30da9394547b5a2883d0d17470b99a93b7d83549b4468ca2db67

Observation e3603938-50c9-446d-9107-271dc73a8ef6 · outbound

This paper cites OpenThoughts: Data Recipes for Reasoning Models.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling OpenThoughts: Data Recipes for Reasoning Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:26.151714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:26.151714Z digest=sha256:ceb6267c15ef8debafffe00b876b3d64d6009a6b5bccb1b6536d25ea45b0e078

Observation 6a2fbe04-2c5c-4787-8fd7-a197b2a4f0e3 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Distilling the Knowledge in a Neural Network

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:26.237527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:26.237527Z digest=sha256:c8cb7b24d504a7c3b226a1f51bee3eb0601ddba22aaa28ba428df6172cb36bd8

Observation 319d01d5-180e-4cec-b10a-5bd0d5c4394d · outbound

This paper cites Babilong: Testing the limits of llms with long context reasoning-in-a-haystack, 2024.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Babilong: Testing the limits of llms with long context reasoning-in-a-haystack, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:25:29.740240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T12:25:26.317897Z digest=sha256:ea86ac8121435ab92a0bb89999ff5764a1815c4adf7960263be0d46154000b67

Observation 9b800b41-3613-4b6f-94ae-bf1e4b6442ab · outbound

This paper cites RACE : Large-scale R e A ding comprehension dataset from examinations.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling RACE : Large-scale R e A ding comprehension dataset from examinations

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:26.460696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:26.460696Z digest=sha256:fe7dbddc9296ed466c3aace71db10cee35e070a16e65fea6999242f8df51904c

Observation 8f8129e7-95e5-4f7f-bfa2-116c3889867b · outbound

This paper cites DataComp-LM: In search of the next generation of training sets for language models.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling DataComp-LM: In search of the next generation of training sets for language models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:26.534257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:26.534257Z digest=sha256:f0b2287885679e8ed28dbaa64e572a7963b067eb7c8cb3736d4a9f5c5a9a0f94

Observation d1bfaa4f-9ee2-4737-9226-6586500a26b9 · outbound

This paper cites Dynamic Knowledge Distillation for Pre-trained Language Models.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Dynamic Knowledge Distillation for Pre-trained Language Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:25:28.999083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T12:25:26.646130Z digest=sha256:892045be5f5b0953ee322e3e1397ff393c51b3a8c59e133e14cced947fd204d1

Observation 75528aac-f872-4454-b9ca-66c4769d2145 · outbound

This paper cites Multi-Agent Verification: Scaling Test-Time Compute with Multiple Verifiers.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Multi-Agent Verification: Scaling Test-Time Compute with Multiple Verifiers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:26.743517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:26.743517Z digest=sha256:cd3a40af230258a96ea6b8d4cdd90e410e35d741dba272ef5548c02ccc0fa19d

Observation da525c05-8ae1-47c4-afa0-52d6583dfb8e · outbound

This paper cites Let's Verify Step by Step.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Let's Verify Step by Step

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:26.814614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:26.814614Z digest=sha256:6bbc120e33a67f862f6037424745cff0fd23592f750c10c5d3236ba1dac0df8e

Observation 1dfec237-e73a-4b93-9c3c-da7387c8bdb7 · outbound

This paper cites Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:26.957355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:26.957355Z digest=sha256:32307633ca41f5b520a05e3869f740139e9e7a61c6bf7d626abb9a9ddec35a56

Observation f66d842b-8d0d-4730-8b0f-226a3dbc9704 · outbound

This paper cites Unifying distillation and privileged information.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Unifying distillation and privileged information

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:27.075470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:27.075470Z digest=sha256:8e884c5ce9df53b3aae7bbefe07a2d06e25d49f71c80b7ec2f27364a8544893f

Observation 11ea56a0-aa12-4035-9a03-e35ba87fecee · outbound

This paper cites Fineweb-edu: the finest collection of educational content, 2024.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Fineweb-edu: the finest collection of educational content, 2024

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:25:29.725818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T12:25:27.160458Z digest=sha256:ef876529f841c3d635f0953ced6835f45cc164976434a98b1f5dd3c9af1bf52e

Observation c30b19de-14d5-441a-bb7b-ba5a9fbb23a2 · outbound

This paper cites A statistical perspective on distillation.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling A statistical perspective on distillation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:25:29.711070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T12:25:27.208168Z digest=sha256:61b29000611448d57cac2c7effa36e21f0f115494690281055132226873d6fb3

Observation a0767308-b056-4601-8786-1bf4d746dc22 · outbound

This paper cites The llama 4 herd: The beginning of a new era of natively multimodal ai innovation.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling The llama 4 herd: The beginning of a new era of natively multimodal ai innovation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:25:29.695815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T12:25:27.320379Z digest=sha256:b7453446e91a4c315238595e6dff0a6b08dd0ed21c232d7f5a46bcc32b33987c

Observation b85ab490-80ee-40a4-a470-343191edce28 · outbound

This paper cites Llama 3.2: Revolutionizing edge ai and vision with open, customizable models.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Llama 3.2: Revolutionizing edge ai and vision with open, customizable models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:25:29.678919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T12:25:27.424089Z digest=sha256:f2eabc50df0948e5b5774113c16d64b39769672a3c3d6aabba09967bff446e81

Observation 557e08d9-99b3-4932-9822-923375655c2b · outbound

This paper cites Improved Knowledge Distillation via Teacher Assistant.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Improved Knowledge Distillation via Teacher Assistant

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:27.560427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:27.560427Z digest=sha256:aa92714ddcefceb10ed9f5d4b6082600ae3ff32f98835f3017ad12cb4bec1850

Observation 905b4434-242f-48ea-a140-62e4deacf45c · outbound

This paper cites Improved knowledge distillation via teacher assistant.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Improved knowledge distillation via teacher assistant

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:25:29.662757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T12:25:27.725332Z digest=sha256:a3441274969a93111bfbb088984f5c751371410f85daf19efb997c13ee2cf490

Observation 8e753b6c-bbe5-4a56-bbd5-639b7da58a5c · outbound

This paper cites Self-Distillation Amplifies Regularization in Hilbert Space.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Self-Distillation Amplifies Regularization in Hilbert Space

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:27.847217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:27.847217Z digest=sha256:649e47f1e697b4629ae7b88d1f579af0713f4ff5302381b6b4953832736c42c1

Observation ad9e1f80-2dc3-4368-91a8-1417616395fd · outbound

This paper cites s1: Simple test-time scaling.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling s1: Simple test-time scaling

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:27.955742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:27.955742Z digest=sha256:6f1f03ee9caafd3389f5d25c3216f44c5798ba318d9c2cad1fb45e3ed36fdcd4

Observation 4abc85d8-de09-4335-b395-1adace19a8e8 · outbound

This paper cites On student-teacher deviations in distillation: does it pay to disobey?.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling On student-teacher deviations in distillation: does it pay to disobey?

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:25:28.870109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T12:25:28.121806Z digest=sha256:9d531551edf2f720733df46f07cdb461d08717914dac7b6f7b25a35eef1c21c9

Observation f0d95de9-d0a1-4a0b-930c-063cb50ab965 · outbound

This paper cites Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:28.192200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:28.192200Z digest=sha256:44762ef34afc2d35eb2d3f5bbecf15c93eb602e7f505200c4afb79065cac8a78

Observation 17dc5935-242b-443f-80fa-80f87e8db6ee · outbound

This paper cites Github code dataset, 2022.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Github code dataset, 2022

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:25:29.649243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T12:25:28.300845Z digest=sha256:12711e47df04857edf52cbd4f7fd8d5327fb91b5c9dc4d0d62c1cc9975898487

Observation 3b8cad1c-5e9b-4692-a04e-fd7d04d2e1d5 · outbound

This paper cites In-context Learning and Induction Heads.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling In-context Learning and Induction Heads

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:28.393471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:28.393471Z digest=sha256:f5e86ebb9deee5e373ce1541f927c6d2a0f0b9244481aa9e63b5fe668e66b482

Observation 1024990e-3795-47dc-8873-a393f8cc9b43 · outbound

This paper cites Towards understanding knowledge distillation.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Towards understanding knowledge distillation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:25:29.635210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T12:25:28.453398Z digest=sha256:a04408ea9d2c8a80c9b02a95c84ccbaa48e2e9c5f9760811c7c27b16339246b3

Observation e5f48848-06da-4b1d-9e02-33e09e48f46b · outbound

This paper cites Knowledge distillation performs partial variance reduction.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Knowledge distillation performs partial variance reduction

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:25:29.619969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T12:25:28.458085Z digest=sha256:551cd97232310b404c92a381e80e8a638aefd5d9b8ed0f404c8fd8516e2657c6

Observation 52a39e37-f54e-4ff0-8073-df32c9f50524 · outbound

This paper cites Analysing Mathematical Reasoning Abilities of Neural Models.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Analysing Mathematical Reasoning Abilities of Neural Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:28.462546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:28.462546Z digest=sha256:b8cb2c3262590200aa76a991825114078c02533bbf1d32bcca009d9ed3b52e8f

Observation f223f575-148c-4056-abc9-c696b265925b · outbound

This paper cites BOND: Aligning LLMs with Best-of-N Distillation.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling BOND: Aligning LLMs with Best-of-N Distillation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:28.468096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:28.468096Z digest=sha256:7ff747f7654188dbf031e7f398342c256b1514279e4d9c7eea208770b686f6b2

Observation 794df7d2-8c04-4523-afba-0c37c3b052a7 · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:28.473347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:28.473347Z digest=sha256:16c301a8e84b5fa4fa5738361872ccf4bc010cd9e0334965ae8d0336b4d3ca72

Observation cadc9b2e-20f5-4ea4-938f-aeb4cdea2e7d · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:28.477571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:28.477571Z digest=sha256:0d5d775149decd780f3544541c0049a9cdf0a7ad7af5a5366ff68de138a39339

Observation b0e3b286-442b-4737-9c7e-39e4b5a9fd4c · outbound

This paper cites Looking beyond the next token.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Looking beyond the next token

Reference 62

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T12:25:28.757794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T12:25:28.481689Z digest=sha256:f07029c0b217847c09c42f3bab398e7d291b9ca064955e52813827c7282dbb6e

Observation 8ff7a89c-ae5c-49ee-849d-0e5cdf7798b2 · outbound

This paper cites Qwen3 Technical Report.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Qwen3 Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:28.485825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:28.485825Z digest=sha256:931a7e0dd9b65a90cecef04c772d93f3dee33abb143f3fb4796c81e18c73849c

Observation 8db1bde8-d7fc-41c4-8ebd-99f2ca04cbfc · outbound

This paper cites Naturalreasoning: Reasoning in the wild with 2.8m challenging questions, 2025.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Naturalreasoning: Reasoning in the wild with 2.8m challenging questions, 2025

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:28.490690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:28.490690Z digest=sha256:56f7b8e44828d4711683518363cfec7962630d4c2c52fb9b8de2f3ca6609cfbb

Observation 2f7f2f36-07a8-4106-bf54-b3fae88e91f3 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:28.494576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:28.494576Z digest=sha256:49082c9a1a7730de96b021a6f37e7125bef7c56bbd094b1fb04916602fb943b5

Observation 45fc63d8-b021-47dd-ad0c-245369d337a8 · outbound

This paper cites Lifting the Curse of Capacity Gap in Distilling Language Models.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Lifting the Curse of Capacity Gap in Distilling Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:28.499159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:28.499159Z digest=sha256:7af5a4980ad492f623ddff1be08aca5f99b8b4a0dc06f18bf915471b9144c4b0

Observation 520022cd-cc73-4caa-9412-06279f1c3ca1 · outbound

This paper cites Towards the Law of Capacity Gap in Distilling Language Models.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Towards the Law of Capacity Gap in Distilling Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:28.503732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:28.503732Z digest=sha256:dc2671416b908a140a94a7c20e6e1edfefa0112347399b7a6f312efd2f4d4d36

Observation 57f74805-85f5-41e1-a581-74370717f587 · outbound

This paper cites Forcing Diffuse Distributions out of Language Models.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Forcing Diffuse Distributions out of Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T12:25:28.508547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:25:28.508547Z digest=sha256:c09d906d8ea25b5aae263080213ae782794b178891dfa110227a6a153c81b3d7

Pith citing papers

Observation 04476116-ecc1-4122-944c-2d0c395bdfa6 · inbound

Ministral 3 cites this paper.

Ministral 3 Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:12:24.696963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T19:12:24.627033Z digest=sha256:ca0365280ccccc4dbac43e260d57cf94db150511f99c8af663145aeeedff772a

Observation bc6c8d4f-1d84-421f-85d3-a20dac254e15 · inbound

Memory Grafting: Scaling Language Model Pre-training via Offline Conditional Memory cites this paper.

Memory Grafting: Scaling Language Model Pre-training via Offline Conditional Memory Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:49:40.781393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T05:49:14.789955Z digest=sha256:d2515ab38e11d2369efa4dfc26e7b818f52059271bad99b9b95f0ec1e6b6afaa

Observation fb3df0d2-ba17-40a7-98f1-6e64c0bf2d94 · inbound

Bridging Compute- and Data-Optimal Pretraining cites this paper.

Bridging Compute- and Data-Optimal Pretraining Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-01T03:02:07.425724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:02:07.425724Z digest=sha256:348e39e02119c9ebd44ed2e791b47508a8d81ee03f00f8c3c005562e88c62ed0