Pith. sign in

Paper Citation Record · LEDGER

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection

As of 14 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 7 inbound Pith citation observations for arXiv:2502.06042.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06042 v2

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:01:36.218426Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T22:30:58.664303Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact2
  • verified fuzzy13
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 0e539dbd-18d5-4e79-8178-391069d27b79 · outbound

This paper cites Scaling laws for generative mixed-modal language models.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Scaling laws for generative mixed-modal language models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.874266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.045067Z digest=sha256:4e05a6e724f779c829b538946ecbab31e2bd85d6d661bbe71396a4bcc19fb30e

Observation 54015f12-73e8-4298-bc75-689f38c3d2a1 · outbound

This paper cites Physics in Next-token Prediction.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Physics in Next-token Prediction

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-08T17:01:36.675840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.049342Z digest=sha256:16018d432307d3de79d14e72e5a2a92594e818cffdf94264c3a449cb7ebca80c

Observation 348d53aa-d60c-4d10-bf07-05829460d3d8 · outbound

This paper cites An Empirical Study of Scaling Laws for Transfer.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection An Empirical Study of Scaling Laws for Transfer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.054331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.054331Z digest=sha256:caf1cbe432fc232aa7afb9e99f204882c870bd91e2336dd4a788f28c3bd7fb79

Observation 5ad549d4-7fac-4b5a-90c7-53a13cc61fbb · outbound

This paper cites Chinchilla Scaling: A replication attempt.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Chinchilla Scaling: A replication attempt

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.058241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.058241Z digest=sha256:210b166af41caf8fb9f6695eee1005b359a93c8ee5aa8cbe328c70d036408fdf

Observation d3d0fa82-5ef3-490b-8ea0-13d01b547b18 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.062327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.062327Z digest=sha256:af30a1d1109da89e876a8e743e0e46a582513a23088525e4d462ab61e03e8193

Observation f5824e7e-4bcb-4e98-be0e-ebfb8295506d · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.066170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.066170Z digest=sha256:d9961b2658a088f54aafd409c2448a6f2555e2bf49ef9e85dedbbdc33a27b23c

Observation 6f0003f1-0c81-4ad5-8ba1-00c8381cdae5 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.070822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.070822Z digest=sha256:a3144b269f1ba11f130c39ec9b3efd44e64823dd0f29533d8bfc6a6c9ebd1e2e

Observation 12aa82ea-4964-4565-8705-4c5e29328f9b · outbound

This paper cites B., Werra, L.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection B., Werra, L

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.864014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.074685Z digest=sha256:749f0b03c60d78e895eca4824a2caa83a1beb6aec2e34fb641a88644ae1b43ba

Observation dd78c85f-f506-471a-ab91-a72768059e65 · outbound

This paper cites The elements of statistical learning: data mining, inference, and prediction, 2017.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection The elements of statistical learning: data mining, inference, and prediction, 2017

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.854328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.078314Z digest=sha256:8991cb33ba2c6e3983183385311a9269ea774bbafcfde251e8014991e83a797b

Observation 2bf46535-cc36-49d7-abf5-1af770924f88 · outbound

This paper cites Towards a Unified View of Parameter-Efficient Transfer Learning.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Towards a Unified View of Parameter-Efficient Transfer Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.082196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.082196Z digest=sha256:171d1103dfd72b9fc0edae3d75a151493cd7e28667facee1c7bca2c27656af5b

Observation df1de12e-323e-4169-a5d7-b9b9546a62ed · outbound

This paper cites Measuring massive multitask language understanding.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Measuring massive multitask language understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.086743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.086743Z digest=sha256:38894b19cc378452a3aa34a6d0cc9d4f4b1a0ca551e5b10841f1f2572a04d3c7

Observation 722dd3be-544e-49c2-891c-b678f9a43b81 · outbound

This paper cites Scaling Laws for Transfer.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Scaling Laws for Transfer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.090967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.090967Z digest=sha256:cef6b1619d23d2cb416282f15e745fc92da64180bbfad4c002a73180ce1fd0e3

Observation 19d8cfee-df61-4bc6-8e13-6b387f11ea40 · outbound

This paper cites Deep Learning Scaling is Predictable, Empirically.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Deep Learning Scaling is Predictable, Empirically

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.095414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.095414Z digest=sha256:7b6a63d69ae3a85ec661c40be5dd7cf96e712d4e6f4f99966bc82c8731057a0d

Observation 670b62d3-df1b-43f4-ac17-63c093676837 · outbound

This paper cites Disentangling and Mitigating the Impact of Task Similarity for Continual Learning.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Disentangling and Mitigating the Impact of Task Similarity for Continual Learning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-08T17:01:36.579501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.099665Z digest=sha256:85caacd16c2191a6839fa851179076b203dfff720168bf7ab6af1fb27333d534

Observation d3629529-3168-4d20-bb8d-ce8902bde973 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Training Compute-Optimal Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.103707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.103707Z digest=sha256:baba003a663d1ae6d235903f5b7a1a9f946a020aa7e67c8647e774a5883eaf07

Observation 7ffca25e-5a4d-4446-8e11-50e9128774da · outbound

This paper cites an unresolved cited work.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-08T17:01:36.836051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.107743Z digest=sha256:4a66878b5ca196e27a3e1913129ee66cfd8896a9ca8f9c5999660267698b4344

Observation 13153a14-8642-4f09-a584-d0d7da39aedc · outbound

This paper cites Parameter-Efficient Transfer Learning for NLP.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Parameter-Efficient Transfer Learning for NLP

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.110759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.110759Z digest=sha256:982f5fe61590c797162daf9b4c45fce0ddd7478cb5f51795af6128f8b04960c8

Observation 2f570756-ec8c-491d-9c7f-683279943162 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection LoRA: Low-Rank Adaptation of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.113930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.113930Z digest=sha256:beda59c740000493f79b5bb741a7291b62b1239190a6c07696133d3ed657b4d7

Observation 5e9eebac-00a9-4b75-9358-e17ecb98a7fd · outbound

This paper cites L., Wang, C., Yao, Y., Zhao, C., Zhou, J., Cai, J., Zhai, Z., Ding, N., Jia, C., Zeng, G., dahai li, Liu, Z., and Sun, M.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection L., Wang, C., Yao, Y., Zhao, C., Zhou, J., Cai, J., Zhai, Z., Ding, N., Jia, C., Zeng, G., dahai li, Liu, Z., and Sun, M

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.825506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.117451Z digest=sha256:9c887a03387466877ccfcec5e635843685440540d6612b6849cdbe8e839135db

Observation 32c3eda3-14ab-4ce6-820c-2666f8d054b3 · outbound

This paper cites Simple and Scalable Strategies to Continually Pre-train Large Language Models.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Simple and Scalable Strategies to Continually Pre-train Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.120434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.120434Z digest=sha256:c9db5a20a17db2dc03185af948afc66753443b9253bc563a6c702241a2950b33

Observation c1c71820-692a-4f1a-a23e-6be4f017bc7c · outbound

This paper cites Scaling laws for downstream task performance of large language models.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Scaling laws for downstream task performance of large language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.123633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.123633Z digest=sha256:73ff5090bfe95bf253f9a98e5a4ab3c2aaf4a317cdb3d47e98c52bc59cddf520

Observation 46cd8a44-6f4f-4b39-bed4-e6aae40e8a25 · outbound

This paper cites Scaling Laws for Forgetting When Fine-Tuning Large Language Models.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Scaling Laws for Forgetting When Fine-Tuning Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.126796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.126796Z digest=sha256:f649793d96a67b41a23f1e510458f817b23e8daa0071e05916db1d78d21a9fcb

Observation eb9e3c6f-cab7-4d45-be90-58871f1ca401 · outbound

This paper cites Get more for less: Principled Data Selection for Warming Up Fine-Tuning in LLMs.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Get more for less: Principled Data Selection for Warming Up Fine-Tuning in LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.130140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.130140Z digest=sha256:cf7dbe53c0c60f247e0b52a509584bd71cd12692a8ae85e2dcf0e0828f8a852b

Observation f69888e6-bb5c-4108-8578-105fac66082d · outbound

This paper cites Scaling Laws for Neural Language Models.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Scaling Laws for Neural Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.133857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.133857Z digest=sha256:3288b7ced3ede1a213641269147dbf2c7195b4d79d1e27db09cb3c9acd8aab4c

Observation d4b50a27-f704-4b8b-b279-6409a8d22899 · outbound

This paper cites SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.137146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.137146Z digest=sha256:ea97733712fd7f63b32a06456c10f537c5c53dabd4af812972b4a1e5c8b55522

Observation 9e551a4f-ba48-4fff-85fc-cd1a357abb79 · outbound

This paper cites Improved fine-tuning by better leveraging pre-training data.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Improved fine-tuning by better leveraging pre-training data

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.814908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.140667Z digest=sha256:93af5ef0bce7b38a22e97ad5282990ad7b3451db24306c1bdf913efaaa25daf0

Observation 5bdfefc6-a351-409c-ac89-22579ccc7eac · outbound

This paper cites and Hutter, F.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection and Hutter, F

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.803295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.144198Z digest=sha256:5bbfc6d130a0e17c36ba1b255f62dcc5e7c4a58485ceb2f3398791a9e8d47ed6

Observation 9fa8498c-dad9-43e8-a60c-a0e333f548b3 · outbound

This paper cites and Hutter, F.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection and Hutter, F

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.147470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.147470Z digest=sha256:a6ddebd6bf41b35f7baaa44d410868c3cc5a5c32c909f9f0926faf04cac98ed2

Observation b0701b48-8b4b-4e01-86cc-847043c43dbd · outbound

This paper cites An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.150551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.150551Z digest=sha256:394a9c9e4e4da9436503ec099ee882c77722ba82d58c86ae327b437acbe7dbec

Observation 89f4ec0a-15e4-410c-9400-07e39d7e2ef8 · outbound

This paper cites LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.154116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.154116Z digest=sha256:437b8681c29ee66790f3b299d515a291560fc756b5acd3bc323de0ba9d1d928a

Observation 22ad3363-1613-43d8-a6f3-8698700c4d9a · outbound

This paper cites Metaicl: Learning to learn in context.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Metaicl: Learning to learn in context

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.786124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.157349Z digest=sha256:640b35c38e3922676bde7bec1792c0b21f57491b1246fe92dcbef270a7c7ced7

Observation b42c65b6-0a14-4b5f-a976-f198e5fb034a · outbound

This paper cites an unresolved cited work.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.160335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.160335Z digest=sha256:3639f980c23aa135288734005f37a91b3ffcf162a8c771d0bf6be76f425e259f

Observation 275cdcbb-c633-49dd-8b0e-b6b879a9b018 · outbound

This paper cites Training language models to follow instructions with human feedback.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Training language models to follow instructions with human feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.163627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.163627Z digest=sha256:3af0296d28f9b21fa53032c941262d48b47c08e3aa8bb31ef02dfde907b82144

Observation 774c7d91-55e4-49be-8a31-36fbb86bdd82 · outbound

This paper cites Resolving Discrepancies in Compute-Optimal Scaling of Language Models.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Resolving Discrepancies in Compute-Optimal Scaling of Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.166405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.166405Z digest=sha256:54aadfac2d0fb184fc488bfe1732f0544dfa8be19e660d9b6455059b999b428b

Observation d1acaf38-48d1-4533-b92d-d1ccf2ea4737 · outbound

This paper cites Self-attention Does Not Need $O(n^2)$ Memory.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Self-attention Does Not Need $O(n^2)$ Memory

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.169940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.169940Z digest=sha256:2479885f1f17c34b563fc5a4241ae36e7375965cb95243693262386b7e74d924

Observation d3f9934a-f0b0-4695-ad68-967cbd97e01a · outbound

This paper cites Language models are unsupervised multitask learners.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Language models are unsupervised multitask learners

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.760912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.174011Z digest=sha256:2f31ed47e4082a3aa971a14980bb3346a8915a4e3975bd5fabcb29f25f0a38ff

Observation 9e25533c-d209-44cc-a08c-28141256f423 · outbound

This paper cites Multitask prompted training enables zero-shot task generalization.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Multitask prompted training enables zero-shot task generalization

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.749186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.177949Z digest=sha256:b6fba17a9e683b93f63ce13e31911fbe5e9ae24c785212e3677f602b771e88a0

Observation 642d8bb8-aa55-41cd-a887-240b8c4f5928 · outbound

This paper cites Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.181915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.181915Z digest=sha256:219c97d0d5fcf7c49784663cf67232a2e7f922077686f76b484f28d8b034ca68

Observation 115c1904-e99d-429f-903e-b3d1bc5afdfe · outbound

This paper cites Sequence to Sequence Learning with Neural Networks.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Sequence to Sequence Learning with Neural Networks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.186433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.186433Z digest=sha256:69c8c0202a03256af4fc7e5948d9a48fc6881aabb5a5456946f919307be9a45a

Observation 3b9816fa-688c-4341-b4a2-73273cece9e3 · outbound

This paper cites Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants, 2023.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants, 2023

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.738109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.190495Z digest=sha256:d31fbe3db35149dab26872f1102ee2393269c2e65cd9d17015713f7011aa738d

Observation db0581da-92ba-4af8-90c0-2553e2f59445 · outbound

This paper cites Scaling Law with Learning Rate Annealing.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Scaling Law with Learning Rate Annealing

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.194110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.194110Z digest=sha256:1038536ffddccf1e34ec06c8efd7d8277e056404c2c02337c293300cd87c9ef5

Observation 702def6f-d4b2-4b6f-9f4f-80b9658d61cf · outbound

This paper cites When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.198148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.198148Z digest=sha256:9c7ec47b02e16edea949f4648ed2f0aca1807e2dee2bf77772242000359b3d90

Observation e4c8af24-347c-4f88-8d86-0e579484ba57 · outbound

This paper cites an unresolved cited work.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-08T17:01:36.727455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.202404Z digest=sha256:946695078f599127c07d813450fceab9a70eadb4db8e06b8d6f4d1454d73de15

Observation 9c8321d0-e366-4e90-b046-338df60caa46 · outbound

This paper cites W., Lester, B., Du, N., Dai, A.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection W., Lester, B., Du, N., Dai, A

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.717023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.205762Z digest=sha256:6e6ffec55bf01bea7ce018c816dcc8f8cb3cb9109ba2830147c610b76c77c5bf

Observation ec7e55ca-b143-495b-bae8-4fbbdeee3771 · outbound

This paper cites What makes a high-quality training dataset for large language models: A practitioners' perspective.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection What makes a high-quality training dataset for large language models: A practitioners' perspective

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.705655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.208846Z digest=sha256:7cbd7c81db45176607b77af5c59c739117e1e8cfda4e49555a4256e07e7c2bdf

Observation 8e58dd0b-5e36-44bd-a36e-76d72e08d32c · outbound

This paper cites When scaling meets LLM finetuning: The effect of data, model and finetuning method.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection When scaling meets LLM finetuning: The effect of data, model and finetuning method

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.693831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.211919Z digest=sha256:15406f752f3696b0a8fd9ac98c3405c8639d18feecd12cd1ab1fd74fb19e5d2c

Observation da611d14-44c9-419b-bc38-a58e6281c029 · outbound

This paper cites Asymmetry in Low-Rank Adapters of Foundation Models.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Asymmetry in Low-Rank Adapters of Foundation Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.214939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.214939Z digest=sha256:82e1e4adb36f8b27c994b76a6971ea256af9e00c8d1e899273fa7a1d7f3678c0

Observation 170695dc-26c0-4cb5-af04-f682e50a23ea · outbound

This paper cites write newline.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection write newline

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.218426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.218426Z digest=sha256:5566d5568d99b011aef1d04c1d321216c932076f74b1ab050688f01987ec56fa

Pith citing papers

Observation 6a545c2e-3c8f-4a92-9ebf-6f75cc838343 · inbound

Diversity in Large Language Models under Supervised Fine-Tuning cites this paper.

Diversity in Large Language Models under Supervised Fine-Tuning Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-09T20:37:32.124103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-09T20:32:37.788283Z digest=sha256:d96bd750a1b4bc4fd45a6a6bcad2f9b3896c3a79e5cd4fd902adcc562b64d5dc

Observation 978121c7-1d7d-4fd1-beeb-272e51c1da1b · inbound

Diversity in Large Language Models under Supervised Fine-Tuning cites this paper.

Diversity in Large Language Models under Supervised Fine-Tuning Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:11:18.148664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-12T03:10:22.314719Z digest=sha256:f2496091fa653109740c13e1b6a73f962ad8e7751bb0b8de076a956d6ca469d2

Observation 85bd389e-4357-4b8c-aa69-868cf41f28f9 · inbound

Early Data Exposure Improves Robustness to Subsequent Fine-Tuning cites this paper.

Early Data Exposure Improves Robustness to Subsequent Fine-Tuning Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:47:58.617673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-14T20:45:27.673290Z digest=sha256:ad8ef56ab271d9fee06a939cbaeb08eb77260c22ca9405e098e983a5dbc06804

Observation 960bddf1-eedd-4e95-98d1-0465059c9ed9 · inbound

Scaling Laws for Mixture Pretraining Under Data Constraints cites this paper.

Scaling Laws for Mixture Pretraining Under Data Constraints Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:48:00.978072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-14T21:44:31.429223Z digest=sha256:5db5351744dfd1711f18d8b3d2c02ff184e38e8a535ed2d01444175b5eec697f

Observation 0701f0ff-3b37-47eb-9a35-895ffbbcad32 · inbound

Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay cites this paper.

Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T23:24:01.970658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-29T23:15:45.174086Z digest=sha256:197b766424ae1378dedfea8cc4bea899188cd0a0b166633817b645ebb0d506a7

Observation 0e5085ee-2d3f-4315-a358-1eabbe3127eb · inbound

Knowledge Editing in Masked Diffusion Language Models cites this paper.

Knowledge Editing in Masked Diffusion Language Models Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:36:29.935755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-28T09:51:42.811856Z digest=sha256:d19c08148cfff287674b6272964eb8dbe608b21650ca76731793deda9d54084e

Observation 85a28b91-5f85-4d9c-af51-c4cff710587a · inbound

A Predict-then-Correct Loop Based on Few-Shot Continuous Contextual Bandit for Demand Forecasting cites this paper.

A Predict-then-Correct Loop Based on Few-Shot Continuous Contextual Bandit for Demand Forecasting Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T22:30:58.664303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:30:58.664303Z digest=sha256:f234691bba756907343fa4a42691d6b6166df476417e88952828fe12e21dad8e