Pith. sign in

Paper Citation Record · LEDGER

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection

As of 15 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 7 inbound Pith citation observations for arXiv:2502.06042.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06042 v2

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:01:36.218426Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T22:30:58.664303Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact2
  • verified fuzzy13
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 0e539dbd-18d5-4e79-8178-391069d27b79 · outbound

This paper cites Scaling laws for generative mixed-modal language models.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Scaling laws for generative mixed-modal language models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.874266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.045067Z digest=sha256:153ec070edd0c4d2f8f44fc4673a2aded1cc60a9b779909150f8418c84257a51

Observation 54015f12-73e8-4298-bc75-689f38c3d2a1 · outbound

This paper cites Physics in Next-token Prediction.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Physics in Next-token Prediction

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-08T17:01:36.675840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.049342Z digest=sha256:9229afdce0bf6058ca79d0b812ea7bdca2453cee00cfb7378dc0d37d0f6aa7fa

Observation 348d53aa-d60c-4d10-bf07-05829460d3d8 · outbound

This paper cites An Empirical Study of Scaling Laws for Transfer.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection An Empirical Study of Scaling Laws for Transfer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.054331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.054331Z digest=sha256:b990854c74bc9cf119c48429398dd780d02510cdf23abc7dbb722fb37fda746c

Observation 5ad549d4-7fac-4b5a-90c7-53a13cc61fbb · outbound

This paper cites Chinchilla Scaling: A replication attempt.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Chinchilla Scaling: A replication attempt

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.058241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.058241Z digest=sha256:1ab41b6531faae1eb6807d39e540ef54047e4a651bf49b9730c6582076e8ef75

Observation d3d0fa82-5ef3-490b-8ea0-13d01b547b18 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.062327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.062327Z digest=sha256:265e5352acf0ff27c0d3f962ace8529d45368c89d8cd094baa73f4fabab8d3fe

Observation f5824e7e-4bcb-4e98-be0e-ebfb8295506d · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.066170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.066170Z digest=sha256:27fed7e99e80a9694fcde41cb11193af896672fb1a5394babcf1708d7e9a23a0

Observation 6f0003f1-0c81-4ad5-8ba1-00c8381cdae5 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.070822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.070822Z digest=sha256:b298b1104ed74e39108b57bfbea7bbc4a96503e1b60034e0a8e81485eddb0015

Observation 12aa82ea-4964-4565-8705-4c5e29328f9b · outbound

This paper cites B., Werra, L.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection B., Werra, L

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.864014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.074685Z digest=sha256:dd2dfc9cf8d940c97ce464420b0975045ce312b323a8e5ca0f443444d1fdadb0

Observation dd78c85f-f506-471a-ab91-a72768059e65 · outbound

This paper cites The elements of statistical learning: data mining, inference, and prediction, 2017.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection The elements of statistical learning: data mining, inference, and prediction, 2017

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.854328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.078314Z digest=sha256:3d4ee1bf3cd4df25cda491178cca482299487dbb5a747b696f0af381a89747bc

Observation 2bf46535-cc36-49d7-abf5-1af770924f88 · outbound

This paper cites Towards a Unified View of Parameter-Efficient Transfer Learning.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Towards a Unified View of Parameter-Efficient Transfer Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.082196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.082196Z digest=sha256:e34ee05a887b532896f24dee5fb9165623f5e9c676924859f605e6be7aa14d42

Observation df1de12e-323e-4169-a5d7-b9b9546a62ed · outbound

This paper cites Measuring massive multitask language understanding.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Measuring massive multitask language understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.086743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.086743Z digest=sha256:e56149d2df8f8ff179ce9d5ce9d88b556683f70685ff1b51041db3959e35f43d

Observation 722dd3be-544e-49c2-891c-b678f9a43b81 · outbound

This paper cites Scaling Laws for Transfer.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Scaling Laws for Transfer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.090967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.090967Z digest=sha256:57313ff0fe1ea0ce6eaabf65c92d6a5703077f9f8859e0a5b5d385473bf288d1

Observation 19d8cfee-df61-4bc6-8e13-6b387f11ea40 · outbound

This paper cites Deep Learning Scaling is Predictable, Empirically.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Deep Learning Scaling is Predictable, Empirically

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.095414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.095414Z digest=sha256:55820bb66535214cd4c64519a1bc974f061900e1583d090597b0213586d68c7a

Observation 670b62d3-df1b-43f4-ac17-63c093676837 · outbound

This paper cites Disentangling and Mitigating the Impact of Task Similarity for Continual Learning.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Disentangling and Mitigating the Impact of Task Similarity for Continual Learning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-08T17:01:36.579501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.099665Z digest=sha256:d8625c0a993309165297a6f99432ee5b93da8bd1f0fd7fe9b784ead584e77fde

Observation d3629529-3168-4d20-bb8d-ce8902bde973 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Training Compute-Optimal Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.103707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.103707Z digest=sha256:41be13790f3910355e63806b2ab0c9f01015a2e083ab89bc1ac9223a52be1b19

Observation 7ffca25e-5a4d-4446-8e11-50e9128774da · outbound

This paper cites an unresolved cited work.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-08T17:01:36.836051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.107743Z digest=sha256:d0f71fe67fe67576f7b518b5ce0fa8728f973d56b14de5062c191042f08c33a5

Observation 13153a14-8642-4f09-a584-d0d7da39aedc · outbound

This paper cites Parameter-Efficient Transfer Learning for NLP.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Parameter-Efficient Transfer Learning for NLP

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.110759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.110759Z digest=sha256:4b0c3bca76630bb9dae95501468cc4e878f33ae91da4de28079e1fa6acd1659c

Observation 2f570756-ec8c-491d-9c7f-683279943162 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection LoRA: Low-Rank Adaptation of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.113930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.113930Z digest=sha256:442d11101a90681144e3d1d31b09703fbdc73a086aa53a34b58b16fdb9a304ac

Observation 5e9eebac-00a9-4b75-9358-e17ecb98a7fd · outbound

This paper cites L., Wang, C., Yao, Y., Zhao, C., Zhou, J., Cai, J., Zhai, Z., Ding, N., Jia, C., Zeng, G., dahai li, Liu, Z., and Sun, M.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection L., Wang, C., Yao, Y., Zhao, C., Zhou, J., Cai, J., Zhai, Z., Ding, N., Jia, C., Zeng, G., dahai li, Liu, Z., and Sun, M

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.825506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.117451Z digest=sha256:e4c2d5da4a9212cc5d864fd1011be9e67e1b90206c84e1f7f60e9848602fefcf

Observation 32c3eda3-14ab-4ce6-820c-2666f8d054b3 · outbound

This paper cites Simple and Scalable Strategies to Continually Pre-train Large Language Models.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Simple and Scalable Strategies to Continually Pre-train Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.120434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.120434Z digest=sha256:7a077c3d1fe9371ff60df15ee844519f50bdaa15a549e59e541e34701973ffc7

Observation c1c71820-692a-4f1a-a23e-6be4f017bc7c · outbound

This paper cites Scaling laws for downstream task performance of large language models.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Scaling laws for downstream task performance of large language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.123633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.123633Z digest=sha256:14b052f7675402f380cc325f585849ca603fdb18b5b30bc4e2d8b12010344ee6

Observation 46cd8a44-6f4f-4b39-bed4-e6aae40e8a25 · outbound

This paper cites Scaling Laws for Forgetting When Fine-Tuning Large Language Models.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Scaling Laws for Forgetting When Fine-Tuning Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.126796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.126796Z digest=sha256:0fe6bbc976e3067358ee8df364d141de4e6537ddbf0d61697831a44ee60402af

Observation eb9e3c6f-cab7-4d45-be90-58871f1ca401 · outbound

This paper cites Get more for less: Principled Data Selection for Warming Up Fine-Tuning in LLMs.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Get more for less: Principled Data Selection for Warming Up Fine-Tuning in LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.130140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.130140Z digest=sha256:a6df85e1ec684c57f3930faedaafdbcf5a29eddff3589af52ec346dd588bbcbd

Observation f69888e6-bb5c-4108-8578-105fac66082d · outbound

This paper cites Scaling Laws for Neural Language Models.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Scaling Laws for Neural Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.133857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.133857Z digest=sha256:f34a094531f2230f6d2e0aa4bca71d845c43d13a68700e96f5f90ea919451dd3

Observation d4b50a27-f704-4b8b-b279-6409a8d22899 · outbound

This paper cites SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.137146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.137146Z digest=sha256:4b8ede2a43d66a1550295db86a70bbca74bfd6ee681b3ad2bcd0ae89ba00f276

Observation 9e551a4f-ba48-4fff-85fc-cd1a357abb79 · outbound

This paper cites Improved fine-tuning by better leveraging pre-training data.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Improved fine-tuning by better leveraging pre-training data

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.814908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.140667Z digest=sha256:b854a20863c6d53202f474c4c63b052fa243612e10acd9dfc8df1faee2673bc7

Observation 5bdfefc6-a351-409c-ac89-22579ccc7eac · outbound

This paper cites and Hutter, F.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection and Hutter, F

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.803295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.144198Z digest=sha256:8ea34b5a2e726d65bf1a184d03b060b2bbd10b15ba0d4376edcc17d2cf4a70a3

Observation 9fa8498c-dad9-43e8-a60c-a0e333f548b3 · outbound

This paper cites and Hutter, F.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection and Hutter, F

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.147470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.147470Z digest=sha256:9f6d0a9dbc1781f6a61675963a4839dc222ca7376ba610d069fa45530c264a2a

Observation b0701b48-8b4b-4e01-86cc-847043c43dbd · outbound

This paper cites An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.150551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.150551Z digest=sha256:7acbc5deff80cb27e1a2ba004de201afdde2aaf16b23de78c21c6eac494f5991

Observation 89f4ec0a-15e4-410c-9400-07e39d7e2ef8 · outbound

This paper cites LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.154116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.154116Z digest=sha256:319a766f9b31dc95ee4693b20e7164ac09108481818487cbaff78b5f074f5f98

Observation 22ad3363-1613-43d8-a6f3-8698700c4d9a · outbound

This paper cites Metaicl: Learning to learn in context.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Metaicl: Learning to learn in context

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.786124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.157349Z digest=sha256:207e2c8004a98ab6c9a9b13e2f023794d50bf5a009948a6018c5cb631edcc536

Observation b42c65b6-0a14-4b5f-a976-f198e5fb034a · outbound

This paper cites an unresolved cited work.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.160335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.160335Z digest=sha256:321419de5bce47219b5d05e112ba05c9a0893be1fe59da02cfd7170c047889c0

Observation 275cdcbb-c633-49dd-8b0e-b6b879a9b018 · outbound

This paper cites Training language models to follow instructions with human feedback.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Training language models to follow instructions with human feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.163627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.163627Z digest=sha256:876681364b05dab05acc3a3a7839548013a1be5c76b4bfa22bc7b69fe1a280f0

Observation 774c7d91-55e4-49be-8a31-36fbb86bdd82 · outbound

This paper cites Resolving Discrepancies in Compute-Optimal Scaling of Language Models.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Resolving Discrepancies in Compute-Optimal Scaling of Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.166405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.166405Z digest=sha256:ed497da8d43b75ca162c4df9bb454c7dc16cfc1534a6b74b46cbed61b20ef8ce

Observation d1acaf38-48d1-4533-b92d-d1ccf2ea4737 · outbound

This paper cites Self-attention Does Not Need $O(n^2)$ Memory.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Self-attention Does Not Need $O(n^2)$ Memory

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.169940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.169940Z digest=sha256:5437ea6173e60f0dc2e5f47c65d2c4fdc3ea55d71e22079f8747df7bdeb761cb

Observation d3f9934a-f0b0-4695-ad68-967cbd97e01a · outbound

This paper cites Language models are unsupervised multitask learners.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Language models are unsupervised multitask learners

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.760912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.174011Z digest=sha256:f2b7feab1905a2c8e61d6d62f18b08f7661d5de1e1c7f825a452868ac5465587

Observation 9e25533c-d209-44cc-a08c-28141256f423 · outbound

This paper cites Multitask prompted training enables zero-shot task generalization.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Multitask prompted training enables zero-shot task generalization

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.749186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.177949Z digest=sha256:fcea7ef47f1b56acb8aa6c264af1a168f6ba1806334b1ca89668cccc23a38e2f

Observation 642d8bb8-aa55-41cd-a887-240b8c4f5928 · outbound

This paper cites Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.181915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.181915Z digest=sha256:f8dc983032660ee7e7e9a3d192af404616019529c6d80e7b3f4bd3617fd86762

Observation 115c1904-e99d-429f-903e-b3d1bc5afdfe · outbound

This paper cites Sequence to Sequence Learning with Neural Networks.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Sequence to Sequence Learning with Neural Networks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.186433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.186433Z digest=sha256:8ec8e55ffca12359098e27e7ae588058d40d59d19837aba8ab034a7fd4ea49c8

Observation 3b9816fa-688c-4341-b4a2-73273cece9e3 · outbound

This paper cites Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants, 2023.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants, 2023

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.738109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.190495Z digest=sha256:ea14e5e661e9b5f12928501ecb93ffd0112aba37bc345b665b2ef04de7c0e2dd

Observation db0581da-92ba-4af8-90c0-2553e2f59445 · outbound

This paper cites Scaling Law with Learning Rate Annealing.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Scaling Law with Learning Rate Annealing

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.194110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.194110Z digest=sha256:82f18b9a863f97799e73d7d7fc71d9272f2aa3e44fe35fff38ceb73b016094bf

Observation 702def6f-d4b2-4b6f-9f4f-80b9658d61cf · outbound

This paper cites When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.198148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.198148Z digest=sha256:b32c8e4320797287a8b9f16da20c9753c9222dfd4fe911848f7082deead515ee

Observation e4c8af24-347c-4f88-8d86-0e579484ba57 · outbound

This paper cites an unresolved cited work.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-08T17:01:36.727455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.202404Z digest=sha256:57cf18dcbd8a3f34cece26ac122dfd53702bdaf3f416765d247860b3fa5307be

Observation 9c8321d0-e366-4e90-b046-338df60caa46 · outbound

This paper cites W., Lester, B., Du, N., Dai, A.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection W., Lester, B., Du, N., Dai, A

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.717023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.205762Z digest=sha256:2e3eb9e5c48567f998dc8a523f915faed5c6a3e0f7ac4ce04de42673155be10c

Observation ec7e55ca-b143-495b-bae8-4fbbdeee3771 · outbound

This paper cites What makes a high-quality training dataset for large language models: A practitioners' perspective.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection What makes a high-quality training dataset for large language models: A practitioners' perspective

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.705655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.208846Z digest=sha256:cace97f52fb75f3516322750531974ef5417dc3eec4fc67c9848fe727f6bebc9

Observation 8e58dd0b-5e36-44bd-a36e-76d72e08d32c · outbound

This paper cites When scaling meets LLM finetuning: The effect of data, model and finetuning method.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection When scaling meets LLM finetuning: The effect of data, model and finetuning method

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:01:36.693831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-08T17:01:36.211919Z digest=sha256:ea0284b4f3fa3388e4ef32245e5cc7398235cb13d4ddae8e074fbf1daba15db7

Observation da611d14-44c9-419b-bc38-a58e6281c029 · outbound

This paper cites Asymmetry in Low-Rank Adapters of Foundation Models.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Asymmetry in Low-Rank Adapters of Foundation Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.214939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.214939Z digest=sha256:cd48d87199a7ccf7cfc564a8caecaadd980ffe6edfc13f4c281df3f1e83766ed

Observation 170695dc-26c0-4cb5-af04-f682e50a23ea · outbound

This paper cites write newline.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection write newline

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.218426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.218426Z digest=sha256:57b9a36d8945f18a0580b874307957ca720df5ed9962fdc01022a79c8a4f2f3f

Pith citing papers

Observation 6a545c2e-3c8f-4a92-9ebf-6f75cc838343 · inbound

Diversity in Large Language Models under Supervised Fine-Tuning cites this paper.

Diversity in Large Language Models under Supervised Fine-Tuning Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-09T20:37:32.124103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-09T20:32:37.788283Z digest=sha256:147b1b0a308ea7f6fc00e049e4c14c1c68a0779fe5fb9e38a16844de704abdd2

Observation 978121c7-1d7d-4fd1-beeb-272e51c1da1b · inbound

Diversity in Large Language Models under Supervised Fine-Tuning cites this paper.

Diversity in Large Language Models under Supervised Fine-Tuning Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:11:18.148664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-12T03:10:22.314719Z digest=sha256:0b70238220fae10843929a906d589aaf0bb844ae630c404f975d74a853b17e1d

Observation 85bd389e-4357-4b8c-aa69-868cf41f28f9 · inbound

Early Data Exposure Improves Robustness to Subsequent Fine-Tuning cites this paper.

Early Data Exposure Improves Robustness to Subsequent Fine-Tuning Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:47:58.617673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-14T20:45:27.673290Z digest=sha256:39eea30a7c3b6c00aff9117fa1550ad2e319372aeb6a091946fa0003e66a65df

Observation 960bddf1-eedd-4e95-98d1-0465059c9ed9 · inbound

Scaling Laws for Mixture Pretraining Under Data Constraints cites this paper.

Scaling Laws for Mixture Pretraining Under Data Constraints Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:48:00.978072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-14T21:44:31.429223Z digest=sha256:ee01507e610ec3a774e1801cea5cd1d70b99878c0c0ce5b2848bd335f8573c8c

Observation 0701f0ff-3b37-47eb-9a35-895ffbbcad32 · inbound

Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay cites this paper.

Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T23:24:01.970658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T23:15:45.174086Z digest=sha256:8cb5f25a3dd2f0c9e97168910738e228dc35849854d029783820dbe52c1a589d

Observation 0e5085ee-2d3f-4315-a358-1eabbe3127eb · inbound

Knowledge Editing in Masked Diffusion Language Models cites this paper.

Knowledge Editing in Masked Diffusion Language Models Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:36:29.935755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T09:51:42.811856Z digest=sha256:148281ed8e136efc81436f2c4f0061967274f8221ca7c52e471edc96ae0115e6

Observation 85a28b91-5f85-4d9c-af51-c4cff710587a · inbound

A Predict-then-Correct Loop Based on Few-Shot Continuous Contextual Bandit for Demand Forecasting cites this paper.

A Predict-then-Correct Loop Based on Few-Shot Continuous Contextual Bandit for Demand Forecasting Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T22:30:58.664303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:30:58.664303Z digest=sha256:41ad42317d650d3500533788be9fb5ed1371acaf14f9989cf4191127e4daf379