Pith. sign in

Paper Citation Record · LEDGER

Learning Dynamics in Continual Pre-Training for Large Language Models

As of 19 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 1 inbound Pith citation observation for arXiv:2505.07796.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.07796 v2

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:14:19.319295Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:19:10.282840Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T04:19:11.804438Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 00532d16-4d10-4e2f-8e6b-d07cc7d4fae4 · outbound

This paper cites G., and Bakshy, E.

Learning Dynamics in Continual Pre-Training for Large Language Models G., and Bakshy, E

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:14:20.411132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T22:14:19.077192Z digest=sha256:4a58bbdfa757edb1c5c341530a9b6d3c1e68142b3000f9629f3636c9d034f94a

Observation 43576439-1807-48fe-b43a-972acdf4a1d6 · outbound

This paper cites An Empirical Study of Scaling Laws for Transfer.

Learning Dynamics in Continual Pre-Training for Large Language Models An Empirical Study of Scaling Laws for Transfer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.082887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.082887Z digest=sha256:40df4baf76d21ac0d904a781e6bd6d5e5670bdbe9fa01b23726696e93961cf84

Observation ebbd0cbc-f12b-499f-a7a1-f2f1a7a5d4e1 · outbound

This paper cites and Bengio, Y.

Learning Dynamics in Continual Pre-Training for Large Language Models and Bengio, Y

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.088591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.088591Z digest=sha256:a56e2f8cae111841cb75329c47956f0d6046ece65a357793771872c73ddbeacf

Observation 5ecb05d9-559d-4b03-9bff-ee522d171bc5 · outbound

This paper cites Lifelong language pretraining with distribution-specialized experts.

Learning Dynamics in Continual Pre-Training for Large Language Models Lifelong language pretraining with distribution-specialized experts

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:14:20.383663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T22:14:19.093529Z digest=sha256:ce1290f00c8fec189a0f964f6c251a5881eaea22a0a807b9a95a89fea838ae02

Observation 008670e1-7564-4fd7-b4f8-c6d11b176cf9 · outbound

This paper cites MEDITRON-70B: Scaling Medical Pretraining for Large Language Models.

Learning Dynamics in Continual Pre-Training for Large Language Models MEDITRON-70B: Scaling Medical Pretraining for Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.098691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.098691Z digest=sha256:52f82ec9675a8e8c5e1419fb589bac4b975c12cd539bc6d7bb1ba30c20843b1d

Observation d739a6d0-6212-474e-926a-bcb0e6c93233 · outbound

This paper cites SaulLM-7B: A pioneering Large Language Model for Law.

Learning Dynamics in Continual Pre-Training for Large Language Models SaulLM-7B: A pioneering Large Language Model for Law

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.104299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.104299Z digest=sha256:ea7ae666710cf118c43899037ea106f5a550bada499f41e877f1aaa2a125af10

Observation 0a8ef6e3-2820-47d0-bef0-7aa2fa8edb96 · outbound

This paper cites DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence.

Learning Dynamics in Continual Pre-Training for Large Language Models DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.110481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.110481Z digest=sha256:1ee20f904369513c9a9cce95c43aab6804f7e44af1992dc7f419576caefcd9dc

Observation d0aa8343-9a9d-43f4-af2d-82d2479e046c · outbound

This paper cites Sailor: Open Language Models for South-East Asia.

Learning Dynamics in Continual Pre-Training for Large Language Models Sailor: Open Language Models for South-East Asia

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.115873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.115873Z digest=sha256:1c6356fec5cb8e8b8860ed9907a919631d80e67e903a74580632e40978e9f4ca

Observation b6215e37-1c44-43f5-812f-d573e6aceed3 · outbound

This paper cites The Llama 3 Herd of Models.

Learning Dynamics in Continual Pre-Training for Large Language Models The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.121164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.121164Z digest=sha256:a1a3b97e758fd2a30d39c5e53cd1034aeb54b13a930bb738d4eb3b95dcd8b858

Observation 2d02dedf-45d9-4870-a968-728aa2f832cb · outbound

This paper cites Unearthing Large Scale Domain-Specific Knowledge from Public Corpora.

Learning Dynamics in Continual Pre-Training for Large Language Models Unearthing Large Scale Domain-Specific Knowledge from Public Corpora

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.126804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.126804Z digest=sha256:cd916a609b9ce5928d692433217e86ddda732be91040e6b8d114715e15a68e7e

Observation 81f49ded-fce9-4a5a-849c-2229406caf03 · outbound

This paper cites an unresolved cited work.

Learning Dynamics in Continual Pre-Training for Large Language Models Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.132108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.132108Z digest=sha256:c9fe51ca5aa289882f3e45bfe47c0a052bc6e18a856c9a879abcf7d4418348f6

Observation 98878f17-b8be-4536-999a-06ac6f02c0c1 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Learning Dynamics in Continual Pre-Training for Large Language Models The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.137187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.137187Z digest=sha256:f96f539cdd4c62d6fe45635eb88be76477970e6d7373f8f94bf24bea0a18f07f

Observation f6f9a9fc-d37a-4aa5-aa03-fb9bb266a78d · outbound

This paper cites CMR scaling law: Predicting critical mixture ratios for continual pre-training of language models.

Learning Dynamics in Continual Pre-Training for Large Language Models CMR scaling law: Predicting critical mixture ratios for continual pre-training of language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.142025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.142025Z digest=sha256:c53721e791977336d0467a01d4299b560acface9ff16464626b8632b54406821

Observation 8f197b83-e306-4a0c-9dcb-19b67cd9a90b · outbound

This paper cites Continual Pre-Training of Large Language Models: How to (re)warm your model?.

Learning Dynamics in Continual Pre-Training for Large Language Models Continual Pre-Training of Large Language Models: How to (re)warm your model?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.146891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.146891Z digest=sha256:22100de51270769bd2a5bc2f46088ebc469e2791270ec1f1186c1a185da85664

Observation aea6de6c-d97c-442a-ac0e-1e1a656a6954 · outbound

This paper cites Data Mixture Inference: What do BPE Tokenizers Reveal about their Training Data?.

Learning Dynamics in Continual Pre-Training for Large Language Models Data Mixture Inference: What do BPE Tokenizers Reveal about their Training Data?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.151913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.151913Z digest=sha256:09d3b4655a84fbc211ad2fe37811f76aae9641ee078a7f63475143c38cbe5296

Observation efb64f27-ea03-4570-afd7-15484ead5ddb · outbound

This paper cites Pile of Law: Learning Responsible Data Filtering from the Law and a 256GB Open-Source Legal Dataset.

Learning Dynamics in Continual Pre-Training for Large Language Models Pile of Law: Learning Responsible Data Filtering from the Law and a 256GB Open-Source Legal Dataset

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.156722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.156722Z digest=sha256:ab30112a80697193d250a25465a2ff7289e1d61922d7822a33bac30f7610e0c5

Observation c955415a-2ab1-4134-822e-bea0556c528a · outbound

This paper cites Scaling Laws for Transfer.

Learning Dynamics in Continual Pre-Training for Large Language Models Scaling Laws for Transfer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.166601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.166601Z digest=sha256:b0b3b60d00a4ddd8b2963e244dbc53c8b3ea1bd44c78a25791dd68ac6d64214e

Observation 924fd2e6-07e5-452f-9712-8f8363a75d86 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Learning Dynamics in Continual Pre-Training for Large Language Models Training Compute-Optimal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.171347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.171347Z digest=sha256:9b49c94957a06ac64982aca6edc4282cb80bff666b11d73a7d637423aadf5382

Observation d121661f-28b3-4877-aeae-050ecf8bd875 · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

Learning Dynamics in Continual Pre-Training for Large Language Models MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.176388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.176388Z digest=sha256:6305e4727535b61fa6bdfe769cf69f6738430bd3d72beae539592783fd94127d

Observation 34833b2e-1659-4506-8777-22dbf7c8949e · outbound

This paper cites an unresolved cited work.

Learning Dynamics in Continual Pre-Training for Large Language Models Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.181367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.181367Z digest=sha256:f034ca195d0ce78ab674bc372220ad5d685758bc7e7a6882078e16998766b3d3

Observation 5fe99d26-80ac-4698-9dac-b1857eb3ffc0 · outbound

This paper cites Qwen2.5-Coder Technical Report.

Learning Dynamics in Continual Pre-Training for Large Language Models Qwen2.5-Coder Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.185976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.185976Z digest=sha256:74e84b6e514c69b1047fce90c0b386896e6030d4679cb083b9400daa9371e969

Observation 2e60ba88-dd35-4f19-aa2c-6ce5fc1e1df3 · outbound

This paper cites L., Anthony, Q.

Learning Dynamics in Continual Pre-Training for Large Language Models L., Anthony, Q

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:14:20.355983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T22:14:19.190850Z digest=sha256:107959759f6ec97c35c6e430c3ed3d3733486adad029087fc24c03ac7a176101

Observation 459ae38a-4482-428d-9128-83b4ab155c1d · outbound

This paper cites Scaling Laws for Neural Language Models.

Learning Dynamics in Continual Pre-Training for Large Language Models Scaling Laws for Neural Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.195460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.195460Z digest=sha256:66956a4a5b3c2bae388f508b922d579d648410f53a2794a2f85e7dc453083ffb

Observation 68069132-8bea-4109-b031-5cbd052be450 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Learning Dynamics in Continual Pre-Training for Large Language Models Adam: A Method for Stochastic Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.200692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.200692Z digest=sha256:7c6b4e60ff3a59251769260139804078d813735d947d65a7e2ca7be969502ba5

Observation e8a86594-ef36-4e22-a423-6908413e4ed9 · outbound

This paper cites D., van de Ven, G.

Learning Dynamics in Continual Pre-Training for Large Language Models D., van de Ven, G

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:14:20.339444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T22:14:19.205364Z digest=sha256:82e5d9f6fac6c88bfb0bd07e5456ca19f41abed9de68ad594cfa9edc4d79114d

Observation 9b3c9d46-f34f-41a5-9d27-fa350dd6b9bd · outbound

This paper cites Crafting papers on machine learning.

Learning Dynamics in Continual Pre-Training for Large Language Models Crafting papers on machine learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.210233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.210233Z digest=sha256:0a1f18820cc729667e7c22243ed982ccdcca98aedff8626b5d954fa78c5b9348

Observation 8d3ac477-eb42-470d-989a-cd4faea88925 · outbound

This paper cites and Hutter, F.

Learning Dynamics in Continual Pre-Training for Large Language Models and Hutter, F

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:14:20.309928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T22:14:19.214951Z digest=sha256:1268da9abeb7dc913ad17330386785ca36b91c60d814e890de5c6a28383f0f5c

Observation 2e51e396-b711-4e09-9e85-55a0c2fbbb37 · outbound

This paper cites Decoupled Weight Decay Regularization.

Learning Dynamics in Continual Pre-Training for Large Language Models Decoupled Weight Decay Regularization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.219560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.219560Z digest=sha256:f7c3deacc7cbc2912771c2361ef5e79b8b40d49b3a41250c5dd23a044e66ae81

Observation 7bd6e633-6962-47b5-b386-475c95976855 · outbound

This paper cites A multi-power law for loss curve prediction across learning rate schedules.

Learning Dynamics in Continual Pre-Training for Large Language Models A multi-power law for loss curve prediction across learning rate schedules

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:14:20.290973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T22:14:19.224165Z digest=sha256:1301190777b1eb06972ec7c060b2dd07e606c63677b4c06d4c1456021f87f09e

Observation 989c7cc4-f2b7-4702-b280-e8eb39b9ad2c · outbound

This paper cites Updating quasi newton matrices with limited storage.

Learning Dynamics in Continual Pre-Training for Large Language Models Updating quasi newton matrices with limited storage

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.229110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.229110Z digest=sha256:c3b0a241663876cb29d31d0d7210d5b632b4dc3ba9964cf39da340bc22a1376e

Observation fdee20b3-088e-4cd9-a63a-432b79da5f9e · outbound

This paper cites GPT-4 Technical Report.

Learning Dynamics in Continual Pre-Training for Large Language Models GPT-4 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.233882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.233882Z digest=sha256:f5477987ab6e24e4d311e0318a5c676fe10bc14bf763c285d80b0ebb26583ef3

Observation f9f86186-71aa-4f82-882d-d1a69542bfc8 · outbound

This paper cites Reuse, Don't Retrain: A Recipe for Continued Pretraining of Language Models.

Learning Dynamics in Continual Pre-Training for Large Language Models Reuse, Don't Retrain: A Recipe for Continued Pretraining of Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.238721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.238721Z digest=sha256:c04879ef4dfb22211d1e051739c5b0fe289d02d250b7da4fd13d6a8fbff605d8

Observation 861df785-cb40-41e5-bb8b-024345a2a11c · outbound

This paper cites D., Azerbayev, Z., and Ba, J.

Learning Dynamics in Continual Pre-Training for Large Language Models D., Azerbayev, Z., and Ba, J

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:14:20.273507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T22:14:19.243666Z digest=sha256:4c315d41999da20a704fbde62a0e50ea7c4b544a055269db9607c46a150b9a4d

Observation 1f09191e-1298-4fc4-aa62-52aa61bb9fad · outbound

This paper cites B., Lozhkov, A., Mitchell, M., Raffel, C., Werra, L.

Learning Dynamics in Continual Pre-Training for Large Language Models B., Lozhkov, A., Mitchell, M., Raffel, C., Werra, L

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.249541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.249541Z digest=sha256:e8117d78fac035d3556c800b02949cb5334b5c9e17dbdac2fc5b1daa82a12513

Observation a2ef0944-f9cd-415d-b159-fe312a74b41a · outbound

This paper cites D- CPT law: Domain-specific continual pre-training scaling law for large language models.

Learning Dynamics in Continual Pre-Training for Large Language Models D- CPT law: Domain-specific continual pre-training scaling law for large language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:14:20.243990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T22:14:19.254361Z digest=sha256:2fdfd3a31014df34c28e66abebfc3b78f7d8a0881438573d82f6aac22b37824f

Observation 85f9ac97-6833-42f3-987d-cddf4ef01bcf · outbound

This paper cites Continual Learning of Large Language Models: A Comprehensive Survey.

Learning Dynamics in Continual Pre-Training for Large Language Models Continual Learning of Large Language Models: A Comprehensive Survey

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.259110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.259110Z digest=sha256:38c9b38f4d90ba9e3368440998c8e1252c771a1184315739a4a5d7cdb9c7dd74

Observation a839d7a2-7450-47b5-9c0b-1a24bd2dda44 · outbound

This paper cites an unresolved cited work.

Learning Dynamics in Continual Pre-Training for Large Language Models Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.264207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.264207Z digest=sha256:1b23b386af3d10d4dc1570f31695f8082b064df65cbbf95b9cf8a07dfb49897a

Observation 321a79f9-b434-444b-8146-0f2629a50b16 · outbound

This paper cites R., Hestness, J., and Dey, N.

Learning Dynamics in Continual Pre-Training for Large Language Models R., Hestness, J., and Dey, N

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.269275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.269275Z digest=sha256:c27300bdb69581e9faf6b0bb72b2516f61cf1892ce9d60c66caba6fb918b877d

Observation 7cbb1862-92dd-443b-af9b-41f1447b764f · outbound

This paper cites Scaling Law with Learning Rate Annealing.

Learning Dynamics in Continual Pre-Training for Large Language Models Scaling Law with Learning Rate Annealing

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.273955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.273955Z digest=sha256:8e7d90cf8bc02ef5bc206ab2972e16ef33c82737c3f8c8b6ea020213f6e98196

Observation 5f722a6d-d6be-452a-a960-a2ed3c795b4c · outbound

This paper cites Learning to prompt for continual learning.

Learning Dynamics in Continual Pre-Training for Large Language Models Learning to prompt for continual learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.279303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.279303Z digest=sha256:9150e6eae9b1d4140b1b5d46ab94d4940d6d7e0e9d6179c941784516231663b3

Observation 58551376-09c2-48f3-af18-a0df838298b2 · outbound

This paper cites A learning rate path switching training paradigm for version updates of large language models.

Learning Dynamics in Continual Pre-Training for Large Language Models A learning rate path switching training paradigm for version updates of large language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:14:20.202006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T22:14:19.284081Z digest=sha256:c266529fc6d37db24381d557663a0b2f7aa8ef923bdbf92584f11e79d3c3716b

Observation 6676a362-fa8a-4ecc-a150-6af55c89abbf · outbound

This paper cites an unresolved cited work.

Learning Dynamics in Continual Pre-Training for Large Language Models Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.289495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.289495Z digest=sha256:bc7e3e28cbf879fcb1358df02a45389a79ef4384f19545baa57d86b7db2d5b97

Observation f304d238-6658-47b8-a1da-c9b68b70feff · outbound

This paper cites Optimization Hyper-parameter Laws for Large Language Models.

Learning Dynamics in Continual Pre-Training for Large Language Models Optimization Hyper-parameter Laws for Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.295472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.295472Z digest=sha256:c2f65543fdd05ed9f20a0e6050bd4632b00cbe72b19b151fd16ee9abf61cb11f

Observation 5971afbb-59ef-4cc7-8894-97b5c7ac19fd · outbound

This paper cites Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance.

Learning Dynamics in Continual Pre-Training for Large Language Models Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.300672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.300672Z digest=sha256:28032b6ea6ab791fd554109f9c75fd594dc45243160ca13d525ab1464f018baa

Observation 9d9b136e-0af0-43a8-8371-126b20a2042f · outbound

This paper cites Domain2vec: Vectorizing datasets to find the optimal data mixture without training, 2025.

Learning Dynamics in Continual Pre-Training for Large Language Models Domain2vec: Vectorizing datasets to find the optimal data mixture without training, 2025

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:14:20.172977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T22:14:19.307989Z digest=sha256:5471c17610a4aed8e0014683d0568a8142185f53660dc0f58e91e93abb991855

Observation 6563d047-ff19-4651-b5fc-a2132e7fe312 · outbound

This paper cites Investigating Continual Pretraining in Large Language Models: Insights and Implications.

Learning Dynamics in Continual Pre-Training for Large Language Models Investigating Continual Pretraining in Large Language Models: Insights and Implications

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.313213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.313213Z digest=sha256:93edb8361977c23dcbfaa7581ddc13bd53855c805de7e98253d5845f3f1f8be3

Observation 396d8970-4cd2-4e07-8d1d-91227aeadf41 · outbound

This paper cites write newline.

Learning Dynamics in Continual Pre-Training for Large Language Models write newline

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:19.319295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:19.319295Z digest=sha256:105a70bcceabf34b3a1ace398e133e3901a463a00a1e2d35bf7a329f4e18fe4c

Pith citing papers

Observation f881e0e2-8d5b-4e42-93fc-65d87c471d6c · inbound

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training cites this paper.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Learning Dynamics in Continual Pre-Training for Large Language Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:19:11.911632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T04:19:10.282840Z digest=sha256:8296ec5c4f3f72fb72016a12b37b35be27206a1de4db99c11c0846b5e4403c9b