Pith. sign in

Paper Citation Record · LEDGER

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning

As of 13 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 1 inbound Pith citation observation for arXiv:2506.05447.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05447 v2

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:31:20.550988Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T03:02:05.516512Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact3
  • verified fuzzy2
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation a381240c-4449-458e-9476-d0579b6f0a85 · outbound

This paper cites online" 'onlinestring :=.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.306883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.306883Z digest=sha256:35d3453de6b517dae6b8facf7adaca09299058a3fd8aa4ffc97a8fa8b56fa7ad

Observation 673f7902-411a-4e2f-bc9f-3480e903d649 · outbound

This paper cites write newline.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.314438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.314438Z digest=sha256:4f0340a79cf88c7f88736162f14731a32b79efb234e37207a06c67eaab4c4a59

Observation 187ee80c-5b8b-4d2e-bde7-363c5159e6eb · outbound

This paper cites Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.328935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.328935Z digest=sha256:d8354aa54281b410e926fd0365828611174ca4d76d39344f1a69e7c605b01c23

Observation 94b2fce2-c6d0-47c3-86a8-d3e2020a40b2 · outbound

This paper cites Atanasov, Jacob A.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Atanasov, Jacob A

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:31:21.353615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.335963Z digest=sha256:e92a58a758a18a5306e95896d2c101ca9ba8ac6e8723eb2c829ed596d0fb3519

Observation 28757c54-f468-4906-b66d-be7fa4b406f8 · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.341972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.341972Z digest=sha256:f94749f18afd24bd2a68b80ab2a4a5af6c989b1dff32ff4ba9525f1cdea95308

Observation 03619f5c-1633-405e-a256-803fcd2d4e81 · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:21.334057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.347469Z digest=sha256:145ba042a20288cf39eb2871ead4f9e20b19917f8c45f3bc2f03543828d05869

Observation bfa26d7c-1cd2-4395-9cd4-0e91a68d634d · outbound

This paper cites Broken Neural Scaling Laws.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Broken Neural Scaling Laws

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.353468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.353468Z digest=sha256:fe35b613ea8aa9a4024e1a904b92c33debd24cd06efc53a556e01ad121249829

Observation 6c24d1c1-a387-45a4-ab02-9f5023e092bc · outbound

This paper cites Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.359142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.359142Z digest=sha256:dc5f7207b102b5eb27866a9f3db012e078792b2127dd59aaf79dfa903753d88e

Observation f225e47c-d49a-40d2-8c6e-9d2916046295 · outbound

This paper cites Scaling Exponents Across Parameterizations and Optimizers.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Scaling Exponents Across Parameterizations and Optimizers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.364789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.364789Z digest=sha256:801846cbd1b761f8134293c76eb23e161c4b1aa6ca8057879e5971e2a13b3473

Observation ab140c9e-642f-4dae-939a-209b848441b4 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.370181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.370181Z digest=sha256:58b975d6ca546bb42697e1c82a63e34876988c57b6f5bd53babdab5a4b183eef

Observation 3bb8d72f-b197-4dc5-883b-01621b12d273 · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 11

Resolution
verified exact
doi, observed 2026-08-07T10:31:20.907800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.375809Z digest=sha256:7ba9ddc5a3f71a65920c6c20ca9083a7840fd4f4d835f442dc5e04edb9c4aff2

Observation 2988fb8e-629f-48d4-ace9-b76f3659ff9c · outbound

This paper cites Why do small language models underperform? Studying Language Model Saturation via the Softmax Bottleneck.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Why do small language models underperform? Studying Language Model Saturation via the Softmax Bottleneck

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.382188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.382188Z digest=sha256:3eb27e1c2294406ff85d3941e0322ecad31b441592c4333812f7a3847e7d09d0

Observation 1027a3ea-4290-441e-9d5e-0e74972add43 · outbound

This paper cites OLMo: Accelerating the Science of Language Models.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning OLMo: Accelerating the Science of Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.387856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.387856Z digest=sha256:22b740b2b255f95a2a965cd735b930bd022f079b95bd9543e2d3b44556725741

Observation c9652438-c0eb-450e-aaa4-e0fe3a553bd3 · outbound

This paper cites Rae, and Laurent Sifre.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Rae, and Laurent Sifre

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:31:21.308776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.394487Z digest=sha256:0c62e43835a40577f1da93bbf96ceb1244a8f0bd1022d064059e3e0ba7e8ab18

Observation 7fbdc2cd-c03f-4aed-8ad7-76703561d66d · outbound

This paper cites Learning Curve Theory.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Learning Curve Theory

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.400174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.400174Z digest=sha256:057feb5f355d95e89b1177b94e42c320da5792467eef592beb3088551e343c7c

Observation 0e645430-c0ac-4d66-8e53-a2ae77ccffed · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:21.288053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.406165Z digest=sha256:ec81a5748b837c2d274fd92f52cb6ea915eca854ec529b277c9a1477af351f0a

Observation 4de56d0a-b8f0-4b67-ba50-8e5e6a1e5294 · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:21.271940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.411836Z digest=sha256:d38cacd33aa0fcf20b5448a90db2902cccb75f8997f0abf89af7b183aaf2b27c

Observation 190826fb-ace0-4549-bb0a-34db1785d24c · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.417724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.417724Z digest=sha256:e84cc6824da03b231b8f04609456e51f1c7cdab9e8d310dd17e3ccd6cde26ed1

Observation ffb6496d-3839-474e-ad3d-49ff68b8cb74 · outbound

This paper cites Scaling Laws for Neural Language Models.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Scaling Laws for Neural Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.423534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.423534Z digest=sha256:ac106434a9f4dc9f04248d6f1ffc2c183158071ebdb55b4193bdf4ab6c5a4fa6

Observation 811dc1f4-9f45-4c92-b30e-d18d83c128aa · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:21.252089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.428496Z digest=sha256:43be32b96190f557f52da1db34b2fbf06afbd945e58205e5ef836350fd4686ae

Observation cb665d93-206a-4ac8-b17e-ec2bffb28845 · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.433609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.433609Z digest=sha256:d29cc2377d246895b5cf4a64ca2bc2b6e195b5d6b8bc0d5711f39576f8ac042c

Observation 2e8a43a8-a092-47cd-8fa6-5eecde1d9bf7 · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:21.235256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.441055Z digest=sha256:7fdc43ab240a8a6ba5ff8de56a7169a15aff9f710250b591650ba481a79faf25

Observation 8e99f216-4cb2-4585-a3ea-18661cbc1fa2 · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.446261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.446261Z digest=sha256:ebd9cd7beed8b80727b36937542189252c36afba09a3bb1592b96452a6fce90b

Observation b5d9cf26-1b55-47ec-b814-6e278a268e92 · outbound

This paper cites Paloma: A Benchmark for Evaluating Language Model Fit.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Paloma: A Benchmark for Evaluating Language Model Fit

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.450742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.450742Z digest=sha256:1c9ba6041501733291b11aba4705cb93e0dc98decb07438ab35595cb0f6c04a0

Observation 64bb2aec-abbc-4d16-aafe-645de5f60d18 · outbound

This paper cites When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.455874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.455874Z digest=sha256:6f7baed136ab18ab4de33ff934bb050e2ccd145944d0c126820b41834e78002a

Observation 0afb2e2b-3515-4df5-8662-a91f2621106d · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:21.202276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.460735Z digest=sha256:1c96664b886c81bc8b44e946d5ceabbb3d5860d34415183bbf7b47563e8a4936

Observation 736de624-1040-4b06-a353-90424ceda43e · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:21.182023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.465412Z digest=sha256:731490d45d7db1b61bef2472a9e1ebb3d3c34897ae0a751d14948e5b1c2699c1

Observation f894a601-7ebb-4e98-9e4d-ad84dd5733de · outbound

This paper cites In-context Learning and Induction Heads.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning In-context Learning and Induction Heads

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.470207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.470207Z digest=sha256:32a00449124fe95535ba7d69d90d7005e30178039483f7173ec38a3141c7a096

Observation fb5891ae-b06e-403b-9648-dfb55cf572dd · outbound

This paper cites The AdEMAMix Optimizer: Better, Faster, Older.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning The AdEMAMix Optimizer: Better, Faster, Older

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.474636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.474636Z digest=sha256:fa3fe2c16557e57a0ebef5672e1b69ebae6ff31530f84470847531daaa8f2987

Observation 9da32be0-03f4-4c1a-b0d1-cb583aa553db · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:21.166392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.479519Z digest=sha256:47f105f700e819926b3ef703242e7ab94ef8edec8276aea9dd62f72c16ced6f2

Observation 77df1296-b4a4-4f56-962b-3c7e4e1842d9 · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:21.149829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.484527Z digest=sha256:830d52a3792a22e13580e38def6aa1f1e4ef0c037f820611e4ce11b3159d6b0f

Observation 07729239-587b-4a1a-ae67-0d987caa8ced · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.490228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.490228Z digest=sha256:33a0c524a4a03974ab1278df1131a76d7e36d3df6752a69fff649207d010cd60

Observation 3d2cf27f-cac3-4065-9290-b35030a91c69 · outbound

This paper cites Outliers with Opposing Signals Have an Outsized Effect on Neural Network Optimization.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Outliers with Opposing Signals Have an Outsized Effect on Neural Network Optimization

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:31:20.681375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.496089Z digest=sha256:477437a641f9e3fa4182da472fbbfa0ca776671d45f997c7fecad7696edc5e81

Observation 6162da8a-0c80-4923-bf02-e979cd7c508b · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:21.133494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.501651Z digest=sha256:1e1d46a941b90856de257206d281eab3ad3b0597bee4f524de5f2eaf9e25bf40

Observation 0acf9fcd-435d-40de-8a1c-bb12889ed4e6 · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:21.117029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.506125Z digest=sha256:6fb2d15b43c6fcb5f5f455e3a08f6eb7b61cd2cbfda752f8bb4249217d61864e

Observation e74392e0-56ae-4ef1-9ee2-b5fde6cddedc · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.511694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.511694Z digest=sha256:1a0577acd8487257fa428ed71aad524223314bbc1f15623a98bed7412b8aaaa3

Observation 2b8ec33d-d619-4193-a0c7-3afc4d71caad · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Gemma 2: Improving Open Language Models at a Practical Size

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.516813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.516813Z digest=sha256:6ee534f268759d12d4aa44c1ba008142a4df434c5176604822cd002a5c706ec7

Observation 8b1f55d0-2325-4943-aadd-c42eb35e4330 · outbound

This paper cites Scaling Law with Learning Rate Annealing.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Scaling Law with Learning Rate Annealing

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.522740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.522740Z digest=sha256:81a8207d8d8954bb511329d85af7eb0ef2d09c1902ae8ff33319a2f20dacf6d0

Observation 0ef2f770-b496-4cd6-ae99-715b7f8e8c68 · outbound

This paper cites The Shape of Learning Curves: a Review.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning The Shape of Learning Curves: a Review

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.528055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.528055Z digest=sha256:7b20699533a587572bb1cc14ad8f3c05912953a4325d06a731ac0b1c163d579a

Observation 22a06aaa-38ec-4f17-98a5-881b1a0aa397 · outbound

This paper cites Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.534009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.534009Z digest=sha256:5e2406d4142ba2e11ac7a0d0d30054d47317d3800f25fadbe6f9ab30e9921a45

Observation a13dbdef-d570-432a-b52a-5c9b2282293a · outbound

This paper cites Data-Dependence of Plateau Phenomenon in Learning with Neural Network --- Statistical Mechanical Analysis.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Data-Dependence of Plateau Phenomenon in Learning with Neural Network --- Statistical Mechanical Analysis

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:31:21.027574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.539480Z digest=sha256:fd933c6ebef2ea64a92d2963076f17302f1c71638d23cebc5fddc9f2f6bd38b4

Observation 380618aa-20e5-4ba1-9f0c-1ad070b497d9 · outbound

This paper cites Opacus: User-Friendly Differential Privacy Library in PyTorch.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Opacus: User-Friendly Differential Privacy Library in PyTorch

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.544798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.544798Z digest=sha256:8802cfb0c8b91854e972ab278de6b51a1ad8e755536fd1359337081ff509b280

Observation 62f9b0f3-c2ac-4530-842f-51ca8ffb82cf · outbound

This paper cites Gradient Surgery for Multi-Task Learning.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Gradient Surgery for Multi-Task Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.550988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.550988Z digest=sha256:0f5d1e9fa657d5e6ed09e170ff402726f5992bfab4258aeabaf14300362ea103

Pith citing papers

Observation 21d0767b-bbd7-44ae-b12e-a8508b3d6055 · inbound

Bridging Compute- and Data-Optimal Pretraining cites this paper.

Bridging Compute- and Data-Optimal Pretraining Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-08-01T03:08:36.328547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-01T03:02:05.516512Z digest=sha256:f0f493bdc1153d6a1515ae22c0d57e59fab3c80ec07bd387ab5cbfa1d44bdf18