Pith. sign in

Paper Citation Record · LEDGER

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning

As of 13 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 1 inbound Pith citation observation for arXiv:2506.05447.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05447 v2

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:31:20.550988Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T03:02:05.516512Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact3
  • verified fuzzy2
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation a381240c-4449-458e-9476-d0579b6f0a85 · outbound

This paper cites online" 'onlinestring :=.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.306883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.306883Z digest=sha256:35d3453de6b517dae6b8facf7adaca09299058a3fd8aa4ffc97a8fa8b56fa7ad

Observation 673f7902-411a-4e2f-bc9f-3480e903d649 · outbound

This paper cites write newline.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.314438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.314438Z digest=sha256:4f0340a79cf88c7f88736162f14731a32b79efb234e37207a06c67eaab4c4a59

Observation 187ee80c-5b8b-4d2e-bde7-363c5159e6eb · outbound

This paper cites Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.328935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.328935Z digest=sha256:d8354aa54281b410e926fd0365828611174ca4d76d39344f1a69e7c605b01c23

Observation 94b2fce2-c6d0-47c3-86a8-d3e2020a40b2 · outbound

This paper cites Atanasov, Jacob A.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Atanasov, Jacob A

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:31:21.353615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.335963Z digest=sha256:08f65da7b5b44861b48443d8db3dfb20b4f5b1732453cd06e470865455bec2e7

Observation 28757c54-f468-4906-b66d-be7fa4b406f8 · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.341972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.341972Z digest=sha256:f94749f18afd24bd2a68b80ab2a4a5af6c989b1dff32ff4ba9525f1cdea95308

Observation 03619f5c-1633-405e-a256-803fcd2d4e81 · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:21.334057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.347469Z digest=sha256:5bfa2c4c5747700fd72234bd82ca1372b0e4cd96cb6c5185cc3bf47d5faf2e88

Observation bfa26d7c-1cd2-4395-9cd4-0e91a68d634d · outbound

This paper cites Broken Neural Scaling Laws.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Broken Neural Scaling Laws

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.353468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.353468Z digest=sha256:9a74b50d8f95d6f9f2fc16bc2f02cc3d755af2044647e4d0b7a039e3e07cc01f

Observation 6c24d1c1-a387-45a4-ab02-9f5023e092bc · outbound

This paper cites Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.359142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.359142Z digest=sha256:dc5f7207b102b5eb27866a9f3db012e078792b2127dd59aaf79dfa903753d88e

Observation f225e47c-d49a-40d2-8c6e-9d2916046295 · outbound

This paper cites Scaling Exponents Across Parameterizations and Optimizers.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Scaling Exponents Across Parameterizations and Optimizers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.364789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.364789Z digest=sha256:801846cbd1b761f8134293c76eb23e161c4b1aa6ca8057879e5971e2a13b3473

Observation ab140c9e-642f-4dae-939a-209b848441b4 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.370181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.370181Z digest=sha256:58b975d6ca546bb42697e1c82a63e34876988c57b6f5bd53babdab5a4b183eef

Observation 3bb8d72f-b197-4dc5-883b-01621b12d273 · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 11

Resolution
verified exact
doi, observed 2026-08-07T10:31:20.907800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.375809Z digest=sha256:c41a3d0c73d7458a9af6ce266c056fe4de5d59da360f6cf9143bd6cd892c9f50

Observation 2988fb8e-629f-48d4-ace9-b76f3659ff9c · outbound

This paper cites Why do small language models underperform? Studying Language Model Saturation via the Softmax Bottleneck.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Why do small language models underperform? Studying Language Model Saturation via the Softmax Bottleneck

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.382188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.382188Z digest=sha256:7ff0de8d2c11c95ad5c7d5b9638fa6fce45bb8e7c992233b154b1ec926856a23

Observation 1027a3ea-4290-441e-9d5e-0e74972add43 · outbound

This paper cites OLMo: Accelerating the Science of Language Models.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning OLMo: Accelerating the Science of Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.387856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.387856Z digest=sha256:22b740b2b255f95a2a965cd735b930bd022f079b95bd9543e2d3b44556725741

Observation c9652438-c0eb-450e-aaa4-e0fe3a553bd3 · outbound

This paper cites Rae, and Laurent Sifre.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Rae, and Laurent Sifre

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:31:21.308776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.394487Z digest=sha256:a53da31feca5717687f262f4563ff6bd1179b6bdf166769185b69212184b77ef

Observation 7fbdc2cd-c03f-4aed-8ad7-76703561d66d · outbound

This paper cites Learning Curve Theory.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Learning Curve Theory

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.400174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.400174Z digest=sha256:057feb5f355d95e89b1177b94e42c320da5792467eef592beb3088551e343c7c

Observation 0e645430-c0ac-4d66-8e53-a2ae77ccffed · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:21.288053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.406165Z digest=sha256:efb9fb1b435b5244975e6f582084f03812d5be9c4d470d43dd482044b13051cf

Observation 4de56d0a-b8f0-4b67-ba50-8e5e6a1e5294 · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:21.271940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.411836Z digest=sha256:09828d1937936e874ccf115192886327278f955c788c958bfff6e52d4f39add3

Observation 190826fb-ace0-4549-bb0a-34db1785d24c · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.417724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.417724Z digest=sha256:e84cc6824da03b231b8f04609456e51f1c7cdab9e8d310dd17e3ccd6cde26ed1

Observation ffb6496d-3839-474e-ad3d-49ff68b8cb74 · outbound

This paper cites Scaling Laws for Neural Language Models.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Scaling Laws for Neural Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.423534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.423534Z digest=sha256:7c860d4f71148e62b122add5cca4d98ccc1948c586ed4469cdf5ae8d186e1099

Observation 811dc1f4-9f45-4c92-b30e-d18d83c128aa · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:21.252089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.428496Z digest=sha256:80135cf7617bc101bed0d242a802514c9d109477f0d0af8d9ca5116eee3f9bb6

Observation cb665d93-206a-4ac8-b17e-ec2bffb28845 · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.433609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.433609Z digest=sha256:d29cc2377d246895b5cf4a64ca2bc2b6e195b5d6b8bc0d5711f39576f8ac042c

Observation 2e8a43a8-a092-47cd-8fa6-5eecde1d9bf7 · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:21.235256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.441055Z digest=sha256:8304b8954683a8cd1dd85e22f5c4ef27251ed43ab96a0c8a38880dafa4f3f15d

Observation 8e99f216-4cb2-4585-a3ea-18661cbc1fa2 · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.446261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.446261Z digest=sha256:ebd9cd7beed8b80727b36937542189252c36afba09a3bb1592b96452a6fce90b

Observation b5d9cf26-1b55-47ec-b814-6e278a268e92 · outbound

This paper cites Paloma: A Benchmark for Evaluating Language Model Fit.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Paloma: A Benchmark for Evaluating Language Model Fit

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.450742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.450742Z digest=sha256:1c9ba6041501733291b11aba4705cb93e0dc98decb07438ab35595cb0f6c04a0

Observation 64bb2aec-abbc-4d16-aafe-645de5f60d18 · outbound

This paper cites When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.455874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.455874Z digest=sha256:206558ffdf3943ba8bcee8b525137ab3795a7815a6079812e353311f96f6531f

Observation 0afb2e2b-3515-4df5-8662-a91f2621106d · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:21.202276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.460735Z digest=sha256:b6e6289bf808905816803f7815557f9e03e34deca97a71f3fe088a3fd55ace3c

Observation 736de624-1040-4b06-a353-90424ceda43e · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:21.182023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.465412Z digest=sha256:db3de4240ed94350eb2831fe1723451d827b414d0c8826a698330847fb8baff2

Observation f894a601-7ebb-4e98-9e4d-ad84dd5733de · outbound

This paper cites In-context Learning and Induction Heads.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning In-context Learning and Induction Heads

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.470207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.470207Z digest=sha256:32a00449124fe95535ba7d69d90d7005e30178039483f7173ec38a3141c7a096

Observation fb5891ae-b06e-403b-9648-dfb55cf572dd · outbound

This paper cites The AdEMAMix Optimizer: Better, Faster, Older.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning The AdEMAMix Optimizer: Better, Faster, Older

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.474636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.474636Z digest=sha256:fa3fe2c16557e57a0ebef5672e1b69ebae6ff31530f84470847531daaa8f2987

Observation 9da32be0-03f4-4c1a-b0d1-cb583aa553db · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:21.166392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.479519Z digest=sha256:2e69a6f176497b3d0d805ef642fdb5a6a02679e588433edbc379557500591357

Observation 77df1296-b4a4-4f56-962b-3c7e4e1842d9 · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:21.149829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.484527Z digest=sha256:a47561913d15a0f646769521ba0a929bfb18a57af0230e2483074cc6cfc37c4c

Observation 07729239-587b-4a1a-ae67-0d987caa8ced · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.490228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.490228Z digest=sha256:33a0c524a4a03974ab1278df1131a76d7e36d3df6752a69fff649207d010cd60

Observation 3d2cf27f-cac3-4065-9290-b35030a91c69 · outbound

This paper cites Outliers with Opposing Signals Have an Outsized Effect on Neural Network Optimization.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Outliers with Opposing Signals Have an Outsized Effect on Neural Network Optimization

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:31:20.681375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.496089Z digest=sha256:63b62acf75e61e77cb91ec573c17f05fa37d14074300a66f0a73a1dc4f74827e

Observation 6162da8a-0c80-4923-bf02-e979cd7c508b · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:21.133494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.501651Z digest=sha256:f846ab3a4f8b14150a40f71482c637c68561342f7404b9f7e3d4120a3837962c

Observation 0acf9fcd-435d-40de-8a1c-bb12889ed4e6 · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:21.117029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.506125Z digest=sha256:9cc7405b75bdc9ac60295d0b907ec9bf1bb8cca4da51f932f8bf333ddc26713b

Observation e74392e0-56ae-4ef1-9ee2-b5fde6cddedc · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.511694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.511694Z digest=sha256:1a0577acd8487257fa428ed71aad524223314bbc1f15623a98bed7412b8aaaa3

Observation 2b8ec33d-d619-4193-a0c7-3afc4d71caad · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Gemma 2: Improving Open Language Models at a Practical Size

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.516813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.516813Z digest=sha256:6ee534f268759d12d4aa44c1ba008142a4df434c5176604822cd002a5c706ec7

Observation 8b1f55d0-2325-4943-aadd-c42eb35e4330 · outbound

This paper cites Scaling Law with Learning Rate Annealing.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Scaling Law with Learning Rate Annealing

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.522740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.522740Z digest=sha256:81a8207d8d8954bb511329d85af7eb0ef2d09c1902ae8ff33319a2f20dacf6d0

Observation 0ef2f770-b496-4cd6-ae99-715b7f8e8c68 · outbound

This paper cites The Shape of Learning Curves: a Review.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning The Shape of Learning Curves: a Review

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.528055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.528055Z digest=sha256:7b20699533a587572bb1cc14ad8f3c05912953a4325d06a731ac0b1c163d579a

Observation 22a06aaa-38ec-4f17-98a5-881b1a0aa397 · outbound

This paper cites Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.534009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.534009Z digest=sha256:5e2406d4142ba2e11ac7a0d0d30054d47317d3800f25fadbe6f9ab30e9921a45

Observation a13dbdef-d570-432a-b52a-5c9b2282293a · outbound

This paper cites Data-Dependence of Plateau Phenomenon in Learning with Neural Network --- Statistical Mechanical Analysis.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Data-Dependence of Plateau Phenomenon in Learning with Neural Network --- Statistical Mechanical Analysis

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:31:21.027574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.539480Z digest=sha256:b1c608f332752efa4f45d0a8baf3a7adcde453c5214062c3e2eabc0f88c37ee7

Observation 380618aa-20e5-4ba1-9f0c-1ad070b497d9 · outbound

This paper cites Opacus: User-Friendly Differential Privacy Library in PyTorch.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Opacus: User-Friendly Differential Privacy Library in PyTorch

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.544798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.544798Z digest=sha256:8802cfb0c8b91854e972ab278de6b51a1ad8e755536fd1359337081ff509b280

Observation 62f9b0f3-c2ac-4530-842f-51ca8ffb82cf · outbound

This paper cites Gradient Surgery for Multi-Task Learning.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Gradient Surgery for Multi-Task Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.550988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.550988Z digest=sha256:0f5d1e9fa657d5e6ed09e170ff402726f5992bfab4258aeabaf14300362ea103

Pith citing papers

Observation 21d0767b-bbd7-44ae-b12e-a8508b3d6055 · inbound

Bridging Compute- and Data-Optimal Pretraining cites this paper.

Bridging Compute- and Data-Optimal Pretraining Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-08-01T03:08:36.328547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-01T03:02:05.516512Z digest=sha256:ed9115285c1571191b613910fb5b913a5b6551130f973a390a24e75f3bc3ef15