Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:31:20.550988Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 1 inbound Pith citation observation for arXiv:2506.05447.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:31:20.550988Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T03:02:05.516512Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
43 of 43 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
Observation a381240c-4449-458e-9476-d0579b6f0a85 · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning online" 'onlinestring :=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 673f7902-411a-4e2f-bc9f-3480e903d649 · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 187ee80c-5b8b-4d2e-bde7-363c5159e6eb · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94b2fce2-c6d0-47c3-86a8-d3e2020a40b2 · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Atanasov, Jacob A
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 28757c54-f468-4906-b66d-be7fa4b406f8 · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03619f5c-1633-405e-a256-803fcd2d4e81 · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bfa26d7c-1cd2-4395-9cd4-0e91a68d634d · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Broken Neural Scaling Laws
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c24d1c1-a387-45a4-ab02-9f5023e092bc · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f225e47c-d49a-40d2-8c6e-9d2916046295 · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Scaling Exponents Across Parameterizations and Optimizers
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab140c9e-642f-4dae-939a-209b848441b4 · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bb8d72f-b197-4dc5-883b-01621b12d273 · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2988fb8e-629f-48d4-ace9-b76f3659ff9c · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Why do small language models underperform? Studying Language Model Saturation via the Softmax Bottleneck
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1027a3ea-4290-441e-9d5e-0e74972add43 · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning OLMo: Accelerating the Science of Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9652438-c0eb-450e-aaa4-e0fe3a553bd3 · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Rae, and Laurent Sifre
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7fbdc2cd-c03f-4aed-8ad7-76703561d66d · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Learning Curve Theory
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e645430-c0ac-4d66-8e53-a2ae77ccffed · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4de56d0a-b8f0-4b67-ba50-8e5e6a1e5294 · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 190826fb-ace0-4549-bb0a-34db1785d24c · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffb6496d-3839-474e-ad3d-49ff68b8cb74 · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Scaling Laws for Neural Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 811dc1f4-9f45-4c92-b30e-d18d83c128aa · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cb665d93-206a-4ac8-b17e-ec2bffb28845 · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e8a43a8-a092-47cd-8fa6-5eecde1d9bf7 · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8e99f216-4cb2-4585-a3ea-18661cbc1fa2 · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5d9cf26-1b55-47ec-b814-6e278a268e92 · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Paloma: A Benchmark for Evaluating Language Model Fit
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64bb2aec-abbc-4d16-aafe-645de5f60d18 · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0afb2e2b-3515-4df5-8662-a91f2621106d · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 736de624-1040-4b06-a353-90424ceda43e · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f894a601-7ebb-4e98-9e4d-ad84dd5733de · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning In-context Learning and Induction Heads
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb5891ae-b06e-403b-9648-dfb55cf572dd · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning The AdEMAMix Optimizer: Better, Faster, Older
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9da32be0-03f4-4c1a-b0d1-cb583aa553db · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 77df1296-b4a4-4f56-962b-3c7e4e1842d9 · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 07729239-587b-4a1a-ae67-0d987caa8ced · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Mixture-of-Depths: Dynamically allocating compute in transformer-based language models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d2cf27f-cac3-4065-9290-b35030a91c69 · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Outliers with Opposing Signals Have an Outsized Effect on Neural Network Optimization
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6162da8a-0c80-4923-bf02-e979cd7c508b · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0acf9fcd-435d-40de-8a1c-bb12889ed4e6 · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e74392e0-56ae-4ef1-9ee2-b5fde6cddedc · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b8ec33d-d619-4193-a0c7-3afc4d71caad · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Gemma 2: Improving Open Language Models at a Practical Size
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b1f55d0-2325-4943-aadd-c42eb35e4330 · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Scaling Law with Learning Rate Annealing
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ef2f770-b496-4cd6-ae99-715b7f8e8c68 · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning The Shape of Learning Curves: a Review
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22a06aaa-38ec-4f17-98a5-881b1a0aa397 · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a13dbdef-d570-432a-b52a-5c9b2282293a · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Data-Dependence of Plateau Phenomenon in Learning with Neural Network --- Statistical Mechanical Analysis
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 380618aa-20e5-4ba1-9f0c-1ad070b497d9 · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Opacus: User-Friendly Differential Privacy Library in PyTorch
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62f9b0f3-c2ac-4530-842f-51ca8ffb82cf · outbound
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Gradient Surgery for Multi-Task Learning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21d0767b-bbd7-44ae-b12e-a8508b3d6055 · inbound
Bridging Compute- and Data-Optimal Pretraining Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.