Pith. sign in

Paper Citation Record · LEDGER

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks

As of 9 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 1 inbound Pith citation observation for arXiv:2507.02119.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02119 v2

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:48:55.792391Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-16T06:58:38.927268Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T07:00:43.255075Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact2
  • verified fuzzy23
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1f8507b7-d662-4d9a-a6f9-a37c495b3814 · outbound

This paper cites GPT-4 Technical Report.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:48.890013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:48.890013Z digest=sha256:85f56761ae898117a4d4faabb77698dc9cf3b8da788b6eafe7e1033bc5473062

Observation ff16d3b2-6874-4029-bc2a-03c5d1a16f45 · outbound

This paper cites and Fisher, D.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks and Fisher, D

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:49:03.332256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:48:48.984871Z digest=sha256:1da5c25e24f3d796babf690e94bf2f1bf51e827dc12ab9cff98d5cc52c608f5c

Observation d210d09d-bd03-4756-84ab-8877e7587d6b · outbound

This paper cites High dimensional analysis reveals conservative sharpening and a stochastic edge of stability.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks High dimensional analysis reveals conservative sharpening and a stochastic edge of stability

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:49.116497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:49.116497Z digest=sha256:c337fef4b14dceeca227c815b0e9af656f01d8b6c02cfe7ac28b1d332274d40e

Observation 803ba54a-c036-4b78-a89f-9aacfaba5bff · outbound

This paper cites Power lines: Scaling laws for weight decay and batch size in llm pre-training.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Power lines: Scaling laws for weight decay and batch size in llm pre-training

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:49.230831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:49.230831Z digest=sha256:876472dbd5d5c36ac80e84f10d9e2ae12aeefeeb3c5c541398cea42f20b24059

Observation 19d23896-26c3-4357-a7ee-da726da3d950 · outbound

This paper cites Finite size scaling analysis of ising model block distribution functions.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Finite size scaling analysis of ising model block distribution functions

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:49:03.061062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:48:49.341620Z digest=sha256:6cfbdd8e98df210206639707f7a0b58f1d2b2d198b5564595d25fb2bd17466ff

Observation 00ddcf25-616c-4d98-97c1-562523c70530 · outbound

This paper cites and Pehlevan, C.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks and Pehlevan, C

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:49:02.877934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:48:49.465906Z digest=sha256:442efbd6418744651a7630f57331a12fa8862f488126888a5512738b7861337c

Observation 7c5eee10-f120-4418-be25-ed004d944912 · outbound

This paper cites Depthwise Hyperparameter Transfer in Residual Networks: Dynamics and Scaling Limit.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Depthwise Hyperparameter Transfer in Residual Networks: Dynamics and Scaling Limit

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:49.634542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:49.634542Z digest=sha256:9ad446e1e944f70032eede973032fb3cfe02ff430bb2667f6e127f4ba4540cd9

Observation 34cf55a9-2ec1-4ee3-979d-31fb7e50ae16 · outbound

This paper cites A Dynamical Model of Neural Scaling Laws.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks A Dynamical Model of Neural Scaling Laws

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:49.730886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:49.730886Z digest=sha256:74a7646daf321351ce1c7758c43daeddf2933696fd6978510f6173bed13b853b

Observation 14832c06-88b9-461a-8b12-2d859f17a77f · outbound

This paper cites How Feature Learning Can Improve Neural Scaling Laws.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks How Feature Learning Can Improve Neural Scaling Laws

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:49.856090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:49.856090Z digest=sha256:b660ea5745baf47873a6437da8e4f03fe5f8c8973943f691175c009968370c08

Observation 03d58112-1883-4f77-943c-d1f8987069cd · outbound

This paper cites Infinite limits of multi-head transformer dynamics.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Infinite limits of multi-head transformer dynamics

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:49:02.507256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:48:49.979212Z digest=sha256:bbb49618bf5a800028449023f42d139ec45254f3a780738a588e951cb000a196

Observation b2a49a53-e9c1-4dc0-bdee-154cb355f9c6 · outbound

This paper cites Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:50.108888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:50.108888Z digest=sha256:5cf432e91670bb0c3550cc267efa9849b7f01809019b5af348e51c98738b5b09

Observation 4d05b138-7864-4521-a851-7c0cf4d60ba8 · outbound

This paper cites Adaptive Gradient Methods at the Edge of Stability.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Adaptive Gradient Methods at the Edge of Stability

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:50.237889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:50.237889Z digest=sha256:80c688a7dda49fa4fb06076bf90d44f97e93cd23a9ca3c4fae71354729c4dbab

Observation 6e76a5f3-8709-4694-8311-a01ccecbc2c8 · outbound

This paper cites M., Damian, A., Talwalkar, A., Kolter, Z., and Lee, J.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks M., Damian, A., Talwalkar, A., Kolter, Z., and Lee, J

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:50.344017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:50.344017Z digest=sha256:4cbcdecdbcaa38bd32dfc6f69f25fbd215bfb298ddcb5af6329583fff7bd8717

Observation 80156d54-db79-46f4-8069-df4302dbccae · outbound

This paper cites Optimal learning rate schedules in high-dimensional non-convex optimization problems.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Optimal learning rate schedules in high-dimensional non-convex optimization problems

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:50.480589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:50.480589Z digest=sha256:d4fbdbe565fa55633390ee3c2f09d6c3fe698563cb98a9610a23663660a293be

Observation 9502e4e6-1a62-40ee-8cb0-5ed9eaf6c731 · outbound

This paper cites C., Noci, L., Li, M., Bordelon, B., Bergsma, S., Pehlevan, C., Hanin, B., and Hestness, J.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks C., Noci, L., Li, M., Bordelon, B., Bergsma, S., Pehlevan, C., Hanin, B., and Hestness, J

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:50.592169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:50.592169Z digest=sha256:efe9bdd68ac4fbfc7d515cba2424d01f2a1fb8e19eac2ce6bb8db672017f0cfe

Observation ad86be4a-4417-4076-9c0f-e4185cd2952a · outbound

This paper cites Kingma, J.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Kingma, J

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:49:02.206565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:48:50.697194Z digest=sha256:467c1b87e801973359d037630130acdea35fddce1c981f1c83a1ed7f47829a96

Observation 0231a66f-5898-4deb-a78f-3976de64ab0d · outbound

This paper cites Scaling Exponents Across Parameterizations and Optimizers.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Scaling Exponents Across Parameterizations and Optimizers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:50.834417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:50.834417Z digest=sha256:25070917fc19fa832a5ba1ea3527f462d0de384d05ebf656354c9c875b5b2bf3

Observation a27cd3c4-c959-4c6c-ade8-6651aebda303 · outbound

This paper cites an unresolved cited work.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:49:01.969127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:48:50.982931Z digest=sha256:e3885df1fe1eb3edc532a3951ca71148fe7292c39a48f2c37ca3853e0d7f1887

Observation 9b472174-aad9-4b29-a6f6-280c995a9936 · outbound

This paper cites Monte Carlo methods in financial engineering, volume 53.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Monte Carlo methods in financial engineering, volume 53

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:49:01.602856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:48:51.135576Z digest=sha256:bf777b3ec8b88075abcd87fcc7cc9340faa1589087470a4615081fc99343e1d1

Observation 894f4596-0b96-4b61-a659-2679eddf39c2 · outbound

This paper cites and Fisher, D.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks and Fisher, D

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:49:01.275473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:48:51.288401Z digest=sha256:f5813a45837fe0da7eff04ba53007d8c9cbc74bf7d2eb103c640658db536a0b0

Observation 01e15fe3-4ed2-492b-8957-bad8c1daf31e · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Gaussian Error Linear Units (GELUs)

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:51.439168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:51.439168Z digest=sha256:6c872d1f2bc9fb1abb9352edfaacc4409958da7490bb36855705b13264f6b217

Observation 4e936d7a-dbe8-4e16-ba21-53b8d78e44a6 · outbound

This paper cites Deep Learning Scaling is Predictable, Empirically.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Deep Learning Scaling is Predictable, Empirically

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:51.541070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:51.541070Z digest=sha256:5986260e12a51acb7c215e27dbeaf85a92fc95ebc7a2fdbbe18c35ac741350e5

Observation a53d9abd-f9f5-46f0-9a61-462a98c5354a · outbound

This paper cites Training Compute-Optimal Large Language Models.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Training Compute-Optimal Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:51.686276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:51.686276Z digest=sha256:8ba457f9638a15922120bbb9d9a604418e440407621867cc2ff7db557b99bde8

Observation 6c3c0699-c49f-492c-a582-7a53930cd057 · outbound

This paper cites Scaling Laws for Neural Language Models.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Scaling Laws for Neural Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:51.813397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:51.813397Z digest=sha256:5dc9b0fa2ab8551aee29e0dea1b726fb135c321d7b6174d3d12b0a8294b26fb4

Observation ff42298c-b9e7-4d0e-9192-5040dfe2c84d · outbound

This paper cites F., Blundell, J.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks F., Blundell, J

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:49:00.969218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:48:51.957254Z digest=sha256:cb5db8c1d7bb50886648400a3297f91706f5a355ec584774cef409b77e4e152e

Observation 285b1657-f9f6-4e76-9a21-83e8c78a68f3 · outbound

This paper cites Stochastic modified equations and adaptive stochastic gradient algorithms.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Stochastic modified equations and adaptive stochastic gradient algorithms

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:49:00.702563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:48:52.040916Z digest=sha256:ea9af8d6863f9a6b7934ca61161198b4ae26cb0b2c54f63f37770c8a06b6f6ad

Observation 1541ff55-00bc-4e67-beb4-0ec9ee48de10 · outbound

This paper cites A Multi-Power Law for Loss Curve Prediction Across Learning Rate Schedules.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks A Multi-Power Law for Loss Curve Prediction Across Learning Rate Schedules

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:48:56.444458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:48:52.152237Z digest=sha256:d4646d5e573391b279325aac7c17d84c58cd7e120a6a25fde322dda0bbf25477

Observation aaa4604b-ec14-4b6c-be1a-efdd3c1a479d · outbound

This paper cites On the sdes and scaling rules for adaptive gradient algorithms.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks On the sdes and scaling rules for adaptive gradient algorithms

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:49:00.420054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:48:52.302674Z digest=sha256:97ac8544ace47b24c2ded195de8a0741017fd286a26b1ed337bf1377c3989f8a

Observation f77b1822-0677-4628-920b-ba8297a7df0a · outbound

This paper cites and Pastur, L.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks and Pastur, L

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:49:00.239937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:48:52.402249Z digest=sha256:18397aac2ad0f8fa2721d54c6d4ad320b5b9781bb47adff45c7389b0d580db58

Observation 3593f90c-ab38-4294-a77a-dbf7cb00324e · outbound

This paper cites An Empirical Model of Large-Batch Training.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks An Empirical Model of Large-Batch Training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:52.527254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:52.527254Z digest=sha256:fbf180ada17cdd8eb7c2b6219dd0c1ce747dacbeb7bb6bd066347639fe64e97a

Observation 9d7c7f74-d112-40fa-87e1-a8fce35df8b4 · outbound

This paper cites Y., Singh, S., Bhatele, A., Goldblum, M., Panda, A., and Goldstein, T.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Y., Singh, S., Bhatele, A., Goldblum, M., Panda, A., and Goldstein, T

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:52.668462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:52.668462Z digest=sha256:928278fbbf826621cc4bf8264ab15fedaa13f48f1764e4349087ea0c52bc424a

Observation 04774ee9-505a-45a6-9571-c86f9dea0c0d · outbound

This paper cites The Deep Bootstrap Framework: Good Online Learners are Good Offline Generalizers.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks The Deep Bootstrap Framework: Good Online Learners are Good Offline Generalizers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:52.826035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:52.826035Z digest=sha256:e2079cb9b782a97d710593daf54bfb370c81231edff6ab2f1ab55564bfbd34b1

Observation 32a46c09-50ef-497d-b2dd-898847fea43b · outbound

This paper cites Super consistency of neural network landscapes and learning rate transfer.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Super consistency of neural network landscapes and learning rate transfer

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:48:59.986705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:48:52.964417Z digest=sha256:195af58ec4391b086dbad8e76c786423ff9b7c5ff5388da498a97f3c77684b6a

Observation f53c09ca-dfd3-44cb-a19f-41b65fc5e507 · outbound

This paper cites Sgd in the large: Average-case analysis, asymptotics, and stepsize criticality.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Sgd in the large: Average-case analysis, asymptotics, and stepsize criticality

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:48:59.812856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:48:53.062783Z digest=sha256:ef0a57a3fd00ce187e65eed85db1e2b2fb19b60ae7ae1781d09f8f98e5b774f4

Observation 6a579c77-06eb-48d0-a7fb-542382b48a89 · outbound

This paper cites Homogenization of sgd in high-dimensions: Exact dynamics and generalization properties.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Homogenization of sgd in high-dimensions: Exact dynamics and generalization properties

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:48:59.564827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:48:53.171825Z digest=sha256:ff9bcd3a48a8ab268e96fa7326df39f203f44c75739b8bc82124214c4becea83

Observation 17e6d158-514a-450d-ad6c-12a4ecc631b6 · outbound

This paper cites 4+3 Phases of Compute-Optimal Neural Scaling Laws.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks 4+3 Phases of Compute-Optimal Neural Scaling Laws

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:53.293116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:53.293116Z digest=sha256:89bdfc17237c5e63e6512f6143afaa515a4b74af5b53ef6963d0061bbb56f0ec

Observation ae4e746d-a68b-4e00-a98f-f1a8a904de60 · outbound

This paper cites T., Agarwala, A., and Fisher, D.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks T., Agarwala, A., and Fisher, D

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:48:59.330413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:48:53.414115Z digest=sha256:7e724b69f634928a0c52e23a5629664cc3cdcdca5ffc27c801c29d1e7389d84c

Observation 5be144fc-b5ef-41b9-bb18-d0db1e2f5d94 · outbound

This paper cites Reconciling Kaplan and Chinchilla Scaling Laws.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Reconciling Kaplan and Chinchilla Scaling Laws

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:53.523148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:53.523148Z digest=sha256:0321900b45c14c22f0b2b29ea81a79985f0be616d2ea60f8051c10300a599ef8

Observation 81426342-a206-4235-9622-0e17160b76f0 · outbound

This paper cites J., Davison, M., Bhaya, D., and Fisher, D.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks J., Davison, M., Bhaya, D., and Fisher, D

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:48:59.064308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:48:53.640808Z digest=sha256:bc8f2f214d88c50e0d3bec7768c88435eee815b369ddf48598a43ec167d2c677

Observation 524ba4f5-5eea-4805-bd48-9394a5ddabe6 · outbound

This paper cites The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:53.764373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:53.764373Z digest=sha256:a8b5d148c994d4a1c9468c91d63c494e033f35927773dbff9f2756a94fa019d3

Observation 671b23b1-1324-45fa-af71-19fe0c84092b · outbound

This paper cites Ubiquitous abundance distribution of non-dominant plankton across the global ocean.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Ubiquitous abundance distribution of non-dominant plankton across the global ocean

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:48:58.772167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:48:53.834070Z digest=sha256:bd2cdba2344ccfb57d5aad43e8e9cfff3aa57b07b7f61d19d8f144b88c574e18

Observation fa376d3b-72b2-4d46-8bf4-f6869b571edc · outbound

This paper cites and Kaplan, J.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks and Kaplan, J

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:48:58.494464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:48:53.926114Z digest=sha256:0ee70f84c4d3f48781a887e1ad705417b91efffdb47c02a07d3d4e8bb41ffb78

Observation 4091f172-60bd-41e7-a799-b1009cb2f517 · outbound

This paper cites Universal Scaling Laws of Absorbing Phase Transitions in Artificial Deep Neural Networks.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Universal Scaling Laws of Absorbing Phase Transitions in Artificial Deep Neural Networks

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:48:56.060606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:48:54.043892Z digest=sha256:a5ffc3eb52d1d47f505dd2b03cab0b56fb102c6a47adb1e90d2aad64cb01d71b

Observation 0d53d82b-059b-4a33-be5a-6c88d9f6f97d · outbound

This paper cites Scaling Law with Learning Rate Annealing.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Scaling Law with Learning Rate Annealing

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:54.159538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:54.159538Z digest=sha256:41c6256964bce811617cb5278bd842291f8e71af8e765cfd021f9000875e074d

Observation 8b73fa26-db25-4502-92c4-30c46691e5dc · outbound

This paper cites R., Geiler-Samerotte, K., H \'e rissant, L., Blundell, J.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks R., Geiler-Samerotte, K., H \'e rissant, L., Blundell, J

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:48:58.140034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:48:54.271428Z digest=sha256:7bc4e5ba2f9ae5e1afaade2801b4b1ca98b5634b8b9410b973b1f2996d3fb586

Observation dbb367e9-4461-4699-9c01-b99aa9f5d739 · outbound

This paper cites Feature-learning networks are consistent across widths at realistic scales.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Feature-learning networks are consistent across widths at realistic scales

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:48:57.822552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:48:54.396692Z digest=sha256:0fbf7c9cff91c0939214e494c8c31e0666d34a55872eb55e29eec42c941abb42

Observation 16c13459-0ce0-4da9-8d2f-dab0399daa4c · outbound

This paper cites How to set AdamW's weight decay as you scale model and dataset size.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks How to set AdamW's weight decay as you scale model and dataset size

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:54.538611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:54.538611Z digest=sha256:5ea25857b4b61fda276541abfdcce9110b593932ad6e9aeb8f88b88fbf78c317

Observation 8e56a2c6-68f4-468a-b2d3-e557308434b4 · outbound

This paper cites Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:54.682273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:54.682273Z digest=sha256:9c2055d52fd8f7323c94ba701934995ad93e00987615e0902f932cc5a6518b70

Observation 675f7208-5acd-4a1d-bed0-13fa8463fd56 · outbound

This paper cites an unresolved cited work.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:48:57.542119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:48:54.825321Z digest=sha256:6ca41212a8074e4f19e5b1ba246d6a0b8c1dcb09334be9b19d1781287afe6ecf

Observation e268f4b4-f837-49ba-a19f-7c5fdae47cf3 · outbound

This paper cites Small-scale proxies for large-scale Transformer training instabilities.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Small-scale proxies for large-scale Transformer training instabilities

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:54.971531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:54.971531Z digest=sha256:4311ccc73ca614df50275bc2970e72dfcb9a8a7baf631414c051f29c3e0e1e35

Observation daf49d49-f471-46bf-8352-787c0b0f9991 · outbound

This paper cites Rethinking Conventional Wisdom in Machine Learning: From Generalization to Scaling.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Rethinking Conventional Wisdom in Machine Learning: From Generalization to Scaling

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:55.089882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:55.089882Z digest=sha256:f8659ba71fb1a553a7c9718bfdab530c17af16af59216e5c3257db263f797e91

Observation cf60c245-08c2-470e-b2e2-307e619503a5 · outbound

This paper cites and Hu, E.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks and Hu, E

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:48:57.192764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:48:55.214821Z digest=sha256:7c38aa37cd7fa301939447f7f2315228c7a286331db8e311e21ec26e3c351c69

Observation 6907316e-e6d2-48a5-97e1-881565c4a1ec · outbound

This paper cites and Littwin, E.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks and Littwin, E

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:48:57.010570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:48:55.349785Z digest=sha256:b169f477fe5d8a7308ccf556417a6ab339f1b90ff284bb56c72d2a60ad98708e

Observation b05f31e8-fb3c-44e1-b053-f8a9f689ea30 · outbound

This paper cites J., Babuschkin, I., Sidor, S., Liu, X., Farhi, D., Ryder, N., Pachocki, J., Chen, W., and Gao, J.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks J., Babuschkin, I., Sidor, S., Liu, X., Farhi, D., Ryder, N., Pachocki, J., Chen, W., and Gao, J

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:48:56.863551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:48:55.507002Z digest=sha256:de8d74e8132a9be261f500fa21335c4303fa39057c23a444fbcad32c5d118c81

Observation d6b96640-4afb-48c5-a528-1ee0a2da6836 · outbound

This paper cites and Sennrich, R.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks and Sennrich, R

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:55.635949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:55.635949Z digest=sha256:5d6ed5f677b022421477e05090d4b4e8834358b523c0368acf57ef7ea1ac70d1

Observation ae15619c-a965-4477-a706-d53476ec9131 · outbound

This paper cites an unresolved cited work.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:55.792391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:55.792391Z digest=sha256:4a34a84e780ad9302c33827c42098b367f7bb8d38b909bfa24136bdfaead9edf

Pith citing papers

Observation 41fd3f8c-d9fe-418e-b9d8-5b79dd298a3f · inbound

Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model cites this paper.

Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:00:43.256809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T06:58:38.927268Z digest=sha256:b591b30b8275325502323b337c8e93a1476b032730c3ae57bd22f0f14608855b