Pith. sign in

Paper Citation Record · LEDGER

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer

As of 9 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 4 inbound Pith citation observations for arXiv:2502.02531.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02531 v3

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:54:36.999715Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-22T10:25:54.649302Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T10:26:24.100616Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5eb35ad6-997a-48be-a296-59725bf22aec · outbound

This paper cites write newline.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T11:54:36.757492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:54:36.757492Z digest=sha256:c73d9e42256852afc1927e825a4804c45e5b134c2fcf231807a5f77390e95d6f

Observation bf6a0a3e-984e-4ad9-bb98-5825e86a2882 · outbound

This paper cites GPT-4 Technical Report.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T11:54:36.763726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:54:36.763726Z digest=sha256:16632a41cc5e7ec41c4e43d3f74bf801d5f8487ea816686723306275dabfd195

Observation 18fab8f8-e431-488d-8ab6-661036f0506d · outbound

This paper cites and Pennington, J.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer and Pennington, J

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:54:37.912140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:54:36.768869Z digest=sha256:29622e554ad6c30dbbebd955a5b8cf6a539327e36b69e3c47a35a37df08d3989

Observation 3974edb7-3f78-4252-871e-453a708d4599 · outbound

This paper cites and Pennington, J.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer and Pennington, J

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:54:37.898058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:54:36.773528Z digest=sha256:331dfbd69f2be28e4007b0b5b7824213149b1c124c46411c9d51b171ff7f5beb

Observation 4491be3e-df65-413b-b1cb-83b4a084524b · outbound

This paper cites S., Saxe, A.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer S., Saxe, A

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T11:54:36.778304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:54:36.778304Z digest=sha256:826ba9acd206fef9785fcd1a57e2757f357d9781b32998e252e44c1e6688a4ba

Observation 888c9cdf-5189-4f96-88ea-3fe798a179af · outbound

This paper cites Out-of-equilibrium dynamical mean-field equations for the perceptron model.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer Out-of-equilibrium dynamical mean-field equations for the perceptron model

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:54:37.874794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:54:36.782921Z digest=sha256:693dd93a6cd2d5e33db4ea7e091f4bb6b519dd43d63fe80cee0321c365340c8d

Observation fb4eb0b4-15c8-457a-98f3-eb03bae36086 · outbound

This paper cites Local kernel renormalization as a mechanism for feature learning in overparametrized convolutional neural networks.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer Local kernel renormalization as a mechanism for feature learning in overparametrized convolutional neural networks

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:54:37.860923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:54:36.787567Z digest=sha256:c793c46abc30c2381522a5753996a77f966ac6edfacdfbf04a58b2e0ef77fa42

Observation f6d688ed-e99c-49e9-aa59-e53305ba142e · outbound

This paper cites Neural networks as kernel learners: The silent alignment effect.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer Neural networks as kernel learners: The silent alignment effect

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:54:37.846327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:54:36.792699Z digest=sha256:12d775635001a6f65af5faa4827c4c4e1965ade462317b002958191a582f969f

Observation e355e0f3-a00b-4864-8445-674f9d001585 · outbound

This paper cites Predictive power of a bayesian effective action for fully connected one hidden layer neural networks in the proportional limit.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer Predictive power of a bayesian effective action for fully connected one hidden layer neural networks in the proportional limit

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:54:37.832142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:54:36.797358Z digest=sha256:1f7363a2aaf7fdea0bea3206951b7041a07a435e52712c26a1cdd0f5b2cc9a21

Observation cf1dc6c6-0932-4ef7-9b4d-a0dbd131e1bf · outbound

This paper cites Feature learning in finite-width Bayesian deep linear networks with multiple outputs and convolutional layers.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer Feature learning in finite-width Bayesian deep linear networks with multiple outputs and convolutional layers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T11:54:36.802045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:54:36.802045Z digest=sha256:2bbd8217e180aa534ea51bb4396cf3473a8e76ac81743a48beeac9edde86fd20

Observation c09b42b8-6058-441a-a813-1c79c1436f24 · outbound

This paper cites Self-Consistent Dynamical Field Theory of Kernel Evolution in Wide Neural Networks.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer Self-Consistent Dynamical Field Theory of Kernel Evolution in Wide Neural Networks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T11:54:36.807310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:54:36.807310Z digest=sha256:0aa2670b8d775c25872ebadb6239a602256906b8662e1ae20574225d03603fa8

Observation 91f53ded-6e7d-4d80-bab1-cb29eff73b45 · outbound

This paper cites Dynamics of Finite Width Kernel and Prediction Fluctuations in Mean Field Neural Networks.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer Dynamics of Finite Width Kernel and Prediction Fluctuations in Mean Field Neural Networks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T11:54:36.812224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:54:36.812224Z digest=sha256:8e65858801a963eecc6c5700cb97e9e4e6463cf47132c3d377d655be153e9225

Observation f1aa050a-3461-4a3a-a485-ffb0c08bc2e9 · outbound

This paper cites A Dynamical Model of Neural Scaling Laws.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer A Dynamical Model of Neural Scaling Laws

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T11:54:36.817140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:54:36.817140Z digest=sha256:f7d85fcda8e15a918bdc78c5bd43d243616c23434c49ce787c90cea851914141

Observation e0677a36-301e-4be4-a930-f0e5157640d7 · outbound

This paper cites How Feature Learning Can Improve Neural Scaling Laws.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer How Feature Learning Can Improve Neural Scaling Laws

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T11:54:36.822209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:54:36.822209Z digest=sha256:14794d76f24a04d7476b1860010e11e7e16dbc505ffa8a975afbbb78440bdd89

Observation 8dd11ef5-724c-4372-ac3a-392f2a3b2628 · outbound

This paper cites B., Hanin, B., and Pehlevan, C.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer B., Hanin, B., and Pehlevan, C

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:54:37.817724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:54:36.827297Z digest=sha256:c32831aa818e4c0313fd34df82271898e992576e08ee997a1b4943b5506bfe5b

Observation 30be389f-6e5d-497e-a1f4-e8621dcec3b2 · outbound

This paper cites Dimension free ridge regression.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer Dimension free ridge regression

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T11:54:36.831994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:54:36.831994Z digest=sha256:299182731f78b07dac7d525f76806a70c5d9650941b9a917c08caf961e33cac1

Observation 39b2a6c4-6922-4c15-af54-acdc5d92246c · outbound

This paper cites Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T11:54:36.836741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:54:36.836741Z digest=sha256:ae66adf5e1582b48d0b6aef9e41ef48d0031c1b084fb0c18a90483d5d09ab8d5

Observation 6cb14acb-408e-4921-b419-20817423b660 · outbound

This paper cites and Sompolinsky, H.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer and Sompolinsky, H

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:54:37.803272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:54:36.841551Z digest=sha256:e84d2eee9697cd00606a22f8258b9f4c83c4fc757811bd81539f149547d80dcf

Observation cb4c6b9c-4353-4c3d-839f-4a67beb73df9 · outbound

This paper cites Bayes-optimal learning of deep random networks of extensive-width.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer Bayes-optimal learning of deep random networks of extensive-width

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:54:37.788026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:54:36.846792Z digest=sha256:25fa98bdb2f22258b7601abe56c2f1d9e8d26bd88e43c5f152f84d3c18ea2a50

Observation bb4af645-3e8a-4f5f-9739-ba494ca1c375 · outbound

This paper cites Error scaling laws for kernel classification under source and capacity conditions.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer Error scaling laws for kernel classification under source and capacity conditions

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:54:37.773047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:54:36.851509Z digest=sha256:1f5efbc46f37a71d40739ab32cc2cd7ace4cc9f471293b4cff363da3e21dd038

Observation e05c33be-1290-497a-924c-41480ba3a4fd · outbound

This paper cites How two-layer neural networks learn, one (giant) step at a time.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer How two-layer neural networks learn, one (giant) step at a time

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:54:37.758227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:54:36.856004Z digest=sha256:1dd4d10723bf957dbf0caaead5c4861fcdba2d3a2442b2b309b7d2cf62523ac2

Observation 487aad8a-f5b0-4bc3-aceb-248014afba66 · outbound

This paper cites From Lazy to Rich: Exact Learning Dynamics in Deep Linear Networks.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer From Lazy to Rich: Exact Learning Dynamics in Deep Linear Networks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T11:54:36.860380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:54:36.860380Z digest=sha256:209a9c2c81db3d89d8bde093cab5bce3e8b3f012e26ae9697e0bdff3ca7a8864

Observation b1a0c616-1509-45ba-b17c-5370de4eaec0 · outbound

This paper cites and Gur-Ari, G.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer and Gur-Ari, G

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:54:37.744328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:54:36.865056Z digest=sha256:2c97a22c2f6b3cbabb8b615aa934a0d69ab001501372eae2fa2f5003a6249010

Observation 4b4a5173-09bc-4304-a3f9-615fcbe86358 · outbound

This paper cites Double trouble in double descent: Bias and variance (s) in the lazy regime.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer Double trouble in double descent: Bias and variance (s) in the lazy regime

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:54:37.730109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:54:36.869482Z digest=sha256:9dd7fc42749cea513d24a6eebbb29f2fdcf18539d342f1a9d1ce498d750867fb

Observation a7e52749-50eb-49f8-8199-6388b3bf013b · outbound

This paper cites Scaling Exponents Across Parameterizations and Optimizers.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer Scaling Exponents Across Parameterizations and Optimizers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T11:54:36.874008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:54:36.874008Z digest=sha256:d2d029ee59f8e6712ff68cd940f8774e29f4f5cab52f0b68f653b6c64e772daa

Observation 51d65a98-8151-425a-9436-df7dddd7fb5e · outbound

This paper cites Rigorous dynamical mean field theory for stochastic gradient descent methods.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer Rigorous dynamical mean field theory for stochastic gradient descent methods

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T11:54:36.878775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:54:36.878775Z digest=sha256:b39ee447bfdf56dae7e79911a536e81f1454e25d1d35642c964eaf0347457fcc

Observation fa877548-18b0-4e8a-9a07-1b3a8ba427f6 · outbound

This paper cites Finite Depth and Width Corrections to the Neural Tangent Kernel.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer Finite Depth and Width Corrections to the Neural Tangent Kernel

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T11:54:36.883247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:54:36.883247Z digest=sha256:7c9d17594d36e2c23bbca243bc3640601cca2a469f09b4b6911431621654fa5f

Observation 69b556e7-2555-4f97-a301-e677cd4c1b19 · outbound

This paper cites Bayesian Inference with Deep Weakly Nonlinear Networks.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer Bayesian Inference with Deep Weakly Nonlinear Networks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T11:54:36.887888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:54:36.887888Z digest=sha256:7d15f1615fbe532b675e50e66a3d04f9153353ff91eb299def7d427956dea1ed

Observation b4495b4e-ae97-43c0-9450-55b568f14903 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer Training Compute-Optimal Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T11:54:36.892600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:54:36.892600Z digest=sha256:008a50003fdd99bb787eac22f1825dfc681cbf0d9d2774743369c7b1d805d458

Observation 9be7a1d6-12c7-4620-9ff0-b1498b2a381a · outbound

This paper cites and Lu, Y.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer and Lu, Y

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T11:54:36.897035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:54:36.897035Z digest=sha256:f3f9d4e8e06c14a72842ba7b059c59acbfdcc9f7fb8af01da9bd85a59a5cffa5

Observation cb006b81-e733-46e9-a44b-8c727f47f8ea · outbound

This paper cites Neural tangent kernel: Convergence and generalization in neural networks.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer Neural tangent kernel: Convergence and generalization in neural networks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T11:54:36.901447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:54:36.901447Z digest=sha256:c8f3864b5b8e74ecd7c0da210f19a94f9990bcb5bad45b0af3eb70f5f36bf430

Observation ff950a05-9e44-4e4d-b021-616f3aefd125 · outbound

This paper cites Get rich quick: exact solutions reveal how unbalanced initializations promote rapid feature learning.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer Get rich quick: exact solutions reveal how unbalanced initializations promote rapid feature learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T11:54:36.905939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:54:36.905939Z digest=sha256:cc194ac7bea16dda6695f22a95892c9079d8d9d57e08933ac9dd2867cba23b39

Observation cb9b95c5-7286-4f86-b37f-a1be6128f2ca · outbound

This paper cites Wide neural networks of any depth evolve as linear models under gradient descent.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer Wide neural networks of any depth evolve as linear models under gradient descent

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:54:37.697018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:54:36.910661Z digest=sha256:cd60f861cfd4ab3e59568a914c7c82d9f39289e93e6ac16c1af3407cdf4a31ae

Observation fef3fdc1-d76b-49cd-b92e-d5cd1c7143cf · outbound

This paper cites and Sompolinsky, H.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer and Sompolinsky, H

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:54:37.682084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:54:36.914950Z digest=sha256:43ce5ad5ffb24508f3a071f0300946d7b0313e60569ab4f9ee8801b2e45cc2b0

Observation e533b552-3d57-4592-8a3a-c2b56294f779 · outbound

This paper cites S., Krzakala, F., Urbani, P., and Zdeborova, L.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer S., Krzakala, F., Urbani, P., and Zdeborova, L

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:54:37.668275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:54:36.919161Z digest=sha256:65752aec0979bf2a180c94f67b3f22f5611df8aca55bb41fbe596d87608ab002

Observation c410c5fe-855b-48e4-bb62-93685066eb7d · outbound

This paper cites C., Siggia, E., and Rose, H.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer C., Siggia, E., and Rose, H

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:54:37.654605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:54:36.923517Z digest=sha256:709528a7177998c2032be816fb82bafc8740cafb4ffb13cca88d2582d07fe53d

Observation 4074d9ba-0652-4f7f-8a0c-f8163ccd94a5 · outbound

This paper cites and Montanari, A.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer and Montanari, A

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:54:37.640026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:54:36.927780Z digest=sha256:f427e7581bfeaecf3eef149f4c9a502914eed19b57bca6c345fc97b3b6b6265d

Observation 0ad92422-bac8-4d8b-9cbb-5d8e2e0527ca · outbound

This paper cites Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:54:37.623786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:54:36.932213Z digest=sha256:250c456b40c0a8429175059ca7139e2087e64fc0b3af16f06d57ed2c05e89b12

Observation c8e199a5-0987-45b1-beaa-7d09c68b14ed · outbound

This paper cites and Urbani, P.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer and Urbani, P

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:54:37.608874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:54:36.936621Z digest=sha256:62e664b77b055fa464438ea4ae98cf1d59f6a18b419705e0e2a376063ef6dab4

Observation 0d56aadf-0172-48a3-bf55-453600e41fc9 · outbound

This paper cites Dynamical mean-field theory for stochastic gradient descent in gaussian mixture classification.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer Dynamical mean-field theory for stochastic gradient descent in gaussian mixture classification

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:54:37.594154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:54:36.940998Z digest=sha256:a21c56d4daed4f5f0b5d5154de090c2d7f7a44946b6e6c0292ebe2978ca3728a

Observation 289b2551-a682-4c73-9c82-303688f13aa2 · outbound

This paper cites Super Consistency of Neural Network Landscapes and Learning Rate Transfer.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer Super Consistency of Neural Network Landscapes and Learning Rate Transfer

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-08-09T11:54:37.219680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:54:36.945315Z digest=sha256:7f41361e81c43827f3b96b1b3e71403dd39ed29c455a372415e9b48a1e2c3e40

Observation 909e0d77-9f8e-4372-ab69-3498556f84ce · outbound

This paper cites A statistical mechanics framework for bayesian deep neural networks beyond the infinite-width limit.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer A statistical mechanics framework for bayesian deep neural networks beyond the infinite-width limit

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:54:37.578804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:54:36.949713Z digest=sha256:6dfe416f610da808ebf0b155fdf30bb16114a2b9648131a18fb6cd0bf58fbc0a

Observation 0b9f5d02-a367-447e-8c95-4e115429eca3 · outbound

This paper cites 4+3 Phases of Compute-Optimal Neural Scaling Laws.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer 4+3 Phases of Compute-Optimal Neural Scaling Laws

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T11:54:36.954538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:54:36.954538Z digest=sha256:289a9272110429d74b6512323645edd4f9fec2fea629db207b671fe8f18827b5

Observation fa9a8356-769a-492d-b60f-d58cd312f38b · outbound

This paper cites an unresolved cited work.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T11:54:36.958962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:54:36.958962Z digest=sha256:6d3547347f6d03ab04a5108610eb26883c0b46353d91523e33f8fd575e0272e3

Observation dd3a2e97-1875-44d4-8a2b-56b426e8e51d · outbound

This paper cites A., Yaida, S., and Hanin, B.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer A., Yaida, S., and Hanin, B

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:54:37.563851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:54:36.963727Z digest=sha256:a0f3928ac9e81b5c4c1949b8a62326600847486ba3b7579618fc3e648fed2413

Observation 5bc77b0f-2c30-4669-8522-aa2f2e219d45 · outbound

This paper cites and Vanden-Eijnden, E.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer and Vanden-Eijnden, E

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T11:54:36.968052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:54:36.968052Z digest=sha256:20b9d85ce84332dc9c4e8cf8978de8bd1d33acbfa6723a493c0b73016bfd480e

Observation 4fce7627-0883-44ff-95e8-ef9cc212008f · outbound

This paper cites Exact solutions to the nonlinear dynamics of learning in deep linear neural networks.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer Exact solutions to the nonlinear dynamics of learning in deep linear neural networks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T11:54:36.972330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:54:36.972330Z digest=sha256:178aa6a1d49dd0aaa0db5fbd794d11888a0d07c480f92f3b617505bacffb2d4c

Observation 53bc0647-54a0-4ec2-86e9-2f2472422dc4 · outbound

This paper cites C., Xu, J., Haas, M., and Cevher, V.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer C., Xu, J., Haas, M., and Cevher, V

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:54:37.540836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:54:36.976896Z digest=sha256:6553ae347584067a847eb7441ecf957b2ea34d689de2b236c43e10710fe674f5

Observation 2cfc7ce2-5863-46df-a64d-8d1d718484a9 · outbound

This paper cites and Hu, E.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer and Hu, E

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T11:54:36.981463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:54:36.981463Z digest=sha256:57af982b173a63c5300574aa0d8687d2c1f6412d7a07910dff20f991916d7f05

Observation a8bf1cc5-f80f-4f56-bf33-f97752193d7b · outbound

This paper cites J., Babuschkin, I., Sidor, S., Liu, X., Farhi, D., Ryder, N., Pachocki, J., Chen, W., and Gao, J.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer J., Babuschkin, I., Sidor, S., Liu, X., Farhi, D., Ryder, N., Pachocki, J., Chen, W., and Gao, J

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-09T11:54:36.985871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:54:36.985871Z digest=sha256:0692cc4ac23d95d8ab2b430351d06947f99e3a954b94101635204fc8a2abebb5

Observation b7c6b5ed-b662-47d4-aeac-825a38105945 · outbound

This paper cites Tensor Programs VI: Feature Learning in Infinite-Depth Neural Networks.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer Tensor Programs VI: Feature Learning in Infinite-Depth Neural Networks

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T11:54:36.990546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:54:36.990546Z digest=sha256:85e0d6fdd99be442033eefc11bd2e58060e67c83b41eefe7d42780ff37eb0261

Observation 53fe2196-153e-4f4f-9a62-ad0481871143 · outbound

This paper cites an unresolved cited work.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:54:37.508761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:54:36.995272Z digest=sha256:bd3259146cb559fb08c2c4b519b29b6cb0ac9aa01d7834a7473cc67731a7a703

Observation d741f9d7-7f6f-4a17-aa21-e2aafadbe8bd · outbound

This paper cites A., Tong, W.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer A., Tong, W

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:54:37.494758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:54:36.999715Z digest=sha256:d351b944c6bbf6a5284f239b8ed167f8b816923913d076af66b6011d8c6e5e44

Pith citing papers

Observation 3f788870-ad61-4919-b3a2-181940488d87 · inbound

There Will Be a Scientific Theory of Deep Learning cites this paper.

There Will Be a Scientific Theory of Deep Learning Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:21:08.980824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T20:11:17.616190Z digest=sha256:1fc727b309c848734e69fdc5d37de855167623dfa01c7d685ae25ca3f6659804

Observation 64686c53-6c31-48df-85d3-0635d4454f8d · inbound

Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer cites this paper.

Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:05:53.487859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T03:02:52.833353Z digest=sha256:0d912f03f33c57c7af00c4e942fe21df132c6410d350112675cb32192241aa1e

Observation cc0eec55-8df5-4051-a7eb-4bd5f1761e89 · inbound

Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer cites this paper.

Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:26:24.104396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T10:25:54.649302Z digest=sha256:b10a5556be733c2100b982217271e97699aaa9c0821913ad0e34333c1f901a58

Observation ac9810ea-8128-4c8c-b11f-5b109f14e692 · inbound

How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization cites this paper.

How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T04:49:44.858804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-15T04:45:20.091598Z digest=sha256:d71046987bd52eb7aead0b2deefc8f2b5867d104f0a694b81ea590450815a042