Pith. sign in

Paper Citation Record · LEDGER

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias

As of 18 August 2026, this Paper Citation Record lists 100 of 117 outbound references and 0 inbound Pith citation observations for arXiv:2605.29152.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.29152 v1

Coverage vector

measured 100 of 117 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T13:20:54.303605Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 117 outbound references displayed

  • verified exact5
  • verified fuzzy0
  • unresolved95
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2c596934-f23a-40f8-9017-b6ef815f24b2 · outbound

This paper cites Under- standing deep learning (still) requires rethinking generalization.Communications of the ACM, 64(3):107–115, 2021.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Under- standing deep learning (still) requires rethinking generalization.Communications of the ACM, 64(3):107–115, 2021

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:88cb8b7036a92f8759b738bbf6ffc75754329f3f2687306e5ab0228c10d1cae9

Observation a1a0c1f8-4739-4246-aef6-ea1ff90c7140 · outbound

This paper cites Strogatz.Nonlinear Dynamics and Chaos: With Applications to Physics, Biology, Chemistry, and Engineering.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Strogatz.Nonlinear Dynamics and Chaos: With Applications to Physics, Biology, Chemistry, and Engineering

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:aef2e2d63c35066634fc158da3b9beaf2ccdeb0095bd92ea614dfb41fdbd0f7e

Observation 365f77c3-23c8-4e91-b5c1-12d02b0c9ab9 · outbound

This paper cites Hirsch, Stephen Smale, and Robert L.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Hirsch, Stephen Smale, and Robert L

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:86b085278d8c900391e5ef15c61e3376463915bd3e5672e808a0e0d5b99d43f6

Observation a8dc0dd6-2ad0-4919-8139-114961bb7023 · outbound

This paper cites On the explicit role of initialization on the convergence and implicit bias of overparametrized linear networks.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias On the explicit role of initialization on the convergence and implicit bias of overparametrized linear networks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:b9fbfbc81059306e0ca57b2b43abb624d3233a6b1dde8ddb3d9d12c27fd8a28f

Observation 9bbd45a2-25a2-4cac-9aba-619c2991b4b8 · outbound

This paper cites On the role of initialization on the implicit bias in deep linear networks, 2024.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias On the role of initialization on the implicit bias in deep linear networks, 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:5eb334ac4a60ad840aede2eab57484040806bbb2b5eaabec1b484e43030a3356

Observation 311b77d0-a8c9-40bf-9851-b9eff4705123 · outbound

This paper cites Camargo, and Ard A.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Camargo, and Ard A

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:23a00692d92c088bdc4c6b51c653d2c08cb723242db738c900604b23bf6287df

Observation 55828dee-4045-40d1-b9c5-687c5e5ed24f · outbound

This paper cites an unresolved cited work.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:6acc95f5029857e5a4975d1067fa152f0abd345c889475d66db976ff809e08c6

Observation a3eae809-5fbf-48e1-ad15-e3f0cec5ed3c · outbound

This paper cites Deep-layered machines have a built-in occam’s razor.arXiv preprint arXiv:2603.01217, 2026.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Deep-layered machines have a built-in occam’s razor.arXiv preprint arXiv:2603.01217, 2026

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:23:28.016757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:7f0209ab6a5e097bee2e4ce1102de112dcc4442d2a8ae23f2da07f6bab4b9a8b

Observation dea8f9d5-5d6e-49f7-a8c9-85716021f7f8 · outbound

This paper cites Understanding the difficulty of training deep feedforward neural networks.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Understanding the difficulty of training deep feedforward neural networks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:0c2997a9d974d96cc9c49861972a6037d5f052ea5ffaed558a9a4c3a24bfcb4e

Observation 1eb131bb-0efb-4b6f-97a1-dec28122bf92 · outbound

This paper cites Delving deep into rectifiers: Surpassing human-level performance on imagenet classification.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Delving deep into rectifiers: Surpassing human-level performance on imagenet classification

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:08798d27571453704a2af99756c43f5481d37d7c1382eb36c81f64ca53792cac

Observation 4dcdbf97-575d-424a-a620-f8d408f6edfc · outbound

This paper cites Exponential expressivity in deep neural networks through transient chaos.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Exponential expressivity in deep neural networks through transient chaos

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:832aeb2aac55eb8748d300db066afe757481dba1f628dca9dbd85c296bbfc016

Observation 6052ce46-ae6c-4d81-9b06-e86ae353e396 · outbound

This paper cites Schoenholz, Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Schoenholz, Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:26b3dcee4745198cab3cf8e794323985ef55c9340fd19472c85cfd96538cb920

Observation 1e3f62e5-b038-43ab-a15e-e8af80274ee3 · outbound

This paper cites Schoenholz, and Surya Ganguli.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Schoenholz, and Surya Ganguli

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:b98ae6c122172b6f9687751b75b6cc8d6b5f75e982968c1aa321da275fa7c39a

Observation 835c2b17-2a4e-4936-8062-399c6210206d · outbound

This paper cites How to start training: The effect of initialization and architecture.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias How to start training: The effect of initialization and architecture

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:d5f08572ea0826cfe1201c0884d0771935a4a303195f7346889583cd3cb05b6c

Observation 1fab6d43-6efe-4589-8dd9-05457b30f321 · outbound

This paper cites Which neural net architectures give rise to exploding and vanishing gradients? InAdvances in Neural Information Processing Systems, volume 31, 2018.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Which neural net architectures give rise to exploding and vanishing gradients? InAdvances in Neural Information Processing Systems, volume 31, 2018

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:d449e11bf9a6d53221390cc6a97e23bea16cc769ac891bb958ffc8b899d47aa8

Observation 7031d63b-4391-429f-a54f-611d2b1f7280 · outbound

This paper cites Schoenholz, and Jeffrey Pennington.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Schoenholz, and Jeffrey Pennington

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:0636fc8d8861b23c2b88f93a506ae9eda9bc664b95248a9ad6d44bfd13ec1c7f

Observation 1ab31dda-82de-44fe-961f-bc675b43c1fd · outbound

This paper cites Schoenholz.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Schoenholz

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:99067c6c40aa500661202278cdcf0739073ea3a42b892bf8eac131bcb1b28729

Observation b53f382f-42ae-4071-8ead-7ed8ac60d03f · outbound

This paper cites an unresolved cited work.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:8529c2b1fddf071f641ae01bf5c6fbfb55adab38ac81f0f7dd276658b8a20fa6

Observation 87ffc089-f4cb-48e6-9e4e-ea0bc163f5d5 · outbound

This paper cites Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:50511d30fffb07eaebad810d6f146704127fecf9e08505e8be712ce43b1a413f

Observation 0aeb3923-10f6-4382-8109-f46aef020baf · outbound

This paper cites Self-consistent dynamical field theory of kernel evo- lution in wide neural networks.Journal of Statistical Mechanics: Theory and Experiment, 2023(11):114009, 2023.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Self-consistent dynamical field theory of kernel evo- lution in wide neural networks.Journal of Statistical Mechanics: Theory and Experiment, 2023(11):114009, 2023

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:bff690a0b2164e2ac44b016c3ceda22855209d51a0d8d90a0ed1ec4a68c6a337

Observation 959669ad-4ff1-42fb-8b8d-85aaaeeedace · outbound

This paper cites Dynamics of finite width kernel and prediction fluctua- tions in mean field neural networks.Journal of Statistical Mechanics: Theory and Experiment, 2024(10):104021, 2024.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Dynamics of finite width kernel and prediction fluctua- tions in mean field neural networks.Journal of Statistical Mechanics: Theory and Experiment, 2024(10):104021, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:e3aad09cd0ad6a56cdb0a6d546ae2117a39ba3ae50c79dd3529218765901382f

Observation e8ad1fe9-e70a-4e8e-8cc3-229c9b3a6e22 · outbound

This paper cites Deep linear network training dynamics from random initialization: Data, width, depth, and hyperparameter transfer.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Deep linear network training dynamics from random initialization: Data, width, depth, and hyperparameter transfer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:781dd9029cb7655dd1ed01e6a5e849b13ddfdf12c572906baa29544faf555f92

Observation 39653b32-d607-4d54-85b1-7e8f50de6c15 · outbound

This paper cites Adaptive kernel predictors from feature-learning infinite limits of neural networks.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Adaptive kernel predictors from feature-learning infinite limits of neural networks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:66a9784916511c9d1d9fba575768e1bf37c0bfb64554ee4943141b69888d24f2

Observation 9dce5377-20ef-4699-b079-4fc60b94cadd · outbound

This paper cites Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:23:28.023366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:77e51a8037f4f0846a3bb49d384fdc7ca16d8c2f1f38ec6c276fec8ce139f31c

Observation 0e87974f-1ddc-4a3c-b3c4-6de0164b5687 · outbound

This paper cites Zuidema, and Stella R.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Zuidema, and Stella R

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:d3d8ed402662532d1a87c97f56f64eb3cedb0b90d16e6f8473248ba3f93b99f1

Observation f5a575c3-3cd2-4a2f-9c56-6de3f6c0d3f6 · outbound

This paper cites Convergence and divergence of language models under different random seeds.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Convergence and divergence of language models under different random seeds

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:1b1a8713d0a365ecf599d2a19063a7babe45e19ef01ed37d46cb49eb2b408319

Observation 52a8d6e3-6a7d-4bc2-9fa7-48bcabe0c4d3 · outbound

This paper cites an unresolved cited work.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:9f273ca23a94292cb767f0d73554424e7d9524c92976675b44ce962f76584b2f

Observation ed066a98-d01a-4a14-b485-c3cb3f699d19 · outbound

This paper cites SeedPrints: Fingerprints can even tell which seed your large language model was trained from.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias SeedPrints: Fingerprints can even tell which seed your large language model was trained from

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:bb604614ed03c6e3f20685e4e724d30830dc07c9ea66d11fabb6906b0a545198

Observation 67844414-9190-4a3a-a8ba-07ad7d680aa5 · outbound

This paper cites Transformers are born biased: Structural inductive biases at random initialization and their practical consequences, 2026.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Transformers are born biased: Structural inductive biases at random initialization and their practical consequences, 2026

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:33258c22549eeeec892c5f564b5964be2ff46589e1c9aa99cb16faec73e06985

Observation 2c0d35fc-4e7e-4645-8698-4d51dcc1099c · outbound

This paper cites Physics of language models: Part 3.1, knowledge storage and extraction.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Physics of language models: Part 3.1, knowledge storage and extraction

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:2b269cbd6e92d53ce0688cc63a83c4ce1f658b680be698893dfcb09a3f7d1283

Observation ddf2782b-5f77-4546-88b9-0c8d117654e8 · outbound

This paper cites Physics of language models: Part 4.1, architecture design and the magic of canon layers.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Physics of language models: Part 4.1, architecture design and the magic of canon layers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:2e9c26fea1db22486905e14a37df09f5a0748cf93e18d7690c87193ff672d374

Observation 6b544f33-ca89-4e22-9208-7ec3b063a18b · outbound

This paper cites an unresolved cited work.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:737881700fadea5bebb8af0a7b2e57d49663070b035f23fe50728f4f7fd3712e

Observation 102e3f34-ffcc-47e2-b4b8-98cc50633790 · outbound

This paper cites Smith, Benoit Dherin, David G.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Smith, Benoit Dherin, David G

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:88abc03ab397cd1b8c0401daa1bc52fc1c69c70aca832a38badba83931cffe85

Observation 340eccbf-0333-4f24-99c4-2f019491ee60 · outbound

This paper cites On the Trajectories of SGD Without Replacement.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias On the Trajectories of SGD Without Replacement

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:23:28.026923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:549aa8870c2a61b80b27a5ff2e9dc4d41af7cd38fcb135d90fb29078c76d2341

Observation 792a5032-7557-4754-b29d-16ae3d49bc63 · outbound

This paper cites How neural networks learn the support is an implicit regularization effect of SGD.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias How neural networks learn the support is an implicit regularization effect of SGD

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:496010b5d1a5e954bc1bb44239f6eb2d392b82992fa2aed0dab794fa61a3a863

Observation bb8430ce-a874-4494-8be8-19900da026c6 · outbound

This paper cites Griffiths and J.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Griffiths and J

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:bcd67879cdf2c59b871a4d3d62bd411e7d0cba1c04f06b2f4a1cec8b7782308f

Observation c2a54f97-0e20-477c-b88a-c01974b367c0 · outbound

This paper cites Springer, Berlin, Heidelberg, 2 edition, 2006.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Springer, Berlin, Heidelberg, 2 edition, 2006

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:dab70420c120313c03f221bf7e5e9fed6541b6a77a74986e02c33d20fb4bd17c

Observation e2f7284f-5f07-47f3-abe4-f65a2c084bee · outbound

This paper cites Implicit regularization in Heavy-ball momentum accelerated stochastic gradient descent.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Implicit regularization in Heavy-ball momentum accelerated stochastic gradient descent

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:23:28.013996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:a13b5a7093ba0cada3973c0096b20143393417d2fb45fbad8263d894a8ed2ee6

Observation 404b6453-3754-445c-a9ed-85da2b1f27b5 · outbound

This paper cites Cattaneo, Jason Matthew Klusowski, and Boris Shigida.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Cattaneo, Jason Matthew Klusowski, and Boris Shigida

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:19cb45b67f345559d3da419e24aa17a67db4f06cb335f7bedc6bd13d664888a7

Observation 4f6bdb28-49e3-4286-9b1c-4c68367e99ac · outbound

This paper cites Implicit regularization in deep matrix factorization.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Implicit regularization in deep matrix factorization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:866e7aac6c038e1e377042ddf1e6c5161ae6332773c50f8cd1bd321b509a2b7e

Observation f1271a0b-d08e-4163-8536-08678091a4bd · outbound

This paper cites Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:23:28.020113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:337a171372205d68ec65ad38c0e7edc63ad44b77e83e4a63a17ade432ff45a87

Observation 4ffcb543-f52f-4a7b-9281-f7295b86d721 · outbound

This paper cites an unresolved cited work.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:ca694d6dd6940ed6f502af57d7a382c9cb2bcd94cf44e2fd6304ab4cf2d3769a

Observation d2a3fd56-1ccb-437a-8a2c-365a15638a5f · outbound

This paper cites Vapnik and Alexey Ya.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Vapnik and Alexey Ya

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:fad9ea58f71ad2a750f0440d74b7ec00bcd142c664ce4f8d2ee7ee483fda26c2

Observation 69668f97-5ebe-4014-99c9-6f4348a28622 · outbound

This paper cites Vapnik.Statistical Learning Theory.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Vapnik.Statistical Learning Theory

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:784c881dbc84cd7ac846b8044f844c02f828e8243c6c7cf4a01df49f40802268

Observation f68f8485-8f33-4102-9bbd-db884b6210e8 · outbound

This paper cites Bartlett.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Bartlett

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:45da234706210735874b078834a26c08cf9c2c38c4bdc241a8f091822f6cecd8

Observation 3fe5b723-7ba7-430c-9839-d8fdf7461f17 · outbound

This paper cites Bartlett and Shahar Mendelson.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Bartlett and Shahar Mendelson

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:eafdacff886a0904741d409d6e0f4d6ee6db83ed591e029f276ebc19a7897c9c

Observation c0731583-5306-4f76-944f-dcf59370b8cf · outbound

This paper cites Norm-based capacity control in neural networks.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Norm-based capacity control in neural networks

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:49fec9e7bfa01cf2f721b4be41fbc7ada256721dfd2eb05008ea7fb086e270b9

Observation 492ff402-a245-4242-b750-52f2f9a9a8b8 · outbound

This paper cites Bartlett, Dylan J.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Bartlett, Dylan J

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:33c4a9b37e07a04c2477f603f6c8fdd2cc5fbdf644023593c525a073d742b380

Observation d872f2b2-7850-41cf-9783-6feb582a0b1f · outbound

This paper cites Understand- ing deep learning requires rethinking generalization.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Understand- ing deep learning requires rethinking generalization

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:32168a72df79d442bb29262fa9b6d6266523c87e6c83e8dc3e11c20337818c8b

Observation 33cc7ff5-7afd-47d2-9d96-1f7646fdc808 · outbound

This paper cites Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, and Simon Lacoste-Julien.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, and Simon Lacoste-Julien

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:eb6ef6d16761979f5576963cb08b84cff3693f15f90425ff01c6619334da89fd

Observation 3dee2d76-6c4b-4dce-90ac-24117bc9507f · outbound

This paper cites Edelman, Fred Zhang, and Boaz Barak.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Edelman, Fred Zhang, and Boaz Barak

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:8bb7a2283ecacb18482e403f4d39fd5b7a810218aee10cfebdc2b3f2725e92bd

Observation 44cf63a5-6c60-41ed-9f6a-5770a84b1bb1 · outbound

This paper cites Train faster, generalize better: Stability of stochastic gradient descent.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Train faster, generalize better: Stability of stochastic gradient descent

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:8efc57b5ec0fbba9aa64c40e83b9cee44aad0098ab335917760c805a7b80f9a7

Observation a3c3d132-634b-42fa-a6e7-99952d24d1e7 · outbound

This paper cites Exploring generalization in deep learning.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Exploring generalization in deep learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:d87bea2556d5801b09afca9faf3c0a79fa507c643cbf79829ec1ff1fc48c0df2

Observation 9572b56d-df5c-4b21-a23d-02476a1b3538 · outbound

This paper cites A PAC- bayesian approach to spectrally-normalized margin bounds for neural networks.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias A PAC- bayesian approach to spectrally-normalized margin bounds for neural networks

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:2b014b01ad875495861a90be1853878a839c915d0ff285d49e5a618d7e8c513b

Observation 54a4c63a-a5a7-401a-b00a-beea8ab6817b · outbound

This paper cites Predicting the generalization gap in deep networks with margin distributions.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Predicting the generalization gap in deep networks with margin distributions

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:3ede4689c1c2e011efe15f49ca4727246f29aa7c91da14db5697cf80a2ed1466

Observation 37ba5b68-355a-4ae4-bd9f-b70bdee85f91 · outbound

This paper cites Fantas- tic generalization measures and where to find them.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Fantas- tic generalization measures and where to find them

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:4fb43b58c9b15d782c669006dd9c9ab37f94f7872c8857150e0ab718f9e4657b

Observation 951c8d20-e516-40a7-9b88-769575bbef0a · outbound

This paper cites Flat minima.Neural Computation, 9(1):1–42, 1997.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Flat minima.Neural Computation, 9(1):1–42, 1997

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:d8970d6e34ef0c1635c9a81339e7db2c8c4366333c7e48aaccd73dce180091f1

Observation 6d44bc74-491a-48d0-9c32-a8be15a015dc · outbound

This paper cites On large-batch training for deep learning: Generalization gap and sharp minima.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias On large-batch training for deep learning: Generalization gap and sharp minima

Reference 59

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:0468ec80ec4bd6d156cbdd3063d36771dd0aecd5b7b41961d678ea760160761c

Observation 0022f79c-12d4-43af-a273-d056702d8447 · outbound

This paper cites Sharp minima can generalize for deep nets.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Sharp minima can generalize for deep nets

Reference 60

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:dc0bb4256aed3f7e032f945d03cf61c0e1e3f5a4add995e66741685b9c909ad1

Observation 85073a1d-1f9c-491b-b6df-b31d1c7d7a5e · outbound

This paper cites Sharpness-aware minimization for efficiently improving generalization.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Sharpness-aware minimization for efficiently improving generalization

Reference 61

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:5dabf9c00ba74d1e5613f77975eceb04265e92f1239a2093d708a954c2478b09

Observation d73c9a05-c160-4b95-9a58-40fa95072139 · outbound

This paper cites McAllester.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias McAllester

Reference 62

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:240298dac0dee94bb044f890b355261e3be8bf8490d5a3ee4421f14364be342e

Observation 9ad66981-77e5-4c2e-b3ba-fb0a2b4d20a9 · outbound

This paper cites an unresolved cited work.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:7e10e02d563f3f661ba69d2145a8a301c44e380a59c580f8d327960dd8ed066c

Observation 4f0f935f-3259-4953-acfd-36887d7c2c9b · outbound

This paper cites Adams, and Peter Orbanz.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Adams, and Peter Orbanz

Reference 64

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:815e6ad187bf9e9bea3d7961ca4d714a3469602748582caa1f9d0cbc02f05f61

Observation 89235dd1-db05-42c0-b5b7-45fb22c95da2 · outbound

This paper cites Stronger generalization bounds for deep nets via a compression approach.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Stronger generalization bounds for deep nets via a compression approach

Reference 65

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:d14517c5d24e210db07c3111cbaf9a658a7fc6fb64352c94c1905258d310cc8d

Observation cee32d27-e80c-4e02-8bea-71c2460319cf · outbound

This paper cites Reconciling modern machine- learning practice and the classical bias–variance trade-off.Proceedings of the National Academy of Sciences, 116(32):15849–15854, 2019.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Reconciling modern machine- learning practice and the classical bias–variance trade-off.Proceedings of the National Academy of Sciences, 116(32):15849–15854, 2019

Reference 66

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:9619cd55803b793f1902be0547fb96f937b0b84813a63a671551623a52f56690

Observation 184ac8ca-3a86-4d9f-b382-48fb077785d0 · outbound

This paper cites Deep double descent: Where bigger models and more data hurt.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Deep double descent: Where bigger models and more data hurt

Reference 67

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:46aed1b2097da35ae373b81ebe64a8bb7c3332050620f1a9392d6544f255a7e1

Observation 5e5d2dc2-eb77-4bd1-8911-6f62172753d2 · outbound

This paper cites Bartlett, Philip M.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Bartlett, Philip M

Reference 68

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:ef0d658a3415b0aac283b922e88b915499a38dc3a7d6c8ba82ee2e0a2384c24a

Observation 293c8f82-3faa-4dc7-944b-de1130079d8a · outbound

This paper cites Neal.Bayesian Learning for Neural Networks, volume 118 ofLecture Notes in Statistics.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Neal.Bayesian Learning for Neural Networks, volume 118 ofLecture Notes in Statistics

Reference 69

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:eae0e6a2f47d1d7a670066df24ad35215a1cb2be39b4a57d7f347aad8dccc41a

Observation 054e625e-72b2-4a31-a638-0b3f96ba4835 · outbound

This paper cites an unresolved cited work.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Unresolved cited work

Reference 70

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:d9cbd30de4a3bf37c87e619d0858b54aaf3bbfdf6d806c1b86ffe8cefeab5332

Observation 6e757fda-d37d-441c-919d-c95247c49b7a · outbound

This paper cites Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein

Reference 71

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:7cf28f669a6e933a94f8121150ea1c0c45f8ecdf192a5514c25eff1df5356c0b

Observation ad1182f6-1d00-4f13-9cbf-0c87261f3761 · outbound

This paper cites Neural tangent kernel: Convergence and generalization in neural networks.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Neural tangent kernel: Convergence and generalization in neural networks

Reference 72

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:7528f67b964136d7696dc5b0ea9bc9d56af8ecdca833d718c1af440463232023

Observation f2dc4f86-eab7-4663-84f1-cd33461828dd · outbound

This paper cites Camargo, and Ard A.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Camargo, and Ard A

Reference 73

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:6fb94bfd919e880bddfb2dfa95a5531685411b49c3e6dc737f1a770383b6ecc5

Observation 06d59cbf-2e4f-49ed-9b2f-d050bc648ff5 · outbound

This paper cites Random deep neural networks are biased towards simple functions.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Random deep neural networks are biased towards simple functions

Reference 74

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:fa31319b16d81be3c557c236683f1fe4a833009cac8758836599a35acae86bdf

Observation 1a58516d-7671-4d56-8fbc-b688353a4082 · outbound

This paper cites On the complexity of finite sequences.IEEE Transactions on Information Theory, 22(1):75–81, 1976.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias On the complexity of finite sequences.IEEE Transactions on Information Theory, 22(1):75–81, 1976

Reference 75

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:ae69d6b8df55fa2dc497e0f1c4b31f01548bd4a5f148d285e6bc7cd4011f0811

Observation efe22477-0dfb-41fe-9b3f-0db4bd858bc7 · outbound

This paper cites A universal algorithm for sequential data compression.IEEE Transactions on Information Theory, 23(3):337–343, 1977.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias A universal algorithm for sequential data compression.IEEE Transactions on Information Theory, 23(3):337–343, 1977

Reference 76

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:a973df5cac245c448f4f3ba4f7646bd5f928d18aca171bdd8eb9242571c3b521

Observation 1c860759-7c1c-4809-8dc0-617d4b1cd3b9 · outbound

This paper cites an unresolved cited work.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Unresolved cited work

Reference 77

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:53e35526a9958c830a190547a31b424888b5412f46b184d820dc513b5968eda7

Observation f38010fe-54d6-4936-a9bd-79b9c8ad3a9d · outbound

This paper cites Simplicity bias in transformers and their ability to learn sparse Boolean functions.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Simplicity bias in transformers and their ability to learn sparse Boolean functions

Reference 78

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:e7e74b78827e96dbec8c505edb0eed4b787b2b93e7a866a606deb76ac631fcf2

Observation f55f61ac-1387-48b1-af04-c458ecbbf9e3 · outbound

This paper cites an unresolved cited work.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Unresolved cited work

Reference 79

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:0fbed4080018c227646ed75b6f0c3a7077b15f1e2c7a9612654108f568c08143

Observation eece98fa-5ad6-465e-b43b-f6304f52742f · outbound

This paper cites Transformers learn low sensitivity functions: Investigations and implications.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Transformers learn low sensitivity functions: Investigations and implications

Reference 80

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:7359839d040666e9ad71c0e32436cf388f7fa9e6f20ac9a19805d494d3f1cbca

Observation e0c8118a-a4b4-42af-9220-8f39358d5e53 · outbound

This paper cites Hamprecht, Yoshua Bengio, and Aaron Courville.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Hamprecht, Yoshua Bengio, and Aaron Courville

Reference 81

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:428b7db5f08f77cad2d86581a85c2aaf0453dcc1e474a6b358c24d38fb6637af

Observation c0a909c4-4503-4be6-a050-86c3adab670c · outbound

This paper cites Abolafia, Jeffrey Pennington, and Jascha Sohl- Dickstein.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Abolafia, Jeffrey Pennington, and Jascha Sohl- Dickstein

Reference 82

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:36c6dabddae302319f7f6e4dbc9edf1a6f28641e45920461e9bd2004f3f5f478

Observation 0816a332-0fc8-4d21-846a-11e50ef39154 · outbound

This paper cites Complexity of linear regions in deep networks.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Complexity of linear regions in deep networks

Reference 83

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:939200cc0e4970edea6cc32ba50c6b64fb8f5e828ecf32d51d87ba89445007c8

Observation 35bb9b61-ad0c-4826-a601-cb840326e886 · outbound

This paper cites Deep ReLU networks have surprisingly few activation patterns.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Deep ReLU networks have surprisingly few activation patterns

Reference 84

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:387e0a6c6386cac6955ebd98cb2a3cd0411ad0ae3bd9588aff270d7be82afaeb

Observation e9526b96-f38d-4fa7-b464-81ee9d34baad · outbound

This paper cites an unresolved cited work.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Unresolved cited work

Reference 85

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:5f22119b727041a4d1da43dd46a64cee8f832d2bc8c53ec940ed80ace50bf3ed

Observation 02b421a0-65e6-42f6-ba40-48f47de19d0d · outbound

This paper cites Neural networks trained with SGD learn distributions of increasing complexity.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Neural networks trained with SGD learn distributions of increasing complexity

Reference 86

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:4eddf1270a9e96154bc4bb68629b5280088cef5ba4cddebcf0d4c2ca42d0dd65

Observation 4a592918-713a-4b0d-a9d9-8c2306f1e8a4 · outbound

This paper cites Simplicity bias and optimization threshold in two-layer ReLU networks, 2024.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Simplicity bias and optimization threshold in two-layer ReLU networks, 2024

Reference 87

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:dfb168031be42bab1a7febcad21538eb3aa350c14f3cfe42f41de5be4317b20d

Observation 85b5d626-2219-480c-9838-2fd867ae5a9f · outbound

This paper cites Saxe, and Peter E.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Saxe, and Peter E

Reference 88

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:c8cbf712493d8e29cecc7217b1b4dfa553cb6915589cfd64bc5f0a5a17f56765

Observation eebaa548-d109-42e1-b268-fc3f84ae757d · outbound

This paper cites Saxe, James L.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Saxe, James L

Reference 89

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:50de87c20d2b288b6bfc07ba9d682f3a0698fd6bf20f3c57f5c69cfc9cd08a3d

Observation 178408fe-3078-4103-838d-85626207a99b · outbound

This paper cites an unresolved cited work.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Unresolved cited work

Reference 90

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:e125fea675ae4375e9a02e265690cb58b0c1847d42f7c320d07f80b95aa37be4

Observation 1a5e7935-6b8e-4660-a503-11a6b3f44acf · outbound

This paper cites Tensor programs I: Wide feedforward or recurrent neural networks of any architecture are gaussian processes, 2020.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Tensor programs I: Wide feedforward or recurrent neural networks of any architecture are gaussian processes, 2020

Reference 91

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:40216121bde79d5b5c9500a1172ae5890f44fcf90d658da99a9f49749ce2d387

Observation 62bf712d-4e7c-4620-83e0-837fa3fe77e2 · outbound

This paper cites Batch normalization: Accelerating deep network training by reducing internal covariate shift.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Batch normalization: Accelerating deep network training by reducing internal covariate shift

Reference 92

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:e7229957333d0caab117d2bbdd61c795670a03a0f349134ffb0b4c3b21e334a7

Observation f3568005-b996-4119-b0f9-e6f2f4ff0287 · outbound

This paper cites Theoretical analysis of auto rate-tuning by batch normalization.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Theoretical analysis of auto rate-tuning by batch normalization

Reference 93

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:a08b88b60cd17c11460f260872bc3fe34263a054d224ea1e66506b23d4a9f1b6

Observation bba10d06-188e-404c-9db6-f7313fd141e2 · outbound

This paper cites Reconciling modern deep learning with tradi- tional optimization analyses: The intrinsic learning rate.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Reconciling modern deep learning with tradi- tional optimization analyses: The intrinsic learning rate

Reference 94

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:9f2b9fff23b1f90ca39788bf385bc64fa6cd97d6872f74042268708d933a4506

Observation 0c9bd7f4-a27b-404c-91ef-1190ab2d06c6 · outbound

This paper cites Number 32 in Notes on Applied Science.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Number 32 in Notes on Applied Science

Reference 95

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:db3acb18f19107c7d198ab1d8ece83398ff48c485af75ec39742b3fa275a6625

Observation d1d9de8f-c6c8-426d-809c-4cda99690ba8 · outbound

This paper cites Monographs on Numerical Analysis.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Monographs on Numerical Analysis

Reference 96

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:5ea96e75762ca3d157fa1ef7d1b651475e085d6ca576f0443c3262dd314978df

Observation 2590ce15-2872-4ec1-a064-b1e0d5b5363e · outbound

This paper cites Higham.Accuracy and Stability of Numerical Algorithms.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Higham.Accuracy and Stability of Numerical Algorithms

Reference 97

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:e681fc4e22c6f8d7db6b4914e87df2c85eecb8b1d3b6efa8614f2888fbd57412

Observation 14ea6b90-f0b4-40f6-b3b8-5819474a3967 · outbound

This paper cites an unresolved cited work.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Unresolved cited work

Reference 98

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:9ac700ea615068dbb4176cfc4d9b058c30c1884ddd3a6ce153e539430ae4a5a5

Observation b1cae5ec-226a-412b-9b15-affc360657b4 · outbound

This paper cites Modified equations for stochastic differential equations.BIT Numerical Mathematics, 46(1):111–125, 2006.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Modified equations for stochastic differential equations.BIT Numerical Mathematics, 46(1):111–125, 2006

Reference 99

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:000cb96f3be66599f83b32233838c4711fafc926b32a80bd7e77c93d0c19a46c

Observation 79a410a8-776d-4d9e-92aa-f9ab6415fbdc · outbound

This paper cites Zygalakis.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Zygalakis

Reference 100

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:b71631a0e00f6355fba1f3a2d356ece22278a30106ee360b19531f3c5af1077e

Observation cac2f792-718c-4331-8a92-62d674eaa89c · outbound

This paper cites Weak backward error analysis for SDEs.SIAM Journal on Numerical Analysis, 50(3):1735–1752, 2012.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Weak backward error analysis for SDEs.SIAM Journal on Numerical Analysis, 50(3):1735–1752, 2012

Reference 101

Resolution
unresolved
no resolver link, observed 2026-06-29T13:20:54.303605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:69f5eecc37796721b7f535f47c2352705c63681b5925f50140a397efaaa8de27

Pith citing papers

No inbound Pith citation observations are available.