Pith. sign in

Paper Citation Record · LEDGER

Saddle-to-Saddle Dynamics in Deep Linear Networks: Small Initialization Training, Symmetry, and Sparsity

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2106.15933.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2106.15933 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:58:26.454460Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T16:07:20.384070Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c5dc415e-0547-4423-bd7c-ffb2aceb4d9d · inbound

Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputs cites this paper.

Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputs Saddle-to-Saddle Dynamics in Deep Linear Networks: Small Initialization Training, Symmetry, and Sparsity

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-24T11:29:24.943153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-24T11:27:58.190151Z digest=sha256:f91b64675d3c1d0eaa8e16ac256b428348bc4908e49c4c5bbca4a9969eb85018

Observation f0835da8-74a9-4d74-b960-dd28cc4b84d8 · inbound

Parameter Symmetry Potentially Unifies Deep Learning Theory cites this paper.

Parameter Symmetry Potentially Unifies Deep Learning Theory Saddle-to-Saddle Dynamics in Deep Linear Networks: Small Initialization Training, Symmetry, and Sparsity

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T19:58:26.454460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:58:26.454460Z digest=sha256:ecdaad95ade9ecd4c88b3ca3748385d7437bea6b6e71b54b0bbac19f6b914913

Observation 769afe19-6e56-443e-aa3b-48ef611a577f · inbound

Adaptive kernel predictors from feature-learning infinite limits of neural networks cites this paper.

Adaptive kernel predictors from feature-learning infinite limits of neural networks Saddle-to-Saddle Dynamics in Deep Linear Networks: Small Initialization Training, Symmetry, and Sparsity

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T11:19:06.913318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:19:06.913318Z digest=sha256:3ddd4b73acc577f02ca921ad3b36d18666025063f35618db9312c9060dcebc4b

Observation 3a8f20e4-0c1c-4f62-8650-bbd5be08d4ff · inbound

Gradient Flow Equations for Deep Linear Neural Networks: A Survey from a Network Perspective cites this paper.

Gradient Flow Equations for Deep Linear Neural Networks: A Survey from a Network Perspective Saddle-to-Saddle Dynamics in Deep Linear Networks: Small Initialization Training, Symmetry, and Sparsity

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T22:31:44.546962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:31:44.546962Z digest=sha256:27130d872e6fc47a54a91308daeb181aabe4cc2747b189579a86372fb977b219

Observation 586fda19-a8a8-40da-b22c-a0822fc908b7 · inbound

Over-Alignment vs Over-Fitting: The Role of Feature Learning Strength in Generalization cites this paper.

Over-Alignment vs Over-Fitting: The Role of Feature Learning Strength in Generalization Saddle-to-Saddle Dynamics in Deep Linear Networks: Small Initialization Training, Symmetry, and Sparsity

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-03T06:00:14.652992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:00:14.652992Z digest=sha256:d15d5532303f818d167e7ef35d33d821be6154e8cbb3b486cf10af3244a6b59d

Observation 9b281b65-6c0a-49ee-bbb4-85a0350aab2a · inbound

There Will Be a Scientific Theory of Deep Learning cites this paper.

There Will Be a Scientific Theory of Deep Learning Saddle-to-Saddle Dynamics in Deep Linear Networks: Small Initialization Training, Symmetry, and Sparsity

Reference 139

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:21:09.144024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-09T20:11:17.616190Z digest=sha256:29bfb7349871c34d3a01284351af2070d2f3fdb417d0076e58d4e50a31d79c5d

Observation 4255e4da-dfcd-46a1-be07-8dca82121aa8 · inbound

Low-Rank Adaptation Redux for Large Models cites this paper.

Low-Rank Adaptation Redux for Large Models Saddle-to-Saddle Dynamics in Deep Linear Networks: Small Initialization Training, Symmetry, and Sparsity

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:26:03.923173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T21:48:48.992712Z digest=sha256:7ae8980328c0c5fb01cdd9f8bb6882f447d81ee05ac8e67d03b5bd7d73ae0269

Observation e986a627-14a1-4570-890d-91fdda2ab573 · inbound

A Theory of Saddle Escape in Deep Nonlinear Networks cites this paper.

A Theory of Saddle Escape in Deep Nonlinear Networks Saddle-to-Saddle Dynamics in Deep Linear Networks: Small Initialization Training, Symmetry, and Sparsity

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:51:05.621784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T14:54:47.763122Z digest=sha256:0878dedb8753ebd3eaea3cce1641f4a8706c44e6759b4fb6ea4d474bca402bd9

Observation 0580cf1d-33f6-493e-a206-c102a1e917a9 · inbound

A Theory of Saddle Escape in Deep Nonlinear Networks cites this paper.

A Theory of Saddle Escape in Deep Nonlinear Networks Saddle-to-Saddle Dynamics in Deep Linear Networks: Small Initialization Training, Symmetry, and Sparsity

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:25:54.524070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T02:22:38.751375Z digest=sha256:57643cf7f6dad74e713ba732643ce38e8aa1e0dfce52c8239863e47c304b9be6

Observation cc13ca84-3b89-473a-b861-b4f28adf5228 · inbound

A Theory of Saddle Escape in Deep Nonlinear Networks cites this paper.

A Theory of Saddle Escape in Deep Nonlinear Networks Saddle-to-Saddle Dynamics in Deep Linear Networks: Small Initialization Training, Symmetry, and Sparsity

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:45:12.010776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T00:37:16.364388Z digest=sha256:66b64e20d99601d629e8719e823c576c12fd5b997562511a722df08331b7a244

Observation 2c200eaa-fce0-4c85-9201-bbd28c6cc082 · inbound

The Implicit Bias of Depth: From Neural Collapse to Softmax Codes cites this paper.

The Implicit Bias of Depth: From Neural Collapse to Softmax Codes Saddle-to-Saddle Dynamics in Deep Linear Networks: Small Initialization Training, Symmetry, and Sparsity

Reference 111

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:26:38.967729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-25T05:26:15.556205Z digest=sha256:dffa248f2a92e314c58400a4859d2c7c10db0afbdabc2aedc0176d4291a666b4

Observation 1ef28b9a-a2f5-49ec-b042-a50300625c0b · inbound

Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention cites this paper.

Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention Saddle-to-Saddle Dynamics in Deep Linear Networks: Small Initialization Training, Symmetry, and Sparsity

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:43:15.235293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T08:39:14.327838Z digest=sha256:561f6978820eaab9f00b6c4db4882c81649b54603fec465599a2f6a5bb8fafc2

Observation c9b9bdf9-7ed9-495c-b542-75333cb6dc10 · inbound

Incremental Learning in Mirror Flows cites this paper.

Incremental Learning in Mirror Flows Saddle-to-Saddle Dynamics in Deep Linear Networks: Small Initialization Training, Symmetry, and Sparsity

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T11:49:50.323221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T07:41:45.660598Z digest=sha256:200b4f21fb160d15cbe3ff0a2254f17936b32d662a5137d365fb90861aa3c327

Observation 5722c635-da17-4046-a543-86ce7abeff20 · inbound

Muon learns balanced solutions in matrix factorization without slow saddle-to-saddle dynamics cites this paper.

Muon learns balanced solutions in matrix factorization without slow saddle-to-saddle dynamics Saddle-to-Saddle Dynamics in Deep Linear Networks: Small Initialization Training, Symmetry, and Sparsity

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:24:21.767685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T07:16:14.269675Z digest=sha256:bf54e4748abd9e2ed8d6cd3b07ee2e1cb1b3186c142fa4596489d2d80649f900

Observation 79880936-ea8e-46f1-8465-aee9ee8437c2 · inbound

Effective dynamics of the Sinkhorn algorithm in the regime of low entropy regularization cites this paper.

Effective dynamics of the Sinkhorn algorithm in the regime of low entropy regularization Saddle-to-Saddle Dynamics in Deep Linear Networks: Small Initialization Training, Symmetry, and Sparsity

Reference 132

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:06:47.629924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-02T08:05:55.418491Z digest=sha256:5a69d7ec980b1d0bf0f8b33a12ca92f03253ec56f52ea0b1a0334d6967e2d651

Observation 30a82447-6799-47ae-b1a5-dddadb43daf8 · inbound

Optimal Learning Rate Scaling Depends on Data in Deep Scalar Linear Networks cites this paper.

Optimal Learning Rate Scaling Depends on Data in Deep Scalar Linear Networks Saddle-to-Saddle Dynamics in Deep Linear Networks: Small Initialization Training, Symmetry, and Sparsity

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-10T16:07:20.385724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-10T16:03:20.699893Z digest=sha256:e1cbce15fabc05be61d8d6cb831fda0ea2396d707ba7d1367ee9f8595049fd81

Observation 784de15a-c03b-4197-8571-f18e7530f66d · inbound

Singular perturbations and hierarchical learning in two-layer neural networks cites this paper.

Singular perturbations and hierarchical learning in two-layer neural networks Saddle-to-Saddle Dynamics in Deep Linear Networks: Small Initialization Training, Symmetry, and Sparsity

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-14T08:40:30.522134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T08:40:30.522134Z digest=sha256:2c302c478828253abf194b227e35b86e9ba0d5aa5c301549509541c4c250864e