Pith. sign in

Paper Citation Record · LEDGER

Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks

As of 11 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 1 inbound Pith citation observation for arXiv:2501.09137.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.09137 v2

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:26:43.530125Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T13:20:54.303605Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T13:23:28.017867Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a3c0857e-9576-4883-b16a-138151270d33 · outbound

This paper cites write newline.

Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T20:26:43.412162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:26:43.412162Z digest=sha256:29eac25fe14354e9416a4e18225bb38c93b1ad54ce8bf18a517726606bcfea3d

Observation c5f48a3d-e7f8-4ddc-a62f-e1616fe57885 · outbound

This paper cites T., Suarez, F., and Zhang, Y.

Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks T., Suarez, F., and Zhang, Y

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:26:43.941693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:26:43.417092Z digest=sha256:768851816f01e1fd36704b528fb0ea1426692107b61e551ffc447e847751a8db

Observation c5b2b8e0-c4fb-4cb4-b295-8567443c7197 · outbound

This paper cites A Convergence Analysis of Gradient Descent for Deep Linear Neural Networks.

Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks A Convergence Analysis of Gradient Descent for Deep Linear Neural Networks

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:26:43.421051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:26:43.421051Z digest=sha256:a75d13c6d7ef2832739c38b7adb3ad315dc7de38a71551f09aeba059318aa2af

Observation da037326-655f-489c-9d80-f24bc6feaf81 · outbound

This paper cites P., Selman, B., and Weinberger, K.

Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks P., Selman, B., and Weinberger, K

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:26:43.928142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:26:43.426425Z digest=sha256:8c9504f790f1617c2509cc0c51712c1c3def363e679dc1a90a4aba8ebf784ab0

Observation e59cfc15-b3c9-487d-88e6-bf090f02520f · outbound

This paper cites Optimization Methods for Large-Scale Machine Learning.

Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks Optimization Methods for Large-Scale Machine Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T20:26:43.430882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:26:43.430882Z digest=sha256:7fc83a0e01391d943d28d19e22f5c25ed7e972abb3a22ccfea192005ba4b7c97

Observation f5d9d63a-5642-4d59-ba11-4f91cf4fb02d · outbound

This paper cites and Bruna, J.

Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks and Bruna, J

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:26:43.908686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:26:43.435109Z digest=sha256:04faecd60c668fa8b9d87446eed17cdbe3762c5593cd2d813dc299cdfe308445

Observation 5ec7451e-75d1-4dd3-9dad-51763c0f20e3 · outbound

This paper cites Z., and Talwalkar, A.

Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks Z., and Talwalkar, A

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:26:43.889619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:26:43.440119Z digest=sha256:738929d696bb15820b8cafaa0770b3c8f0bb0bfefdc6fed623fe498033061b99

Observation 98fce254-a364-4024-8a8d-cf374f5f1177 · outbound

This paper cites Implicit Regularization of Discrete Gradient Dynamics in Linear Neural Networks.

Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks Implicit Regularization of Discrete Gradient Dynamics in Linear Neural Networks

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:26:43.873290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:26:43.445461Z digest=sha256:5ae04277081c7776842635d4c0dc72c442e30768e52a0dbf25082e1884131270

Observation a66da8c3-f9f6-4ce4-8668-f04cfd0b3d25 · outbound

This paper cites and Schmidhuber, J.

Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks and Schmidhuber, J

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:26:43.857854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:26:43.449369Z digest=sha256:976ca09d9a3f67482413e4737243c9e3eed9700442b1be50f9b7efe2daf93e91

Observation 836d0d80-4c5a-441e-ba53-19e71e36882a · outbound

This paper cites The Break-Even Point on Optimization Trajectories of Deep Neural Networks.

Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks The Break-Even Point on Optimization Trajectories of Deep Neural Networks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:26:43.454094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:26:43.454094Z digest=sha256:eb3d543fb89342676bfaa9a9ecd311de8308a3d47a0cc57c7b28e7d07c1de11d

Observation fb4d54e9-6740-4744-9a80-a8b7f500874f · outbound

This paper cites S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P.

Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:26:43.836855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:26:43.457845Z digest=sha256:eaa063e7cc58f67ee424e96e4151b16d614cdfcc6755ad817c97bbf6e2ccd4cd

Observation be8bfd10-f67b-4c79-a432-26eb978412e4 · outbound

This paper cites B., and Müller, K.-R.

Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks B., and Müller, K.-R

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:26:43.817831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:26:43.461625Z digest=sha256:8c54baf8f310bb0ebd41da21c027ccb9933be6fff94cf4afe9794f7bb1275959

Observation 7f64b7da-712d-495f-8161-69a21ad71df1 · outbound

This paper cites The large learning rate phase of deep learning: the catapult mechanism.

Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks The large learning rate phase of deep learning: the catapult mechanism

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T20:26:43.466024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:26:43.466024Z digest=sha256:8bd222d63ab6a98151a830334bdd2f1514245df584bc233fbf946de0d6611639

Observation 9d8049ef-d133-44ef-920f-a71bdfe7ac22 · outbound

This paper cites Towards explaining the regularization effect of initial large learning rate in training neural networks.

Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks Towards explaining the regularization effect of initial large learning rate in training neural networks

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:26:43.802916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:26:43.470884Z digest=sha256:efc3be3ae4dc64c3c0125ebe444f83e9710218e5685166b43ffc9c177bac3012

Observation 4181a5dd-1f4d-4c31-a8c4-174f953c33b4 · outbound

This paper cites M., Rauhut, H., and Terstiege, U.

Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks M., Rauhut, H., and Terstiege, U

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:26:43.784615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:26:43.475099Z digest=sha256:66b6e4e9ba52104724eb2565a4448bbbdcaa1a39198457bbfb0aa48859a5bc39

Observation b2a8ab38-6d47-4557-8a67-bc6e331773df · outbound

This paper cites The Effect of Network Width on Stochastic Gradient Descent and Generalization : an Empirical Study.

Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks The Effect of Network Width on Stochastic Gradient Descent and Generalization : an Empirical Study

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:26:43.767367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:26:43.480674Z digest=sha256:8692dd01e550ed2ead5f8422aa38b79c3052344709f9a1d8caa6c509769addf5

Observation ca90ba2d-03ac-4964-967f-6c2b1c79ed06 · outbound

This paper cites an unresolved cited work.

Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:26:43.752819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:26:43.485432Z digest=sha256:8eeebe7aa7b10ba19cd8d779e2f6fff5e1fc919f560114b02a5ca7dd8ff2367f

Observation 2dd077db-079a-4f7d-bebe-0efc2e47a981 · outbound

This paper cites Exact solutions to the nonlinear dynamics of learning in deep linear neural networks.

Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks Exact solutions to the nonlinear dynamics of learning in deep linear neural networks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T20:26:43.492055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:26:43.492055Z digest=sha256:c71445a0ae4cb4c146dab3a9e057bb580f5a9fb4df39d6fb6cdb535d26607bd9

Observation fef5958b-0a7e-4721-b896-35d878e34995 · outbound

This paper cites A Bayesian Perspective on Generalization and Stochastic Gradient Descent.

Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks A Bayesian Perspective on Generalization and Stochastic Gradient Descent

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T20:26:43.497375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:26:43.497375Z digest=sha256:867384599f16d7ab45cefdbfd1dc34f03121d674999b8dc6178888ec4ffa7c3f

Observation a9308a71-ecfb-4aa5-bc96-0aba767f8dbf · outbound

This paper cites D., and Vidal, R.

Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks D., and Vidal, R

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:26:43.737100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:26:43.502061Z digest=sha256:7512affa607123e246003e30a098ae069a04884b4596a50f05a1e37c30741d2b

Observation a26044fc-ce5c-4f30-bf6e-eeab91f6b694 · outbound

This paper cites Large Learning Rate Tames Homogeneity : Convergence and Balancing Effect.

Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks Large Learning Rate Tames Homogeneity : Convergence and Balancing Effect

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:26:43.717299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:26:43.508487Z digest=sha256:cf3d3d6abfee5c3435aa304506d3fd18621663fc5e70ec1229524db1f72928bb

Observation edd11cbf-48dc-408a-9fe7-0f43af6fa797 · outbound

This paper cites Three Mechanisms of Feature Learning in a Linear Network.

Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks Three Mechanisms of Feature Learning in a Linear Network

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:26:43.512820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:26:43.512820Z digest=sha256:1951f041450546aee8427ad6526506c7eb0dbd60d4b9dde10b236ff0ee9daed5

Observation 0952d60f-06b8-453c-b4c8-c883720e3fd0 · outbound

This paper cites Linear convergence of gradient descent for finite width over-parametrized linear networks with general initialization.

Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks Linear convergence of gradient descent for finite width over-parametrized linear networks with general initialization

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:26:43.700426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:26:43.516613Z digest=sha256:2b978ac981a1378a24451f6d5255c1c67c08d4fe4202c3e7a4d9c13fb8cc5ef9

Observation 7aa214f3-1596-4b92-a4ec-3480c1750796 · outbound

This paper cites @esa (Ref.

Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks @esa (Ref

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T20:26:43.521281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:26:43.521281Z digest=sha256:5a0251d449ff3b21ec2c673fea54fa64c71e6d8f9bcff9ff5afba531d5a54b7d

Observation 5e389e42-4f56-4eaa-bd43-7f8f4921c8b6 · outbound

This paper cites an unresolved cited work.

Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T20:26:43.525592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:26:43.525592Z digest=sha256:fe0199e9b22c4c96e15ca33b89af574b549261a15edbb630abb18f6e09f545e7

Observation b494e2f5-df1c-419d-8b51-6f59a2a6230a · outbound

This paper cites ,# (7),01444 '9=82<.342C 2! !22222222222222222222222222222222222222222222222222.

Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks ,# (7),01444 '9=82<.342C 2! !22222222222222222222222222222222222222222222222222

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:26:43.666612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T20:26:43.530125Z digest=sha256:5cd143b01a6ae358d0d334d330750271ceddd7f8f2b757ddf7c1577e710dfa0a

Pith citing papers

Observation f1271a0b-d08e-4163-8536-08678091a4bd · inbound

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias cites this paper.

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:23:28.020113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T13:20:54.303605Z digest=sha256:89e3d754acd6f57f75f3253ded0c128a7b78dd410be13617339b73bec167fb7d