Pith. sign in

Paper Citation Record · LEDGER

The Break-Even Point on Optimization Trajectories of Deep Neural Networks

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2002.09572.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2002.09572 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:26:43.454094Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:56:29.475038Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 836d0d80-4c5a-441e-ba53-19e71e36882a · inbound

Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks cites this paper.

Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks The Break-Even Point on Optimization Trajectories of Deep Neural Networks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:26:43.454094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:26:43.454094Z digest=sha256:eb3d543fb89342676bfaa9a9ecd311de8308a3d47a0cc57c7b28e7d07c1de11d

Observation 4cf84440-425f-4b07-9ebf-96c81e38221a · inbound

Momentum Further Constrains Sharpness at the Edge of Stochastic Stability cites this paper.

Momentum Further Constrains Sharpness at the Edge of Stochastic Stability The Break-Even Point on Optimization Trajectories of Deep Neural Networks

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:20:26.375256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T13:12:58.210967Z digest=sha256:cd93f09565617d7ec44df6c12049d6c5ce2493c699bc3af44a3c805694982417

Observation ae0a759a-e3d1-4286-b518-e006e3b8598c · inbound

Generalization at the Edge of Stability cites this paper.

Generalization at the Edge of Stability The Break-Even Point on Optimization Trajectories of Deep Neural Networks

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:38:17.394321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T02:37:13.715607Z digest=sha256:60fff897687024946e88588622724dc31333caa1add79e065d3c9220fe850191

Observation 97bd1e6f-cd99-4618-abe9-eb424b44922d · inbound

There Will Be a Scientific Theory of Deep Learning cites this paper.

There Will Be a Scientific Theory of Deep Learning The Break-Even Point on Optimization Trajectories of Deep Neural Networks

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:21:09.127698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-09T20:11:17.616190Z digest=sha256:8327eaa7fed71049e3dfdf4bb73eb85b8ff6d5fab5d08ad2f76d07a25e5fb475

Observation f7516ce2-91eb-4742-9f20-7476e596311b · inbound

The Spectral Lifecycle of Transformer Training: Transient Compression Waves, Persistent Spectral Gradients, and the Q/K--V Asymmetry cites this paper.

The Spectral Lifecycle of Transformer Training: Transient Compression Waves, Persistent Spectral Gradients, and the Q/K--V Asymmetry The Break-Even Point on Optimization Trajectories of Deep Neural Networks

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:58:15.679799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T20:56:56.460027Z digest=sha256:dadd861bac9d4a57a48806f9fbdf7e59e83171c2d0dc1c849e0ae87b2d60bfb4

Observation 224e3133-eb4b-470f-b08a-b74405f67348 · inbound

Does Weight Decay Enhance Training Stability? cites this paper.

Does Weight Decay Enhance Training Stability? The Break-Even Point on Optimization Trajectories of Deep Neural Networks

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:53:42.954328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T19:49:01.351717Z digest=sha256:eb707593250b75b7e5e8b80e2516baecb0cb09549342a79a2c11e597422de10b

Observation 22c3a08f-5f3e-4c5c-8d7e-7d96d717b208 · inbound

Edge of Stability Selectively Shapes Learning Across the Data Distribution cites this paper.

Edge of Stability Selectively Shapes Learning Across the Data Distribution The Break-Even Point on Optimization Trajectories of Deep Neural Networks

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:56:29.476938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T10:27:39.029895Z digest=sha256:440df0ec1fda427f785f26f3e1aa25a80b1caacf6553f2d2026ec256f6fd7e99

Observation ff1d4b52-464e-4c3c-a238-677888512589 · inbound

Weight-norm Criticality: A Mechanism for Loss Spikes Induced by the Normalization and Weight Decay cites this paper.

Weight-norm Criticality: A Mechanism for Loss Spikes Induced by the Normalization and Weight Decay The Break-Even Point on Optimization Trajectories of Deep Neural Networks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T08:50:41.552689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T08:50:41.552689Z digest=sha256:af217616f3aaf56cc2988904bc92adba6fe1e95fc80b1eda3125d6d59ce6f1ee

Observation b5a819b3-65cf-4b6b-914e-3f3421486652 · inbound

A Defense of the Quadratic Model cites this paper.

A Defense of the Quadratic Model The Break-Even Point on Optimization Trajectories of Deep Neural Networks

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-01T06:58:09.738474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:58:09.738474Z digest=sha256:5fb2cde28e6a1ba0a3501313d6209618379b782eae287e34a2f3bf83d492d1bd