Pith. sign in

Paper Citation Record · LEDGER

The Break-Even Point on Optimization Trajectories of Deep Neural Networks

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2002.09572.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2002.09572 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:26:43.454094Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:56:29.475038Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 836d0d80-4c5a-441e-ba53-19e71e36882a · inbound

Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks cites this paper.

Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks The Break-Even Point on Optimization Trajectories of Deep Neural Networks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:26:43.454094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:26:43.454094Z digest=sha256:eb3d543fb89342676bfaa9a9ecd311de8308a3d47a0cc57c7b28e7d07c1de11d

Observation 4cf84440-425f-4b07-9ebf-96c81e38221a · inbound

Momentum Further Constrains Sharpness at the Edge of Stochastic Stability cites this paper.

Momentum Further Constrains Sharpness at the Edge of Stochastic Stability The Break-Even Point on Optimization Trajectories of Deep Neural Networks

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:20:26.375256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T13:12:58.210967Z digest=sha256:e55ceeffe2163f171e938241e915230f98e07bbe732918919fc872c7d8bf29ac

Observation ae0a759a-e3d1-4286-b518-e006e3b8598c · inbound

Generalization at the Edge of Stability cites this paper.

Generalization at the Edge of Stability The Break-Even Point on Optimization Trajectories of Deep Neural Networks

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:38:17.394321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T02:37:13.715607Z digest=sha256:f98f8c0b51a00b24cc02448ce31bf33be70008dac7494973e42629d49da17134

Observation 97bd1e6f-cd99-4618-abe9-eb424b44922d · inbound

There Will Be a Scientific Theory of Deep Learning cites this paper.

There Will Be a Scientific Theory of Deep Learning The Break-Even Point on Optimization Trajectories of Deep Neural Networks

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:21:09.127698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-09T20:11:17.616190Z digest=sha256:5e494fc654da5bf9bd432443f909adf8eb301e6cfc8e6a337a87fac1c504bc9e

Observation f7516ce2-91eb-4742-9f20-7476e596311b · inbound

The Spectral Lifecycle of Transformer Training: Transient Compression Waves, Persistent Spectral Gradients, and the Q/K--V Asymmetry cites this paper.

The Spectral Lifecycle of Transformer Training: Transient Compression Waves, Persistent Spectral Gradients, and the Q/K--V Asymmetry The Break-Even Point on Optimization Trajectories of Deep Neural Networks

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:58:15.679799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T20:56:56.460027Z digest=sha256:0d5598a92ba719dd7b130ac564cbd95a56b83b52f91ff65309fba48c90680566

Observation 224e3133-eb4b-470f-b08a-b74405f67348 · inbound

Does Weight Decay Enhance Training Stability? cites this paper.

Does Weight Decay Enhance Training Stability? The Break-Even Point on Optimization Trajectories of Deep Neural Networks

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:53:42.954328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T19:49:01.351717Z digest=sha256:829c297edb11f323b0c0c63c7fab1b41d356424c7b77b13a0efd0ae983e2f47a

Observation 22c3a08f-5f3e-4c5c-8d7e-7d96d717b208 · inbound

Edge of Stability Selectively Shapes Learning Across the Data Distribution cites this paper.

Edge of Stability Selectively Shapes Learning Across the Data Distribution The Break-Even Point on Optimization Trajectories of Deep Neural Networks

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:56:29.476938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T10:27:39.029895Z digest=sha256:286da7e229de1f7124adbb4f4c06406ef27b0ba3f7355dcc70d4d7a8b87af097

Observation ff1d4b52-464e-4c3c-a238-677888512589 · inbound

Weight-norm Criticality: A Mechanism for Loss Spikes Induced by the Normalization and Weight Decay cites this paper.

Weight-norm Criticality: A Mechanism for Loss Spikes Induced by the Normalization and Weight Decay The Break-Even Point on Optimization Trajectories of Deep Neural Networks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T08:50:41.552689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T08:50:41.552689Z digest=sha256:af217616f3aaf56cc2988904bc92adba6fe1e95fc80b1eda3125d6d59ce6f1ee

Observation b5a819b3-65cf-4b6b-914e-3f3421486652 · inbound

A Defense of the Quadratic Model cites this paper.

A Defense of the Quadratic Model The Break-Even Point on Optimization Trajectories of Deep Neural Networks

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-01T06:58:09.738474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:58:09.738474Z digest=sha256:5fb2cde28e6a1ba0a3501313d6209618379b782eae287e34a2f3bf83d492d1bd