Pith. sign in

Paper Citation Record · LEDGER

The large learning rate phase of deep learning: the catapult mechanism

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 36 inbound Pith citation observations for arXiv:2003.02218.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2003.02218 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 36 of 36 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T16:31:36.516166Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

60
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation aa995168-f597-4e21-9019-e9e4d2652a75 · inbound

Scaling Laws for Autoregressive Generative Modeling cites this paper.

Scaling Laws for Autoregressive Generative Modeling The large learning rate phase of deep learning: the catapult mechanism

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:49:43.819446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T07:49:43.711653Z digest=sha256:faaf915bf93d659b3bbc43deee1d067982f26ce63fc0a98a2823f51eaa4f0bda

Observation 19ece7e6-6614-46d0-b7da-52a487f6eaa4 · inbound

Scaling Laws for Transfer cites this paper.

Scaling Laws for Transfer The large learning rate phase of deep learning: the catapult mechanism

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:58:13.693119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-18T00:58:13.116663Z digest=sha256:1481ef48ec67feaae4631381e23fee71ee1f3f237d7e9425708230c42a22c492

Observation 617e79ba-f6d4-4378-84b8-de100a8cfc73 · inbound

A General Language Assistant as a Laboratory for Alignment cites this paper.

A General Language Assistant as a Laboratory for Alignment The large learning rate phase of deep learning: the catapult mechanism

Reference 94

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:22:58.981470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-11T14:22:57.925354Z digest=sha256:498f0e25dd22eda5f8a78592321ceb2df0eedf11dc2e8bc5e5eedbefe0b44951

Observation 6261559f-2e0e-4b22-b45a-4d8df32b9d36 · inbound

Language Models (Mostly) Know What They Know cites this paper.

Language Models (Mostly) Know What They Know The large learning rate phase of deep learning: the catapult mechanism

Reference 152

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T15:42:47.810297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T15:42:47.274448Z digest=sha256:053e23ae2dfa3da89cb177f4751585b17758bf71778e01f063ec2709d6a0a11a

Observation a15c71d2-eff2-4c84-a603-85886ace9913 · inbound

A ghost mechanism: An analytical model of abrupt learning in recurrent networks cites this paper.

A ghost mechanism: An analytical model of abrupt learning in recurrent networks The large learning rate phase of deep learning: the catapult mechanism

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:32:38.851832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T06:31:21.006691Z digest=sha256:15f3a4ebb513eb0c2750dc5f2bc8b7b1cbf068b7f5bbafb3312a5610389a9015

Observation 2b72b369-736d-41cd-b45f-90aec27f8d62 · inbound

Right Time to Learn:Promoting Generalization via Bio-inspired Spacing Effect in Knowledge Distillation cites this paper.

Right Time to Learn:Promoting Generalization via Bio-inspired Spacing Effect in Knowledge Distillation The large learning rate phase of deep learning: the catapult mechanism

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-08T16:31:36.516166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:31:36.516166Z digest=sha256:09aaf474f3b78ed9a90206baa656926f6837ad5d56b3558f24e52780c2af9f5e

Observation 9f45d863-9706-4ebf-8ea7-1d696367aee8 · inbound

Adaptive Preconditioners Trigger Loss Spikes in Adam cites this paper.

Adaptive Preconditioners Trigger Loss Spikes in Adam The large learning rate phase of deep learning: the catapult mechanism

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:42:54.707684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:42:54.707684Z digest=sha256:3a7653cd3380a6e36452d72afe99667855aea2364f32db533a10b58d50c8b653

Observation 2d4679b0-9100-4b52-8d3a-6bd1049139c2 · inbound

Large Spikes in Stochastic Gradient Descent: A Large-Deviations View cites this paper.

Large Spikes in Stochastic Gradient Descent: A Large-Deviations View The large learning rate phase of deep learning: the catapult mechanism

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:30:51.043994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T13:30:04.366752Z digest=sha256:f54312940aab863db2ac9ffba9a7d34893ce35ca2f7cefcafc608ac99947be60

Observation 6f90fd1c-8a12-40b0-bdc5-5f8197777bfc · inbound

(How) Learning Rates Regulate Catastrophic Overtraining cites this paper.

(How) Learning Rates Regulate Catastrophic Overtraining The large learning rate phase of deep learning: the catapult mechanism

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:40:26.497872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:40:19.845409Z digest=sha256:6d195a2b5dde0dc58ef31a5e511fac109605382cb595e17f6fa44fe3beb39b7e

Observation 758cdf53-4afe-46cd-94f4-7c0e3a9d543b · inbound

(How) Learning Rates Regulate Catastrophic Overtraining cites this paper.

(How) Learning Rates Regulate Catastrophic Overtraining The large learning rate phase of deep learning: the catapult mechanism

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T05:30:35.676773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:30:35.676773Z digest=sha256:445f7a18e1716c846e5d8793679174857b2c6f3aee819054f48a9e163c4531c1

Observation abbf4d66-e621-4a0d-b426-2cab6f18c99e · inbound

Zeroth-Order Optimization at the Edge of Stability cites this paper.

Zeroth-Order Optimization at the Edge of Stability The large learning rate phase of deep learning: the catapult mechanism

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:20:10.144195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T11:18:30.486247Z digest=sha256:ffde23983df532aef29864149bae927ca9d099eebb97142e93a3a108ef03743c

Observation f1f63203-12e3-4bab-a3b4-d460b16298df · inbound

Zeroth-Order Optimization at the Edge of Stability cites this paper.

Zeroth-Order Optimization at the Edge of Stability The large learning rate phase of deep learning: the catapult mechanism

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T20:06:39.488359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:06:39.488359Z digest=sha256:d6166fa031864375f8f70abe88e094ceb7825d1962387fc657d44b5253f6b9d1

Observation 9f1484d4-1a70-4671-b587-2d69ce24461c · inbound

The Origin of Edge of Stability cites this paper.

The Origin of Edge of Stability The large learning rate phase of deep learning: the catapult mechanism

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:36:07.948407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T01:24:55.474104Z digest=sha256:8ddfab4c00d7cc624312fe1faf291e01c82708c91dff3ca72ba3f5de4a537bb3

Observation 9602fa3f-807d-4725-8944-0d29c099e53f · inbound

SGD at the Edge of Stability: The Stochastic Sharpness Gap cites this paper.

SGD at the Edge of Stability: The Stochastic Sharpness Gap The large learning rate phase of deep learning: the catapult mechanism

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:59:49.391917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T00:59:39.800662Z digest=sha256:4f827502976b148d0df5671acb537bb1bcc199851083fe04ed0b60320e6b86e0

Observation 82f3b924-b8a0-420e-8775-d27a2173f4d6 · inbound

There Will Be a Scientific Theory of Deep Learning cites this paper.

There Will Be a Scientific Theory of Deep Learning The large learning rate phase of deep learning: the catapult mechanism

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:21:08.996891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-09T20:11:17.616190Z digest=sha256:10eced8045c9a836e7139554e2eae86324abf9e5eba0358b33be17de4c0e4cd6

Observation fa3bea90-131e-48d7-94a2-4a0d1c1a8b4d · inbound

Finite-Size Gradient Transport in Large Language Model Pretraining: From Cascade Size to Intensive Transport Efficiency cites this paper.

Finite-Size Gradient Transport in Large Language Model Pretraining: From Cascade Size to Intensive Transport Efficiency The large learning rate phase of deep learning: the catapult mechanism

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T15:00:31.040454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T14:47:47.383941Z digest=sha256:4efb802dabf40304fe66a894a8e20a2014f3be255ba5fe72b2dc55200bb13197

Observation f0c3f25a-a64b-4ad5-9fd7-6b6673039343 · inbound

Endogenous Regime Switching Driven by Scalar-Irreducible Learning Dynamics cites this paper.

Endogenous Regime Switching Driven by Scalar-Irreducible Learning Dynamics The large learning rate phase of deep learning: the catapult mechanism

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:36:01.856918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:34:29.652388Z digest=sha256:73586664b06aef1e217c19087fdaec565d0d372046e92e8a84a9f7495ba1e9a7

Observation 6bbaf906-a3d7-44ce-9fde-e2b624d0f59e · inbound

A Rod Flow Model for Adam at the Edge of Stability cites this paper.

A Rod Flow Model for Adam at the Edge of Stability The large learning rate phase of deep learning: the catapult mechanism

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:00:56.647680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-11T00:54:24.615112Z digest=sha256:6f93317d1d3dbd08a5b43ed33fcc922db9420e6b01c8469e272f05ab193bdd4e

Observation 22c671cb-62d2-45e7-a830-c0bd1247e408 · inbound

Can Muon Fine-tune Adam-Pretrained Models? cites this paper.

Can Muon Fine-tune Adam-Pretrained Models? The large learning rate phase of deep learning: the catapult mechanism

Reference 77

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:51:28.677058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T03:53:11.469583Z digest=sha256:1b4f18174dcf2ef82e6a97c8c0cf99f01057ed8e0e9780ef3d0ffd7b94af24e1

Observation dcff6084-c81d-47e3-a253-5b8e87ae07eb · inbound

How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization cites this paper.

How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization The large learning rate phase of deep learning: the catapult mechanism

Reference 90

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T04:49:44.826858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-15T04:45:20.091598Z digest=sha256:281fe0c1c4a97b339f3b57c890f83181f60e080153fd4ca334ab765d10e06c2f

Observation 49dd278e-4c23-412f-8a62-8c5246844818 · inbound

Fine-Tuning Without Forgetting via Loss-Adaptive Learning Rates cites this paper.

Fine-Tuning Without Forgetting via Loss-Adaptive Learning Rates The large learning rate phase of deep learning: the catapult mechanism

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:18:07.041459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T07:14:59.396900Z digest=sha256:100e1611269f6d10f56611f703776cc7cc82c86df37315a0a11207b643a33a1d

Observation d622889e-88bb-4c7b-a0d8-efc3a61dd255 · inbound

Large-Step Training Dynamics of a Two-Factor Linear Transformer Model cites this paper.

Large-Step Training Dynamics of a Two-Factor Linear Transformer Model The large learning rate phase of deep learning: the catapult mechanism

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-21T03:53:56.329810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-21T03:50:43.744189Z digest=sha256:7740d130043872e12bd1b7df9b38b877c7dc1ccf682ca3b96a43550b02df7791

Observation a38262b2-cb62-44b6-a9ba-85ffa08037ec · inbound

Accurate Large-sample Uncertainty Quantification using Stochastic Gradient Markov Chain Monte Carlo cites this paper.

Accurate Large-sample Uncertainty Quantification using Stochastic Gradient Markov Chain Monte Carlo The large learning rate phase of deep learning: the catapult mechanism

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:16:01.182666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T22:47:17.603350Z digest=sha256:bbf437c0a95a0811bab63be48a414ac1dbb7ea4de59b6c14ae396df7e62dba82

Observation aea36370-95e3-4e8b-a880-45dff5b07730 · inbound

Edge of Stability Selectively Shapes Learning Across the Data Distribution cites this paper.

Edge of Stability Selectively Shapes Learning Across the Data Distribution The large learning rate phase of deep learning: the catapult mechanism

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:56:29.501282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:27:39.029895Z digest=sha256:cc775e3ffe9a0a19d1f94fa77925c9d24ca35d89cb35db8c2ce3351b72c899df

Observation 5c7b22c6-ee91-4a3a-b8ca-083ab82e857a · inbound

Gradient Descent with Large Step Size Restores Symmetry in Deep Linear Networks with Multi-Pathway cites this paper.

Gradient Descent with Large Step Size Restores Symmetry in Deep Linear Networks with Multi-Pathway The large learning rate phase of deep learning: the catapult mechanism

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:22:46.406450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T23:15:59.394256Z digest=sha256:9c0e61573596b49a8a658e4dcd3f0a099b9ae94fe1e4016031adbe31f7d5a875

Observation 3f38c97f-f7c8-48c7-bcb0-6ee12b2e4c95 · inbound

Edge Flow: A Tractable and Predictive Continuous-Time Model for Gradient Descent at the Edge of Stability cites this paper.

Edge Flow: A Tractable and Predictive Continuous-Time Model for Gradient Descent at the Edge of Stability The large learning rate phase of deep learning: the catapult mechanism

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:28:55.905813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T01:17:37.449377Z digest=sha256:f4b5e47853f8618ac4d0c54aa36d8c841c7b165ce172e9428175d5c2ad333595

Observation af75fc6f-3063-471f-97e7-19aaa4a2d3d2 · inbound

Gradient-Descent Steps to Success over Mean Accuracy: A Paradigm Shift for ML cites this paper.

Gradient-Descent Steps to Success over Mean Accuracy: A Paradigm Shift for ML The large learning rate phase of deep learning: the catapult mechanism

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:49:39.677056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T12:37:25.115951Z digest=sha256:f97bacf44c35fa9f76115ccc11c6a9a0f9049b85c020e833d55d40df01651170

Observation 147cf3b2-b3b7-46b3-a6fb-387931b0fe8e · inbound

Convergence of Gradient Descent for General Neural Network Architectures Beyond the NTK Regime cites this paper.

Convergence of Gradient Descent for General Neural Network Architectures Beyond the NTK Regime The large learning rate phase of deep learning: the catapult mechanism

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:29:44.290800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T08:53:46.285233Z digest=sha256:3f1a87a6e10633aafa56b8d60fdf995fe6cf74b20b688fe5b8a4e4230d3786c5

Observation f5590f2d-b115-4234-8a37-eca57540006c · inbound

Implicit Bias of SGD in Multivariate ReLU Networks: Effective Width Collapse cites this paper.

Implicit Bias of SGD in Multivariate ReLU Networks: Effective Width Collapse The large learning rate phase of deep learning: the catapult mechanism

Reference 189

Resolution
unresolved
no resolver link, observed 2026-07-12T01:11:50.444218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:11:50.444218Z digest=sha256:8ef906988155f1e5e38a479d83db2f1c8da5783588ead4e7dd1b885311fc5382

Observation 99af2c4f-03dd-481a-b5fc-2eeb432f6685 · inbound

Directional Curvature from Armijo Backtracking: A Low-Cost Sharpness Probe and a Calibration-Free Learning-Rate Safeguard for Adam cites this paper.

Directional Curvature from Armijo Backtracking: A Low-Cost Sharpness Probe and a Calibration-Free Learning-Rate Safeguard for Adam The large learning rate phase of deep learning: the catapult mechanism

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-11T22:26:03.646701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:26:03.646701Z digest=sha256:721e0c9a93f4a7417733c6f68d4d51a250bdf624e000aa190c9ca79fad1f2646

Observation 69446ef4-97dc-40ba-8097-aa60a77e02ff · inbound

Directional Curvature from Armijo Backtracking: A Low-Cost Sharpness Probe and a Calibration-Free Learning-Rate Safeguard for Adam cites this paper.

Directional Curvature from Armijo Backtracking: A Low-Cost Sharpness Probe and a Calibration-Free Learning-Rate Safeguard for Adam The large learning rate phase of deep learning: the catapult mechanism

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T16:26:22.217462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:26:22.217462Z digest=sha256:0f36d47d677986494ac9914f7dfa542285828036010987b508e3fb7b618c33b0

Observation bcb3f7be-9112-4054-a632-51be920dd956 · inbound

Directional Curvature from Armijo Backtracking: A Low-Cost Sharpness Probe and a Calibration-Free Learning-Rate Safeguard for Adam cites this paper.

Directional Curvature from Armijo Backtracking: A Low-Cost Sharpness Probe and a Calibration-Free Learning-Rate Safeguard for Adam The large learning rate phase of deep learning: the catapult mechanism

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T08:48:43.532111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:48:43.532111Z digest=sha256:a238d9629767c7e9062557a298879b708abf41d64acd1c750817e23fadf49a11

Observation 8ed05444-e79c-417c-b07c-cc62055f153e · inbound

The Map Behind the Flow: Finite-Step Gradient Descent as a Dynamical System cites this paper.

The Map Behind the Flow: Finite-Step Gradient Descent as a Dynamical System The large learning rate phase of deep learning: the catapult mechanism

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-11T10:24:00.719150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:24:00.719150Z digest=sha256:6c99c438a3e2bb3d2f77dc5258b94542047b6547d4106b504436c17d6d36a64d

Observation 71e7f561-54f2-4a81-ad47-b6e249b99043 · inbound

Weight-norm Criticality: A Mechanism for Loss Spikes Induced by the Normalization and Weight Decay cites this paper.

Weight-norm Criticality: A Mechanism for Loss Spikes Induced by the Normalization and Weight Decay The large learning rate phase of deep learning: the catapult mechanism

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T08:50:41.724748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T08:50:41.724748Z digest=sha256:235f01c33898613a752f7ae16ca5a287210d847911c66f773a1222759df3f4ae

Observation ce33ad15-ee68-4472-b792-648ee5d9b15b · inbound

A Defense of the Quadratic Model cites this paper.

A Defense of the Quadratic Model The large learning rate phase of deep learning: the catapult mechanism

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T06:58:09.772532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:58:09.772532Z digest=sha256:3d068dd041b0d3ffe25dbd6625f8d2ed3617080a03336e3ee3c7db80fe3662f6

Observation 54c77455-4e96-4490-aced-f7b75d138663 · inbound

The Fourth Quadrant: A Stylized View of Benign Misfitting cites this paper.

The Fourth Quadrant: A Stylized View of Benign Misfitting The large learning rate phase of deep learning: the catapult mechanism

Reference 118

Resolution
unresolved
no resolver link, observed 2026-08-06T00:43:14.981584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:43:14.981584Z digest=sha256:18bda9e86450ce0feb8b42656d663a43d4dce0e0d33dbf50056977ddd7a09d68