Pith. sign in

Paper Citation Record · LEDGER

Why Gradients Rapidly Increase Near the End of Training

As of 16 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 10 inbound Pith citation observations for arXiv:2506.02285.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02285 v2

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:32:04.384284Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:27:56.628221Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T05:40:24.372091Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8667ed2e-7c56-4e37-a221-106fbab8bc69 · outbound

This paper cites Layer Normalization.

Why Gradients Rapidly Increase Near the End of Training Layer Normalization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:02.886552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:32:02.886552Z digest=sha256:f56356071bad579ec31fb5ebbeacb48f3843ab45f833e889e9a0bf075f4d7719

Observation 7672aa3e-7fed-494e-ab7b-b28e33ba24e4 · outbound

This paper cites an unresolved cited work.

Why Gradients Rapidly Increase Near the End of Training Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:32:06.594928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:32:02.936151Z digest=sha256:371990df2e934342d673dcdcd3ab115bee4cff479bb63c4991e65dcb651ec933

Observation 5531a407-aaf0-4ace-8ac2-0c91d91a7ef6 · outbound

This paper cites an unresolved cited work.

Why Gradients Rapidly Increase Near the End of Training Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:32:06.390649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:32:03.009411Z digest=sha256:f62817f72f235a7dbed30fb2b3daa4c1d57ebd11f303488414c3d67377eb08e3

Observation c6c93d21-9bbd-48e5-bd22-a63995c07b03 · outbound

This paper cites an unresolved cited work.

Why Gradients Rapidly Increase Near the End of Training Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:32:06.122567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:32:03.095688Z digest=sha256:f37bb134aeade59775e898d111d07b713e0e17effe8337e95ee2fd82aa8ddd55

Observation 08db36d8-a0df-4c7c-8936-ef46e243c71d · outbound

This paper cites and Bottou, L.

Why Gradients Rapidly Increase Near the End of Training and Bottou, L

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:32:05.877245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:32:03.175874Z digest=sha256:da1f5c108d133fb34ae48c6fe6723d5ee2b1a388b0e429c6b2adf07f4ead6b12

Observation 4b0b303b-89ab-4fd6-bd30-7809a85fbde7 · outbound

This paper cites an unresolved cited work.

Why Gradients Rapidly Increase Near the End of Training Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:32:05.627056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:32:03.230459Z digest=sha256:0d9aa831a0e886202ff4e16bd62c7cc74da838c72bb199caaf7ae2d2be1e4daa

Observation 2be33ce0-a8e0-4f4e-a52d-2e5194feacd1 · outbound

This paper cites and Gower, R.

Why Gradients Rapidly Increase Near the End of Training and Gower, R

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:03.275141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:32:03.275141Z digest=sha256:c20567f307c6272d47661a7f9aa0e904c033774c8207603e965cee40123bd986

Observation 9a60bbc3-005f-40d8-9af8-f4a85fee0872 · outbound

This paper cites and Mishchenko, K.

Why Gradients Rapidly Increase Near the End of Training and Mishchenko, K

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:03.377669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:32:03.377669Z digest=sha256:e172a10740f0479deeb04b2bbadfcc8dcc6bbc95b799da1c8163cb03ab558c29

Observation 4f5789c7-82a3-4da5-974f-41a7fc2e7191 · outbound

This paper cites an unresolved cited work.

Why Gradients Rapidly Increase Near the End of Training Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:32:05.400559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:32:03.474851Z digest=sha256:152376e2ee704d008a1f12ae32f884440480deb9536fa8af1bd2f6a2c60c8d3c

Observation 27c72ec4-65e5-4f5a-a0ea-94f42066bec4 · outbound

This paper cites an unresolved cited work.

Why Gradients Rapidly Increase Near the End of Training Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:32:05.334360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:32:03.579676Z digest=sha256:ea17aa65438945ca3de65f63a9ebbbdf0ba368c4884053a776a78d48649db2c6

Observation ef705d8e-9ab2-4a36-a65c-b673f8576b72 · outbound

This paper cites and Szegedy, C.

Why Gradients Rapidly Increase Near the End of Training and Szegedy, C

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:32:05.238177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:32:03.676898Z digest=sha256:fdb88d49920787f53ae6d3cd74df25811b07400725854a05eb36f48dad381d2d

Observation 9263c93b-63c6-4827-a537-8a53c4918d37 · outbound

This paper cites an unresolved cited work.

Why Gradients Rapidly Increase Near the End of Training Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:03.723875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:32:03.723875Z digest=sha256:dc1afc012d5f5eb15712ab5a15493853540982c337ab5e9fb33aa3cf4e003fb6

Observation 636ae290-6d76-4cf3-b12e-1e9760ad81d5 · outbound

This paper cites and Hutter, F.

Why Gradients Rapidly Increase Near the End of Training and Hutter, F

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:03.811301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:32:03.811301Z digest=sha256:2416c5caf8552dd4d5b179d61623615efb473786c026453c138b24fbf4797b26

Observation 4611d213-a767-43d1-a745-7f4e4bce64e4 · outbound

This paper cites Online Learning: A Modern Introduction Using Convex Optimization.

Why Gradients Rapidly Increase Near the End of Training Online Learning: A Modern Introduction Using Convex Optimization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:03.877272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:32:03.877272Z digest=sha256:27716f4209a4d1b6c22922332ff90abe6e226f63bd43056b9e2f28bbfffbfe5e

Observation 28337d32-f3c9-4675-ba61-592f81b7276a · outbound

This paper cites B., Lozhkov, A., Mitchell, M., Raffel, C., Werra, L.

Why Gradients Rapidly Increase Near the End of Training B., Lozhkov, A., Mitchell, M., Raffel, C., Werra, L

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:32:05.093205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:32:03.915870Z digest=sha256:cb44be9c2fed4ef4d43ac8dea886e2bfcf89c2aa89c8014af258e3798c3527d7

Observation 66ba23e5-347c-45bc-af2c-85bb6040e140 · outbound

This paper cites C., and Fei-Fei, L.

Why Gradients Rapidly Increase Near the End of Training C., and Fei-Fei, L

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:32:04.951450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:32:03.994499Z digest=sha256:4062d69918c84151fbed0265e77d68ac9261ccac32bd73e898174239acec8ec9

Observation 817e3174-5ca4-4ea4-81ba-0167e048b52c · outbound

This paper cites an unresolved cited work.

Why Gradients Rapidly Increase Near the End of Training Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:32:04.884836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:32:04.050469Z digest=sha256:3f1c417e60e4f3800dfe4f331152f3b26ff6e52bab6264d1df419730152ae3c6

Observation 7c43b120-6677-4a8d-b09a-739151b93b4c · outbound

This paper cites L2 Regularization versus Batch and Weight Normalization.

Why Gradients Rapidly Increase Near the End of Training L2 Regularization versus Batch and Weight Normalization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:04.103673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:32:04.103673Z digest=sha256:37296b5a2c92882b3b44d36b941f0c39dbd20faa5b3e6fc7bdfe47472ff926e2

Observation 69a1258c-d2bc-4057-a133-b27204da1a16 · outbound

This paper cites Implicit Bias of AdamW: $\ell_\infty$ Norm Constrained Optimization.

Why Gradients Rapidly Increase Near the End of Training Implicit Bias of AdamW: $\ell_\infty$ Norm Constrained Optimization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:04.192151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:32:04.192151Z digest=sha256:11f0f53bdfe17d0d5b1938eebac507ab017b1fc206538935d5bc239486b49c0a

Observation 6b97a015-9c9b-499f-b25b-f1fbb9092572 · outbound

This paper cites an unresolved cited work.

Why Gradients Rapidly Increase Near the End of Training Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:32:04.764571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:32:04.246337Z digest=sha256:306bde93aec813964fe291a8d53e366a5510e13bddb00dc14d61be93571530b7

Observation 7db33698-d989-43f0-955a-82cd872ea666 · outbound

This paper cites an unresolved cited work.

Why Gradients Rapidly Increase Near the End of Training Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:32:04.653154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:32:04.332003Z digest=sha256:039f4267568d93d4766237c21ba4a0e63afc5e2020d1b543566466875d722ccf

Observation 4e3c916a-4668-49dc-9a72-591cd53e0501 · outbound

This paper cites an unresolved cited work.

Why Gradients Rapidly Increase Near the End of Training Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:32:04.541396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:32:04.384284Z digest=sha256:79d768d8467d8959753f5c8bf24107d96705157d6f6a938e41ea0d374b33d500

Pith citing papers

Observation 57c806d0-96f4-4a92-99f6-aedbce155fcc · inbound

Why Do We Need Warm-up? A Theoretical Perspective cites this paper.

Why Do We Need Warm-up? A Theoretical Perspective Why Gradients Rapidly Increase Near the End of Training

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:53.777630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:53.777630Z digest=sha256:9ced6623c15037d34ee809dbf7793fe92642c643acc6ab261b1a6eb0db4f730c

Observation 574e5e0b-156e-4985-8126-dcf919846b31 · inbound

Safeguarded Stochastic Polyak Step Sizes for Non-smooth Optimization: Robust Performance Without Small (Sub)Gradients cites this paper.

Safeguarded Stochastic Polyak Step Sizes for Non-smooth Optimization: Robust Performance Without Small (Sub)Gradients Why Gradients Rapidly Increase Near the End of Training

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-03T19:10:48.092792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:10:48.092792Z digest=sha256:e8f31380f4333ad63d0b40e35da313179319162090358afe2a89739ed3ba710e

Observation 4f6d60b2-9096-4fc1-ae1f-5a2baabaab1e · inbound

Rethinking Language Model Scaling under Transferable Hypersphere Optimization cites this paper.

Rethinking Language Model Scaling under Transferable Hypersphere Optimization Why Gradients Rapidly Increase Near the End of Training

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:53:02.525835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-14T21:51:14.678941Z digest=sha256:969cb445669af3fcbc2877628bf4a97d07eeddffa9d0173ddfbb51489e864c4a

Observation f0d07f9d-136b-4fa3-8284-a26bcbc926ab · inbound

Broximal Alignment for Global Non-Convex Optimization cites this paper.

Broximal Alignment for Global Non-Convex Optimization Why Gradients Rapidly Increase Near the End of Training

Reference 1

Resolution
malformed identifier
arxiv_id, observed 2026-05-10T13:00:24.674807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T12:57:52.201675Z digest=sha256:146e1cef65f27aca6aec3e50ce3ada99622d80f7e04df6c83ac7cec20d0d5bf4

Observation 366a7266-7668-49cf-a79c-f180b3c32392 · inbound

Demystifying Manifold Constraints in LLM Pre-training cites this paper.

Demystifying Manifold Constraints in LLM Pre-training Why Gradients Rapidly Increase Near the End of Training

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:16:08.661283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T17:44:44.438637Z digest=sha256:ee889e7f750244be9c14d28f3801f89d53becb609de62421776c461eebd89b4b

Observation 2c37b115-b48a-4f6b-83e9-cfa5e1166495 · inbound

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less cites this paper.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Why Gradients Rapidly Increase Near the End of Training

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.465260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:ab3c11f6c345f41f771a138f0b92e3118def743d43aeb4cb71928a389b15f18b

Observation 41ab2342-e1d9-4e23-96d5-104b291b08d0 · inbound

Anytime Training with Schedule-Free Spectral Optimization cites this paper.

Anytime Training with Schedule-Free Spectral Optimization Why Gradients Rapidly Increase Near the End of Training

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:40:24.375625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-25T05:38:16.958574Z digest=sha256:ed32eeeaa9b80d66324592b74e46c3430d248dbeff7b5b5c773f4557143a65d4

Observation 75168a28-59d3-473b-a6c4-d52b80b9be66 · inbound

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers cites this paper.

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers Why Gradients Rapidly Increase Near the End of Training

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-11T22:10:49.683444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:10:49.683444Z digest=sha256:ca7579b5ec9d1064118de9746a5ec9d527b02148a090504183def014f6f9c316

Observation cf743017-35ec-4b6e-a066-38f82f0bb173 · inbound

Scale Weight Decay and Train Better cites this paper.

Scale Weight Decay and Train Better Why Gradients Rapidly Increase Near the End of Training

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.954445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.954445Z digest=sha256:f3be84eb93919ab4398673e9ac786927e96a6ab3ba52a54c8e64ca418ad9a52a

Observation adbc6f64-beda-4a03-9c7f-03f92a182b1a · inbound

Full-bandwidth transformer cites this paper.

Full-bandwidth transformer Why Gradients Rapidly Increase Near the End of Training

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:56.628221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:27:56.628221Z digest=sha256:323bba50719ba3d8e609a6890b5bfe3f959aea941b52039506de12a013204a9a