Pith. sign in

Paper Citation Record · LEDGER

Why Gradients Rapidly Increase Near the End of Training

As of 17 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 10 inbound Pith citation observations for arXiv:2506.02285.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02285 v2

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:32:04.384284Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:27:56.628221Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T05:40:24.372091Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8667ed2e-7c56-4e37-a221-106fbab8bc69 · outbound

This paper cites Layer Normalization.

Why Gradients Rapidly Increase Near the End of Training Layer Normalization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:02.886552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:32:02.886552Z digest=sha256:f56356071bad579ec31fb5ebbeacb48f3843ab45f833e889e9a0bf075f4d7719

Observation 7672aa3e-7fed-494e-ab7b-b28e33ba24e4 · outbound

This paper cites an unresolved cited work.

Why Gradients Rapidly Increase Near the End of Training Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:32:06.594928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:32:02.936151Z digest=sha256:12c6aff14a4666b2848dc360e5bb29e151c7cf727f11ae89399775ac8da23d10

Observation 5531a407-aaf0-4ace-8ac2-0c91d91a7ef6 · outbound

This paper cites an unresolved cited work.

Why Gradients Rapidly Increase Near the End of Training Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:32:06.390649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:32:03.009411Z digest=sha256:461fc9070d7ebdda80542de40b93adff00e7d4b6d24bc12cb35e737acbafd979

Observation c6c93d21-9bbd-48e5-bd22-a63995c07b03 · outbound

This paper cites an unresolved cited work.

Why Gradients Rapidly Increase Near the End of Training Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:32:06.122567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:32:03.095688Z digest=sha256:31a244e51e65278c44d4f600726aacb1fe94ee107fff74c30a2c84f1b7a48cb8

Observation 08db36d8-a0df-4c7c-8936-ef46e243c71d · outbound

This paper cites and Bottou, L.

Why Gradients Rapidly Increase Near the End of Training and Bottou, L

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:32:05.877245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:32:03.175874Z digest=sha256:5248d0baebccaab949c926e33da023a84ffed170a4b5ff92a510ac878ca97fe7

Observation 4b0b303b-89ab-4fd6-bd30-7809a85fbde7 · outbound

This paper cites an unresolved cited work.

Why Gradients Rapidly Increase Near the End of Training Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:32:05.627056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:32:03.230459Z digest=sha256:74befedf0f08ef4f90a5fa789354ca0fed75f1c0830a133c928597042482fc60

Observation 2be33ce0-a8e0-4f4e-a52d-2e5194feacd1 · outbound

This paper cites and Gower, R.

Why Gradients Rapidly Increase Near the End of Training and Gower, R

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:03.275141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:32:03.275141Z digest=sha256:c20567f307c6272d47661a7f9aa0e904c033774c8207603e965cee40123bd986

Observation 9a60bbc3-005f-40d8-9af8-f4a85fee0872 · outbound

This paper cites and Mishchenko, K.

Why Gradients Rapidly Increase Near the End of Training and Mishchenko, K

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:03.377669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:32:03.377669Z digest=sha256:e172a10740f0479deeb04b2bbadfcc8dcc6bbc95b799da1c8163cb03ab558c29

Observation 4f5789c7-82a3-4da5-974f-41a7fc2e7191 · outbound

This paper cites an unresolved cited work.

Why Gradients Rapidly Increase Near the End of Training Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:32:05.400559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:32:03.474851Z digest=sha256:e56f22dd2f71cae37d3a14aea90706eb2d2070151fad3740406500f2ff0398ab

Observation 27c72ec4-65e5-4f5a-a0ea-94f42066bec4 · outbound

This paper cites an unresolved cited work.

Why Gradients Rapidly Increase Near the End of Training Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:32:05.334360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:32:03.579676Z digest=sha256:df8a9e72fdf654c82b6770a641403de0a049b41f2fba5f7f42e366df94d88cc6

Observation ef705d8e-9ab2-4a36-a65c-b673f8576b72 · outbound

This paper cites and Szegedy, C.

Why Gradients Rapidly Increase Near the End of Training and Szegedy, C

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:32:05.238177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:32:03.676898Z digest=sha256:45cd763d82f4492ec37b0954d47951b0ccc3fbf54ae46713a4fcb1bc63f412ac

Observation 9263c93b-63c6-4827-a537-8a53c4918d37 · outbound

This paper cites an unresolved cited work.

Why Gradients Rapidly Increase Near the End of Training Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:03.723875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:32:03.723875Z digest=sha256:dc1afc012d5f5eb15712ab5a15493853540982c337ab5e9fb33aa3cf4e003fb6

Observation 636ae290-6d76-4cf3-b12e-1e9760ad81d5 · outbound

This paper cites and Hutter, F.

Why Gradients Rapidly Increase Near the End of Training and Hutter, F

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:03.811301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:32:03.811301Z digest=sha256:2416c5caf8552dd4d5b179d61623615efb473786c026453c138b24fbf4797b26

Observation 4611d213-a767-43d1-a745-7f4e4bce64e4 · outbound

This paper cites Online Learning: A Modern Introduction Using Convex Optimization.

Why Gradients Rapidly Increase Near the End of Training Online Learning: A Modern Introduction Using Convex Optimization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:03.877272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:32:03.877272Z digest=sha256:27716f4209a4d1b6c22922332ff90abe6e226f63bd43056b9e2f28bbfffbfe5e

Observation 28337d32-f3c9-4675-ba61-592f81b7276a · outbound

This paper cites B., Lozhkov, A., Mitchell, M., Raffel, C., Werra, L.

Why Gradients Rapidly Increase Near the End of Training B., Lozhkov, A., Mitchell, M., Raffel, C., Werra, L

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:32:05.093205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:32:03.915870Z digest=sha256:4f8e01b4650e8b3d359162f7d8559130ecc0386351ca6dd3597ab2ba17d4c27b

Observation 66ba23e5-347c-45bc-af2c-85bb6040e140 · outbound

This paper cites C., and Fei-Fei, L.

Why Gradients Rapidly Increase Near the End of Training C., and Fei-Fei, L

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:32:04.951450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:32:03.994499Z digest=sha256:a0f3783b99586453ccb58df0dfa383e9a46360335a0d035d15c6a61ff24acc36

Observation 817e3174-5ca4-4ea4-81ba-0167e048b52c · outbound

This paper cites an unresolved cited work.

Why Gradients Rapidly Increase Near the End of Training Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:32:04.884836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:32:04.050469Z digest=sha256:161814f2ab25c07e5682b6268c5e2407ca36459e903f8c5964c9a552ea2aef60

Observation 7c43b120-6677-4a8d-b09a-739151b93b4c · outbound

This paper cites L2 Regularization versus Batch and Weight Normalization.

Why Gradients Rapidly Increase Near the End of Training L2 Regularization versus Batch and Weight Normalization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:04.103673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:32:04.103673Z digest=sha256:3c93ea428ffa6df4a462c9b211028a22affbe2933c2deb68196fcb13705cc1f5

Observation 69a1258c-d2bc-4057-a133-b27204da1a16 · outbound

This paper cites Implicit Bias of AdamW: $\ell_\infty$ Norm Constrained Optimization.

Why Gradients Rapidly Increase Near the End of Training Implicit Bias of AdamW: $\ell_\infty$ Norm Constrained Optimization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:04.192151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:32:04.192151Z digest=sha256:32dbeb873618418c28af31edd4387d4ec8305fd875109f927449892d3db6efd6

Observation 6b97a015-9c9b-499f-b25b-f1fbb9092572 · outbound

This paper cites an unresolved cited work.

Why Gradients Rapidly Increase Near the End of Training Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:32:04.764571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:32:04.246337Z digest=sha256:48a40410794ef712947686649d602a3dc582ab3e22e26447ab178d4f9b22cf13

Observation 7db33698-d989-43f0-955a-82cd872ea666 · outbound

This paper cites an unresolved cited work.

Why Gradients Rapidly Increase Near the End of Training Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:32:04.653154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:32:04.332003Z digest=sha256:b18390e361ed746575748a71e56cc8073d6b858df9a042690c2db2256ad97530

Observation 4e3c916a-4668-49dc-9a72-591cd53e0501 · outbound

This paper cites an unresolved cited work.

Why Gradients Rapidly Increase Near the End of Training Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:32:04.541396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T11:32:04.384284Z digest=sha256:2637b2d923472220e358901bb69011cf0f3a653c86d9c0a3b5c947135b0e8e06

Pith citing papers

Observation 57c806d0-96f4-4a92-99f6-aedbce155fcc · inbound

Why Do We Need Warm-up? A Theoretical Perspective cites this paper.

Why Do We Need Warm-up? A Theoretical Perspective Why Gradients Rapidly Increase Near the End of Training

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:53.777630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:53.777630Z digest=sha256:9ced6623c15037d34ee809dbf7793fe92642c643acc6ab261b1a6eb0db4f730c

Observation 574e5e0b-156e-4985-8126-dcf919846b31 · inbound

Safeguarded Stochastic Polyak Step Sizes for Non-smooth Optimization: Robust Performance Without Small (Sub)Gradients cites this paper.

Safeguarded Stochastic Polyak Step Sizes for Non-smooth Optimization: Robust Performance Without Small (Sub)Gradients Why Gradients Rapidly Increase Near the End of Training

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-03T19:10:48.092792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:10:48.092792Z digest=sha256:e8f31380f4333ad63d0b40e35da313179319162090358afe2a89739ed3ba710e

Observation 4f6d60b2-9096-4fc1-ae1f-5a2baabaab1e · inbound

Rethinking Language Model Scaling under Transferable Hypersphere Optimization cites this paper.

Rethinking Language Model Scaling under Transferable Hypersphere Optimization Why Gradients Rapidly Increase Near the End of Training

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:53:02.525835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-14T21:51:14.678941Z digest=sha256:58eca4bdfca090738e89017ce6be0adb73a68c52aa17b3ab837808317b60c2b5

Observation f0d07f9d-136b-4fa3-8284-a26bcbc926ab · inbound

Broximal Alignment for Global Non-Convex Optimization cites this paper.

Broximal Alignment for Global Non-Convex Optimization Why Gradients Rapidly Increase Near the End of Training

Reference 1

Resolution
malformed identifier
arxiv_id, observed 2026-05-10T13:00:24.674807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T12:57:52.201675Z digest=sha256:d9168f0ae66a1653aa437f769fae9c36f243216558c208a02098addf76346514

Observation 366a7266-7668-49cf-a79c-f180b3c32392 · inbound

Demystifying Manifold Constraints in LLM Pre-training cites this paper.

Demystifying Manifold Constraints in LLM Pre-training Why Gradients Rapidly Increase Near the End of Training

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:16:08.661283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T17:44:44.438637Z digest=sha256:f6be51150b7948f16ebe2a0ebcb801d976ec000db20ab9053c9fb6cc6c6cca5c

Observation 2c37b115-b48a-4f6b-83e9-cfa5e1166495 · inbound

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less cites this paper.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Why Gradients Rapidly Increase Near the End of Training

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.465260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:8f240c61fbd61d82f715fb20a7bdaa5bea2914529180a7ffdd1d5212d203801d

Observation 41ab2342-e1d9-4e23-96d5-104b291b08d0 · inbound

Anytime Training with Schedule-Free Spectral Optimization cites this paper.

Anytime Training with Schedule-Free Spectral Optimization Why Gradients Rapidly Increase Near the End of Training

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:40:24.375625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-25T05:38:16.958574Z digest=sha256:698a1fe9b8a2ea5f542af0ec81cac8bca9ed85dbf081efddf6275f09cc7b63e0

Observation 75168a28-59d3-473b-a6c4-d52b80b9be66 · inbound

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers cites this paper.

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers Why Gradients Rapidly Increase Near the End of Training

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-11T22:10:49.683444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:10:49.683444Z digest=sha256:ca7579b5ec9d1064118de9746a5ec9d527b02148a090504183def014f6f9c316

Observation cf743017-35ec-4b6e-a066-38f82f0bb173 · inbound

Scale Weight Decay and Train Better cites this paper.

Scale Weight Decay and Train Better Why Gradients Rapidly Increase Near the End of Training

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:40.954445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:40.954445Z digest=sha256:f3be84eb93919ab4398673e9ac786927e96a6ab3ba52a54c8e64ca418ad9a52a

Observation adbc6f64-beda-4a03-9c7f-03f92a182b1a · inbound

Full-bandwidth transformer cites this paper.

Full-bandwidth transformer Why Gradients Rapidly Increase Near the End of Training

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:56.628221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:27:56.628221Z digest=sha256:323bba50719ba3d8e609a6890b5bfe3f959aea941b52039506de12a013204a9a