Pith. sign in

Paper Citation Record · LEDGER

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence

As of 13 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 3 inbound Pith citation observations for arXiv:2509.07972.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.07972 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T21:32:36.011814Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T12:38:58.891916Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T20:15:04.523466Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact1
  • verified fuzzy22
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d4718808-baa3-4b27-ac4b-061ff6581a49 · outbound

This paper cites Qsgd: Communication- Efficient SGD via Gradient Quantization and Encoding.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Qsgd: Communication- Efficient SGD via Gradient Quantization and Encoding

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:41.308989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-04T21:32:32.871158Z digest=sha256:ee62a867a73e411a86eba635bca4640a853c2c45831986f5db445211d789fe9d

Observation dc6aff50-060d-44e7-9c7a-b86e3e5caf38 · outbound

This paper cites Duchi, Dylan J.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Duchi, Dylan J

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:40.966682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-04T21:32:32.970569Z digest=sha256:f7c82b97f43781faf6afd014b8ed4c1095af9112922ce738a7714cef7e1f103d

Observation ca417af7-faa8-406f-bd36-1b52e48d41ac · outbound

This paper cites Curtis, and Jorge Nocedal.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Curtis, and Jorge Nocedal

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:40.749870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-04T21:32:33.084868Z digest=sha256:4394dea20cb4bdd0c0c3585ca4bc1e1acdcfb3f300a91fdc493fb5e8eb42de07

Observation ec9e8a2e-59b2-4cf0-993b-e3cbc9711f6b · outbound

This paper cites Convex optimization.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Convex optimization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:33.151389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:33.151389Z digest=sha256:4e5aa82180420d41ffc5be49a795c788934cfcf9bacb979dcd337bfcc09fad65

Observation 9b97a3c6-abf5-4c56-9ca9-c68d227c8443 · outbound

This paper cites Gradient descent on neural networks typically occurs at the edge of stability.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Gradient descent on neural networks typically occurs at the edge of stability

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:40.409316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-04T21:32:33.230781Z digest=sha256:bb2041464c17f57532f5481840a9b07b3f69ce3f2cb1eda6913f8323904e7984

Observation 4ee23d18-6703-4cd9-ac07-1b29ba8e725c · outbound

This paper cites Robustness to unbounded smoothness of generalized signsgd.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Robustness to unbounded smoothness of generalized signsgd

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:40.183963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-04T21:32:33.317211Z digest=sha256:d0b31c9ffdc604b2819e482e2eb9cc18be53ad548e761445814b626cd64932b5

Observation ff52d472-7ba2-4df3-b8ab-a87f8c73c2f9 · outbound

This paper cites A loss curvature perspective on training instabilities of deep learning models.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence A loss curvature perspective on training instabilities of deep learning models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:39.910840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-04T21:32:33.412706Z digest=sha256:3a591217ffef7ba8077657deb6341a134d2b85ef7797f781d434ecbb52d795f9

Observation 6ebaf8dc-8753-4f3b-817b-50221d9502ab · outbound

This paper cites A Closer Look at Deep Learning Heuristics: Learning rate restarts, Warmup and Distillation.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence A Closer Look at Deep Learning Heuristics: Learning rate restarts, Warmup and Distillation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:33.488985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:33.488985Z digest=sha256:cd54f7c3e493b4a1d1dfbff9c5134b6d7c5b83491744ef7903c3ca9313774413

Observation ff8a23f2-32d0-4b8f-b34d-cdcf5d9b0fcd · outbound

This paper cites A closer look at deep learning heuristics: Learning rate restarts, warmup and distillation.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence A closer look at deep learning heuristics: Learning rate restarts, warmup and distillation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:39.688153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-04T21:32:33.601486Z digest=sha256:5a098ec2ce11aa7c9431c80711a3c144eca919bbed7bc033112c03fc0b92d569

Observation cbcc7641-dd17-465f-87d8-f237f7ed0fbc · outbound

This paper cites SGD: General Analysis and Improved Rates.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence SGD: General Analysis and Improved Rates

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:33.743464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:33.743464Z digest=sha256:a332f8b7c89411298321fef43bc0d0e17896bfaaa673449731ddd1b142ca6d4e

Observation d984b7e2-2db1-4594-be0e-ee657e728883 · outbound

This paper cites Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:33.855290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:33.855290Z digest=sha256:c87a1fff42da74f5444b890ec4388cb088942f7e1852441db9e19036178ea792

Observation 7474a037-c114-41b5-a4b6-1033745c136a · outbound

This paper cites Deep residual learning for image recognition.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Deep residual learning for image recognition

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:33.998114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:33.998114Z digest=sha256:120eb0bcbad40605fdefaac790779e1f0fb2af72558ab206e1131edeecd042cb

Observation 98a5cef4-ef28-45f2-9dd9-f922df3f1c56 · outbound

This paper cites Three factors influencing minima in SGD , 2018.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Three factors influencing minima in SGD , 2018

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:39.329972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-04T21:32:34.144952Z digest=sha256:cbfd4e02c9a8029dd8bb95b533641936531fb72844dd7e2aef50a6162c33d81e

Observation b47f25e8-7dcc-479d-8228-28461a978702 · outbound

This paper cites Why warmup the learning rate? underlying mechanisms and improvements.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Why warmup the learning rate? underlying mechanisms and improvements

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:39.070350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-04T21:32:34.271114Z digest=sha256:80c1b51efc6784d15694cde4276a4e315a31387dd029e48250312084b5ce7fc0

Observation b7c5919c-8153-4c8a-b900-931a12fdc66c · outbound

This paper cites Better Theory for SGD in the Nonconvex World.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Better Theory for SGD in the Nonconvex World

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:38.824116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-04T21:32:34.414677Z digest=sha256:6f07f55be55627b475a767828f7802f039c705d05ced5ec74c58ea9442b3745a

Observation afccdaec-29b4-4ae4-b533-bbf112fe3a6f · outbound

This paper cites Distributed learning with compressed gradients.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Distributed learning with compressed gradients

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-04T21:32:36.237186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-04T21:32:34.555629Z digest=sha256:72e239471c447afb7a91e099fbaf55470b39016985d1dd1ca3049a378a9c3c2e

Observation d3180bf2-e1fc-4b2d-82f1-9378c2df9ebb · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Adam: A Method for Stochastic Optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:34.617678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:34.617678Z digest=sha256:44af71bd5627566e89d8e64fd571f71d33ddc0e86e12d0799f1f32ff316ffcc5

Observation 07c72ed4-452d-47d7-8528-37717622a62c · outbound

This paper cites Analyzing & reducing the need for learning rate warmup in gpt training.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Analyzing & reducing the need for learning rate warmup in gpt training

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:38.616868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-04T21:32:34.680021Z digest=sha256:33c1e82f41cf86064b018df758acfc9b430848e5d5920c5968936eee0db4056b

Observation 314450eb-82f1-4de0-9886-746ee0df8559 · outbound

This paper cites Convex and Non -convex Optimization Under Generalized Smoothness.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Convex and Non -convex Optimization Under Generalized Smoothness

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:38.378429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-04T21:32:34.742379Z digest=sha256:22a72aee5874215118344cb6f42562f188b301765d5f82a31a6b5a3bdf4ab0d5

Observation 4ef6a172-b240-4d5d-8efd-4f75da586e6b · outbound

This paper cites Convergence of Adam Under Relaxed Assumptions.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Convergence of Adam Under Relaxed Assumptions

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:38.127455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-04T21:32:34.805621Z digest=sha256:642aa9309ac24b8ce538bcbf92e6ded5c743050e5b9163b84b58d236d0fe224f

Observation 8f0ad54b-13d1-488d-9177-1fb185bedf73 · outbound

This paper cites On the Variance of the Adaptive Learning Rate and Beyond.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence On the Variance of the Adaptive Learning Rate and Beyond

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:37.879480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-04T21:32:34.892504Z digest=sha256:c21aa5dbeacb99dabafd1117a279f9adc9032e7e32cc55831ac2f83209e31694

Observation 95096954-aea3-4c7a-9389-6f4fe541cb3c · outbound

This paper cites AdaGrad under Anisotropic Smoothness.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence AdaGrad under Anisotropic Smoothness

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:34.950133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:34.950133Z digest=sha256:1f50f65b2698f2df34bc853909ea1ac7f269272e981b34f24be7a1107825818e

Observation 5c9a1aa6-d925-4778-a765-04781f2f3c79 · outbound

This paper cites Revisiting the last-iterate convergence of stochastic gradient methods.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Revisiting the last-iterate convergence of stochastic gradient methods

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:35.018805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:35.018805Z digest=sha256:3517cfcbedb38982dcaa17ec119909ce1f7bb9a949609a5abf38621337402627

Observation fc032b57-383b-4fde-906c-ae001ac6363c · outbound

This paper cites SGDR : Stochastic gradient descent with warm restarts.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence SGDR : Stochastic gradient descent with warm restarts

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:35.082983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:35.082983Z digest=sha256:ca39578ceb6fe424a4ed721db54ccf03c81ea699e9f11106cdbd7b4e6db9a84b

Observation 06ea6c50-4a2e-49d3-b6bc-fbbec7f9a9ad · outbound

This paper cites Adaptive Gradient Descent without Descent.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Adaptive Gradient Descent without Descent

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:37.598245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-04T21:32:35.141570Z digest=sha256:684fca941b15e0de230911cf0a1d87843def503c06bf62332190afc33d52f166

Observation 0d3c2c69-fdca-4b8f-afcf-3954453339b0 · outbound

This paper cites Lectures on convex optimization, volume 137.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Lectures on convex optimization, volume 137

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:35.205786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:35.205786Z digest=sha256:68d111907bdf70db0b2193a5668d9a329c68891625cb3dc251df699bc9f689aa

Observation d1d7617d-4e8b-4f70-bdbf-b8be3886a8d8 · outbound

This paper cites Global Convergence and Stability of Stochastic Gradient Descent.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Global Convergence and Stability of Stochastic Gradient Descent

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:37.355620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-04T21:32:35.260121Z digest=sha256:229f45c70bce6dd4aa2b85c70158356f349b14c927a7cb6dcfb389f39e7723ec

Observation 2604dd10-241a-420a-9448-984bf0d78c36 · outbound

This paper cites Understanding Gradient Clipping In Incremental Gradient Methods.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Understanding Gradient Clipping In Incremental Gradient Methods

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:37.177582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-04T21:32:35.337374Z digest=sha256:79bae07cae6bbec7f4d9be57bd5981ec15fda720fc062a80965d7109c09cd2f3

Observation a48916cd-bc8a-49db-94b1-08c7f8f83388 · outbound

This paper cites Smith, Pieter-Jan Kindermans, and Quoc V.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Smith, Pieter-Jan Kindermans, and Quoc V

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:37.062095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-04T21:32:35.416029Z digest=sha256:a3f13da443d2a07888ef6ae85e621b536c66048817f9a9da988e1ea80fc3fe1c

Observation cbb4a111-831c-45db-83f1-ce519472d5c2 · outbound

This paper cites An elementary approach to tight worst case complexity analysis of gradient based methods.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence An elementary approach to tight worst case complexity analysis of gradient based methods

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:36.979005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-04T21:32:35.483322Z digest=sha256:9feda99da0b6be28e7778ec47172bb36583017eea19e1b39eaad5ac4c313dc1a

Observation b6f8f43b-f42d-48b6-99c4-75379aa3e0fd · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:35.560961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:35.560961Z digest=sha256:b4dd19472ba1426eb3b3afd042de49ecac9cf42bdddf85c87db6857b8a77d0c3

Observation 3b624295-8e1f-4767-99bf-3ffbf28c1104 · outbound

This paper cites Toward a Unified Theory of Gradient Descent under Generalized Smoothness.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Toward a Unified Theory of Gradient Descent under Generalized Smoothness

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:36.794213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-04T21:32:35.645274Z digest=sha256:924df5c365d727a0a95870331a304656319447ef83e29985b6327137c52561bf

Observation 55ec8de5-bbd5-4992-8004-caac4d780bb8 · outbound

This paper cites Attention is all you need.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Attention is all you need

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:35.738963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:35.738963Z digest=sha256:1c095f21d8ebddae5db13a305e125e1be94d9a25347db9c065c47d260983e060

Observation 3e405299-3eb1-4fca-9f5b-0bc3331cc7b8 · outbound

This paper cites Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:35.801321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:35.801321Z digest=sha256:18c94cfc77a2201ad2421cef9a0d2da45d9791bb5c59ce7f7949784063d9e824

Observation 099ab271-6b10-4276-9f52-cfbd1bd47165 · outbound

This paper cites Improved analysis of clipping algorithms for non-convex optimization.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Improved analysis of clipping algorithms for non-convex optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T21:32:35.878410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:32:35.878410Z digest=sha256:2733d3f0805b58c2447306201734d901a86e3c24110087a4f47e2d431d875887

Observation 34286f8d-6728-41e9-8beb-e90599e97273 · outbound

This paper cites Why Gradient Clipping Accelerates Training : A Theoretical Justification for Adaptivity.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence Why Gradient Clipping Accelerates Training : A Theoretical Justification for Adaptivity

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:36.561393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-04T21:32:35.927559Z digest=sha256:3eabd5fb86d27b4e1192953be088ae53e45dff7603c485e0d70ad1b05d130cc2

Observation 4b4e9247-268c-4084-a859-25ca2fecfb87 · outbound

This paper cites On the convergence and improvement of stochastic normalized gradient descent.

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence On the convergence and improvement of stochastic normalized gradient descent

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:32:36.404112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-04T21:32:36.011814Z digest=sha256:ca7d8fe76bdb55f8ab77d8ea2def091e045525593307ef4ac737c7cbfaa268ba

Pith citing papers

Observation 89caaa2e-7bee-452a-9079-db1bb129da9e · inbound

Why Do We Need Warm-up? A Theoretical Perspective cites this paper.

Why Do We Need Warm-up? A Theoretical Perspective Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:58.891916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:58.891916Z digest=sha256:7690d21ebf73e443f56b42dcb3291a6b7a8b3ef859e490e249f7dd98c59b33c2

Observation 1a80413b-21af-4eee-9d4f-a1b31709a47f · inbound

A Provably Robust Multi-Jet Framework applied to Active Flow Control of an Airfoil in Weakly Compressible Flow cites this paper.

A Provably Robust Multi-Jet Framework applied to Active Flow Control of an Airfoil in Weakly Compressible Flow Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:26:25.972480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-07T10:57:04.360455Z digest=sha256:deed13c76f2c99cc245f28eba6d264043a91f3ccca20c191c0427fc9ed0609b5

Observation 9c487ebe-6959-4395-86ea-d46ccef35102 · inbound

Avoiding Bias in Clipped SGD for Overparameterized Models under Generalized Smoothness cites this paper.

Avoiding Bias in Clipped SGD for Overparameterized Models under Generalized Smoothness Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-30T20:15:04.525674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-30T20:13:36.440752Z digest=sha256:b36759e77c67dd8c27f3ac637f9fd09d0fcfbf83326d7a5d234af3987906971f