Pith. sign in

Paper Citation Record · LEDGER

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size

As of 20 August 2026, this Paper Citation Record lists 100 of 100 outbound references and 0 inbound Pith citation observations for arXiv:2508.15071.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.15071 v1

Coverage vector

measured 100 of 100 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:02:00.305004Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 100 outbound references displayed

  • verified exact4
  • verified fuzzy28
  • unresolved68
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 75ed954f-d7e5-4a45-9329-3157aa52ecc1 · outbound

This paper cites Why Do We Need Weight Decay in Modern Deep Learning?.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Why Do We Need Weight Decay in Modern Deep Learning?

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T04:01:59.813603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:01:59.813603Z digest=sha256:6598645f26796c900e495d7a9bd9bfaed7ba331da628b2cae331e410c9914a9f

Observation 403f152a-6a6f-486d-9200-0c8b3aecbd8e · outbound

This paper cites Complexity guarantees for polyak steps with momentum.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Complexity guarantees for polyak steps with momentum

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T04:01:59.821581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:01:59.821581Z digest=sha256:82eb2654a9ab5e8197e7117f0ccd9189cfa4e3d80309843a9facd6358b3957ef

Observation 362af00c-b526-42c9-add2-0b2671c6380f · outbound

This paper cites signsgd: Compressed optimisation for non-convex problems.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size signsgd: Compressed optimisation for non-convex problems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T04:01:59.827166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:01:59.827166Z digest=sha256:47f926829ce9d27fefae11362795649a7fc45ecc81836f867edb29ab133e8c1a

Observation b080a475-917a-41b5-8c32-d6e03869f1e2 · outbound

This paper cites GPT-NeoX-20B: An Open-Source Autoregressive Language Model.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size GPT-NeoX-20B: An Open-Source Autoregressive Language Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T04:01:59.832448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:01:59.832448Z digest=sha256:e4f86bb6403d4c0eaa439e9378c4e33ce26f97dd9e67a2f72c5fbe73e5fd09cb

Observation be69c4a0-1819-4cb2-aea0-b7a7d68dd687 · outbound

This paper cites On the fast convergence of minibatch heavy ball momentum.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size On the fast convergence of minibatch heavy ball momentum

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-16T04:02:18.867921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:01:59.837418Z digest=sha256:4882b8399366ce58e56c2b3f449adc48bdf4b1549f42e4981ad093e837ea5628

Observation e3c185a7-48f6-4117-9798-1eb8a2a2cf89 · outbound

This paper cites Depthwise Hyperparameter Transfer in Residual Networks: Dynamics and Scaling Limit.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Depthwise Hyperparameter Transfer in Residual Networks: Dynamics and Scaling Limit

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T04:01:59.842207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:01:59.842207Z digest=sha256:039df2c1920cbf0cf2f9edc429d38ba60c26d48602dc94e027b91c124ae3676c

Observation ec149d4d-e97f-4c2f-a593-e7c4e1f79bd5 · outbound

This paper cites Language Models are Few-Shot Learners.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Language Models are Few-Shot Learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T04:01:59.847291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:01:59.847291Z digest=sha256:c61c774077d5c7bf3d109debcfd07ff9381f3f5b387ba6c782563b843b4dbd0e

Observation 4adb68de-4d40-4e94-b9b6-b7a87b0704b8 · outbound

This paper cites Symbolic discovery of optimization algorithms.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Symbolic discovery of optimization algorithms

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T04:01:59.851826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:01:59.851826Z digest=sha256:c4f2ea8567d2b722920205581b0992eb089fcfacf9871c63a8d3bdeea0f5c255

Observation ab9bba67-7bb4-46a8-bab4-aa9eddd05a3c · outbound

This paper cites On Empirical Comparisons of Optimizers for Deep Learning.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size On Empirical Comparisons of Optimizers for Deep Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T04:01:59.856422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:01:59.856422Z digest=sha256:91973083164b6d0f0f9f63aa009777bddd52d201878f978ece334dbe0a5a62c6

Observation 8c2d8e97-46a8-4663-b9af-cd72051f532a · outbound

This paper cites Palm: Scaling language modeling with pathways.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Palm: Scaling language modeling with pathways

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T04:01:59.861183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:01:59.861183Z digest=sha256:814d93142d79d37596446e5d80476e34e338307bf777f78349309a229aad553e

Observation 056d1561-8fed-4550-ace4-5ec377bcc419 · outbound

This paper cites Adaptive methods through the lens of SDE s: Theoretical insights on the role of noise.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Adaptive methods through the lens of SDE s: Theoretical insights on the role of noise

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T04:01:59.870213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:01:59.870213Z digest=sha256:eb98634c5e7c8349ba6b754da01257bfd66cd79eba38557f85e380eadeb71367

Observation 6aec30bb-a9c7-48b2-aea7-063f9ac1c952 · outbound

This paper cites Momentum improves normalized sgd.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Momentum improves normalized sgd

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T04:01:59.875646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:01:59.875646Z digest=sha256:bd3275d90dfbda8d5ed7b1f90fbd000de63f46241cda72f3cdc73a89ad5604d7

Observation cbb35b1a-b3c7-4fc3-b693-0ea56fb8e94f · outbound

This paper cites Momentum-based variance reduction in non-convex sgd.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Momentum-based variance reduction in non-convex sgd

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T04:01:59.880529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:01:59.880529Z digest=sha256:0c8010d1ef6fc8eecaad6f04ecb54afed43d78f6590ed7c0c863069bf09fa7d2

Observation b1fafb98-2a8c-49ae-a669-7c614b5c7fec · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Imagenet: A large-scale hierarchical image database

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T04:01:59.885792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:01:59.885792Z digest=sha256:629991b141fca7d9ef11d71b8e7102b03457e43046624ff1eae72be368d743a5

Observation 7ed8d0da-db96-404d-8af0-9520e9eae34a · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size An image is worth 16x16 words: Transformers for image recognition at scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T04:01:59.890642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:01:59.890642Z digest=sha256:5cc46fd5f05376b8000696a894392011a79cf18732431e8d545f99bc456d70d5

Observation e4f410e4-c06e-45ab-bdea-13841bd7d5bc · outbound

This paper cites Adaptive subgradient methods for online learning and stochastic optimization.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Adaptive subgradient methods for online learning and stochastic optimization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T04:01:59.895580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:01:59.895580Z digest=sha256:384146e096ef447b987c8cd7baff994b79c777cff39fbf2f6112bf773345ccf0

Observation d84542db-9db5-4983-abdd-a66beee89ab0 · outbound

This paper cites Momentum provably improves error feedback! Advances in Neural Information Processing Systems, 2024.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Momentum provably improves error feedback! Advances in Neural Information Processing Systems, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T04:01:59.900512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:01:59.900512Z digest=sha256:1311c86d16a1726efa93d8d2d9b2265c40dac99120eea9638d93086ea5afb50d

Observation 8e646e69-db21-49d1-8120-17a36663ad91 · outbound

This paper cites Lecture 6: Matrix norms and spectral radii.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Lecture 6: Matrix norms and spectral radii

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T04:01:59.905247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:01:59.905247Z digest=sha256:c4f01e5277c364de0373a0e082967b04d254095ab62219c3f548733b46f56a53

Observation 0fa091cf-f408-4d75-92e0-ad6a9aff91aa · outbound

This paper cites When and Why Momentum Accelerates SGD:An Empirical Study.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size When and Why Momentum Accelerates SGD:An Empirical Study

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T04:01:59.910390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:01:59.910390Z digest=sha256:cc3c262c8407d7b1c85ee797da77f01cbe1ee6f616837cf4eedb08c80bb55c89

Observation 41b36890-f508-48fa-b27f-58435cd84597 · outbound

This paper cites Handbook of Convergence Theorems for (Stochastic) Gradient Methods.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Handbook of Convergence Theorems for (Stochastic) Gradient Methods

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T04:01:59.915311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:01:59.915311Z digest=sha256:841c129f78310dca757c846baa3c3a284a7f9129a318cdc425b490b5e75818e4

Observation 210fce15-16f0-4a12-9aa8-ed2918e0991b · outbound

This paper cites Global convergence of the heavy-ball method for convex optimization.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Global convergence of the heavy-ball method for convex optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T04:01:59.920518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:01:59.920518Z digest=sha256:f589889996f5abbfef4f014be7a43854720812f38ae9f07d7966d7ea299bee21

Observation 174f9570-3a11-466f-bd32-01401c60714a · outbound

This paper cites Goodfellow, Yoshua Bengio, and Aaron Courville.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Goodfellow, Yoshua Bengio, and Aaron Courville

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T04:01:59.925929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:01:59.925929Z digest=sha256:2bf3bcccad9a1ccbbabf96cbf0233889c8edefc5165f7ef6b481cdb13244ca56

Observation 1a535e1c-61a7-43f9-9dab-add5e3f9a65d · outbound

This paper cites Sgd for structured nonconvex functions: Learning rates, minibatching and interpolation.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Sgd for structured nonconvex functions: Learning rates, minibatching and interpolation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T04:01:59.931521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:01:59.931521Z digest=sha256:a139bda072303ad6804358ab175b34f4463301f3c431812a9c853c0b28d90a1b

Observation aba92dde-7f24-4e3d-af17-47e17f2e1447 · outbound

This paper cites Analysis of an Idealized Stochastic Polyak Method and its Application to Black-Box Model Distillation.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Analysis of an Idealized Stochastic Polyak Method and its Application to Black-Box Model Distillation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T04:01:59.936451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:01:59.936451Z digest=sha256:a9304a07828f6c2c14ae4f9f20735cc1216a0423ca93dd636764b71b8cced651

Observation a04d8da8-b71d-4f58-beef-26b0f54f82ba · outbound

This paper cites Deep residual learning for image recognition.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Deep residual learning for image recognition

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T04:01:59.941928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:01:59.941928Z digest=sha256:edb8d8dcb73d55cba677612979d5996b9c888d18997eec28046526a964a3135a

Observation 740fac5a-ff24-4f31-b598-2a1d46973180 · outbound

This paper cites Neural networks for machine learning lecture 6a overview of mini-batch gradient descent.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Neural networks for machine learning lecture 6a overview of mini-batch gradient descent

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T04:01:59.946402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:01:59.946402Z digest=sha256:d8422a4641ce77de22d8625b7a8685ae26688eca6cb2abc842d31aaf6e44c449

Observation d1f5bdb0-8cca-47e6-ab26-36e3c8eaa708 · outbound

This paper cites Empirical tests of optimization assumptions in deep learning.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Empirical tests of optimization assumptions in deep learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T04:01:59.951291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:01:59.951291Z digest=sha256:278d393f384efef33072e91fe2b19b21ecd2439d470f9702e6c928cce459f960

Observation eef4b47c-299c-4139-80e3-e8daf80496bd · outbound

This paper cites Long short-term memory.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Long short-term memory

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T04:01:59.955639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:01:59.955639Z digest=sha256:6086c1b1a000b7a99b7be585eb184b288aba0afbdb64e98cbcfd37d8b8bdea18

Observation 1f1cea8b-8643-49ee-a258-c263f5e1f1e5 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Training Compute-Optimal Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T04:01:59.960894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:01:59.960894Z digest=sha256:4e38f78e8b233098d3a64d965c7e2bd0470c42ba6a38eb6964a2e716ab7e835b

Observation f70f5270-f2a1-4395-b745-c75121890d9a · outbound

This paper cites Loss landscape characterization of neural networks without over-parametrization.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Loss landscape characterization of neural networks without over-parametrization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T04:01:59.965887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:01:59.965887Z digest=sha256:a005f37e66ccfbaeca676ec47b0dc93dc99e37fcaee47113a6c51a9b032d86de

Observation 987c96ca-a8f8-40eb-a401-db5c1f729c39 · outbound

This paper cites Near optimal decentralized optimization with compression and momentum tracking.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Near optimal decentralized optimization with compression and momentum tracking

Reference 31

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-16T04:02:09.781779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:01:59.970596Z digest=sha256:4ab7166659feedaa1e49c8faf6c9519a275a34c1cbc30fcc610e9522844752a2

Observation 341e3e36-1be9-490f-8431-30bc80cc2828 · outbound

This paper cites Double momentum and error feedback for clipping with fast rates and differential privacy.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Double momentum and error feedback for clipping with fast rates and differential privacy

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T04:01:59.975331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:01:59.975331Z digest=sha256:d54becb7376ee7c5f4e8d9f7f7a6273ad8ee8a15f608116212673f0e034caf45

Observation 523b001d-2000-4b46-86c3-5e2cd957af75 · outbound

This paper cites Accelerating stochastic gradient descent for least squares regression.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Accelerating stochastic gradient descent for least squares regression

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T04:01:59.980295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:01:59.980295Z digest=sha256:83dcf878846a1732a2287cdc46a772ca816e9e92d82550dc51b9f35bc37e22e6

Observation 932d303f-1c6d-4887-90ec-66713514045b · outbound

This paper cites Towards understanding how momentum improves generalization in deep learning.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Towards understanding how momentum improves generalization in deep learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:02:19.563390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:01:59.984757Z digest=sha256:9d2385e36eaf1b8ed3a9b6c041f0f66ccaaa87ad39e7ce508adfae23a91814fd

Observation 45fd837e-82ef-4cbd-9b40-d3c71febe758 · outbound

This paper cites Adaptive sgd with polyak stepsize and line-search: Robust convergence and variance reduction.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Adaptive sgd with polyak stepsize and line-search: Robust convergence and variance reduction

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:02:19.544944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:01:59.989655Z digest=sha256:198ce5e02ef46c3a70abf9b13fa3d826f8050fe0b18e62182c40db56512f5e76

Observation 88be9c17-a0e5-4a9e-9246-969d022d624d · outbound

This paper cites Muon: An optimizer for hidden layers in neural networks, 2024.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Muon: An optimizer for hidden layers in neural networks, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T04:01:59.993693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:01:59.993693Z digest=sha256:d235b917c12b31d2cb64fecd74a22f2fd30140c446ea7574554f8372504cf6e0

Observation 68c044d0-912a-4bcc-9e9a-1a85a711fd78 · outbound

This paper cites Scaling Laws for Neural Language Models.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Scaling Laws for Neural Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T04:01:59.997766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:01:59.997766Z digest=sha256:552891f81d6108ef240e684ba73dbc30426658ced6ddd7daef010852c4668dde

Observation 7c572350-404a-4c5f-bdeb-6006d6621b67 · outbound

This paper cites char-rnn.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size char-rnn

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T04:02:00.003115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:02:00.003115Z digest=sha256:cd3bbad9a16ae58a28cd3c812fcd1d5e93399269e1b1c6f02a21b96cc9acea2b

Observation 9b554f5b-db03-4b85-98e8-02eccc2519b9 · outbound

This paper cites an unresolved cited work.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:02:19.504798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:02:00.007599Z digest=sha256:4a15e080b2552226fa42d83f1f8cf32987312b28e2c882598d9137a1f7e33d53

Observation 388b9629-0fed-4001-b2bc-bc687205efef · outbound

This paper cites Adam: A method for stochastic optimization.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Adam: A method for stochastic optimization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T04:02:00.012031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:02:00.012031Z digest=sha256:84903ebd14d2f35b4b6650fcf8e8328463d7eed64668b56c7e1de66681714920

Observation 5ac3088c-d89f-4ff8-ac68-3382266e53a3 · outbound

This paper cites Learning multiple layers of features from tiny images.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Learning multiple layers of features from tiny images

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:02:19.476462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:02:00.016264Z digest=sha256:601ceabbe7537cd6860dbe889044256a2bbe6e49fa242f4312255d0c3e1b1044

Observation e209251b-0801-4741-8c7c-8769f6a3d36f · outbound

This paper cites Noise is not the main factor behind the gap between sgd and adam on transformers, but sign descent might be.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Noise is not the main factor behind the gap between sgd and adam on transformers, but sign descent might be

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:02:19.459652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:02:00.021590Z digest=sha256:bf817be0411294bb686a4be84a3a1719a51024f0298f311228228a377da2e8b4

Observation 9c2a846a-b032-4017-be1b-a5bf448f13cf · outbound

This paper cites Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T04:02:00.026561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:02:00.026561Z digest=sha256:f859e90810d0edf7d42a948ae62a1dedcabdbf8e9132f152725ebf5f60a50082

Observation 6de7c598-9e9f-4512-987d-7e6c5e651a3d · outbound

This paper cites Trajectory of mini-batch momentum: batch size saturation and convergence in high dimensions.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Trajectory of mini-batch momentum: batch size saturation and convergence in high dimensions

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:02:19.442827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:02:00.031611Z digest=sha256:1ec5752549c55501a3a707fbff797918c9126d05d57167e1b44e781e44f5accf

Observation 31b14a5f-5c45-49ca-b369-2bcaec6c32c5 · outbound

This paper cites Convergence of adam under relaxed assumptions.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Convergence of adam under relaxed assumptions

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:02:19.426091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:02:00.036593Z digest=sha256:a1c9297fe653bc486086f8fad02f9b9e1c245b3e6bbab8a01d0f7cb5b4dcfc62

Observation e61cf730-7b8b-4446-95ce-e02c9e944240 · outbound

This paper cites PyTorch Distributed: Experiences on Accelerating Data Parallel Training.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size PyTorch Distributed: Experiences on Accelerating Data Parallel Training

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T04:02:00.041685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:02:00.041685Z digest=sha256:91054ffa533ec3bb85197c0ef8fe9f679b31ed49418cbb6501791ba6b437c6c7

Observation 4edf6a43-4863-4074-8d7b-ddcc4de46123 · outbound

This paper cites SP2: A Second Order Stochastic Polyak Method.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size SP2: A Second Order Stochastic Polyak Method

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-16T04:02:00.723166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:02:00.047612Z digest=sha256:a489f147705e21b139e5cdcd28fe8912e064e31c4b57908d04e7dd52e756e6f7

Observation cf55416c-46a5-4871-b7c6-d5275d9cfc81 · outbound

This paper cites An improved analysis of stochastic gradient descent with momentum.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size An improved analysis of stochastic gradient descent with momentum

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:02:19.408006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:02:00.052681Z digest=sha256:a819b50c4bc259fc2abe17a51f4c26d8e04d9f319cb4a94194f931fbadcec9c0

Observation 889c4601-5666-484e-b683-3169e5147704 · outbound

This paper cites Stochastic polyak step-size for sgd: An adaptive learning rate for fast convergence.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Stochastic polyak step-size for sgd: An adaptive learning rate for fast convergence

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:02:19.389977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:02:00.057766Z digest=sha256:d3514c9965b698484e777b91bf496c77ef3e1fb0525404e716e7b52aee215344

Observation dc80494d-6782-4e33-b6c0-980beb2680d4 · outbound

This paper cites Decoupled Weight Decay Regularization.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Decoupled Weight Decay Regularization

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T04:02:00.062565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:02:00.062565Z digest=sha256:beaab74f916e5218b2aed57e798e0514e6136e8d785a116f2aaa45e410dd8103

Observation 6599f62c-3d3d-48f4-bf8e-ffdca5fb9e6f · outbound

This paper cites Adaptive Gradient Methods with Dynamic Bound of Learning Rate.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Adaptive Gradient Methods with Dynamic Bound of Learning Rate

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T04:02:00.067938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:02:00.067938Z digest=sha256:50d468c1af0b7ae61499d6de196e03388e9f714dfe1280959755a7f336c88d60

Observation 382dbff7-06db-4725-b7e0-9ac0c0ceebec · outbound

This paper cites Quasi-hyperbolic momentum and Adam for deep learning.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Quasi-hyperbolic momentum and Adam for deep learning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T04:02:00.073103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:02:00.073103Z digest=sha256:e53b0fda17e63aea90f4bf22b401753359fd822e60ad104e2729a53eb0f26b76

Observation 42f52e0d-7ed3-4e5a-88eb-10d5e0d82333 · outbound

This paper cites Pointer sentinel mixture models.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Pointer sentinel mixture models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:02:19.372278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:02:00.078044Z digest=sha256:20b59ebd654713e5bad3e5a0c9d7a6b630f0589252ec2455ae3b6c02ba31bb9a

Observation 85655926-f5e6-4422-b187-89b0ba5fb808 · outbound

This paper cites Recurrent neural network based language model.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Recurrent neural network based language model

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:02:19.354580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:02:00.082597Z digest=sha256:19c303e9b1f55eb6536e7f371d0c087caae60120d78d9771d380b063418400d5

Observation 10c9c299-6125-41cc-a4af-e253b7948e81 · outbound

This paper cites A Theory on Adam Instability in Large-Scale Machine Learning.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size A Theory on Adam Instability in Large-Scale Machine Learning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T04:02:00.088151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:02:00.088151Z digest=sha256:c20d872a18412b5c2b250ae6380fa7c6b7036e8422004bd15e1b015b9aa2f517

Observation 5a0de51c-9f30-43d8-9c1c-649b2f45e034 · outbound

This paper cites Signal propagation in transformers: Theoretical perspectives and the role of rank collapse.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Signal propagation in transformers: Theoretical perspectives and the role of rank collapse

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:02:19.337854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:02:00.092951Z digest=sha256:6c44701ca3e557a992aaa6a2f77c8f19f03c7185d3aaf0e95e7ed8d6951e23ef

Observation 75b19c71-4040-4864-bd6d-b4cb27e0ecd2 · outbound

This paper cites Stochastic Polyak Step-sizes and Momentum: Convergence Guarantees and Practical Performance.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Stochastic Polyak Step-sizes and Momentum: Convergence Guarantees and Practical Performance

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T04:02:00.097462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:02:00.097462Z digest=sha256:0a776a8de1ffe3dc574747ee4d2530e45ffea91981fdda78c98a9037f3ebc083

Observation c71116f5-2cd0-4fbf-9957-ef0213b41407 · outbound

This paper cites The Cost of Training NLP Models: A Concise Overview.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size The Cost of Training NLP Models: A Concise Overview

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T04:02:00.102551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:02:00.102551Z digest=sha256:a84e55a69feb17dce73593f5521bdbbc0845e43ad7ba31e2c964beace9382048

Observation 4bfdd543-bdff-4703-afe0-40d4ce78a86f · outbound

This paper cites An Adaptive Stochastic Gradient Method with Non-negative Gauss-Newton Stepsizes.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size An Adaptive Stochastic Gradient Method with Non-negative Gauss-Newton Stepsizes

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T04:02:00.107734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:02:00.107734Z digest=sha256:2ccbf47130b66762040a7608e0b27909dc03d8f8ad94cb690427f86ac92685a1

Observation 6108796e-ab14-4c86-8d43-7288afb57cfa · outbound

This paper cites Dynamics of sgd with stochastic polyak stepsizes: Truly adaptive variants and convergence to exact solution.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Dynamics of sgd with stochastic polyak stepsizes: Truly adaptive variants and convergence to exact solution

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:02:19.320234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:02:00.112532Z digest=sha256:51483bb730d9b383e941724061dac1bcedc21dc29bda5b473bc19bddc4b0fd7b

Observation 75e87561-2ebc-40cc-9cba-582c63add8d1 · outbound

This paper cites Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T04:02:00.117089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:02:00.117089Z digest=sha256:ec18358664a0f880ec806dc4e9e7814e941aeac2f6a49ada558170b784349385

Observation 1c955981-b673-40dd-8fd3-4725b85e56e4 · outbound

This paper cites Automatic differentiation in pytorch.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Automatic differentiation in pytorch

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:02:19.291055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:02:00.121205Z digest=sha256:3e0ed4a8db5147fd5d6105f041c593e06b8de541266599ed732b00e47be48343

Observation f589f836-559e-44b3-a0a1-215ae6486243 · outbound

This paper cites The fineweb datasets: Decanting the web for the finest text data at scale.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size The fineweb datasets: Decanting the web for the finest text data at scale

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:02:19.274427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:02:00.126086Z digest=sha256:201045192c1aa35ac3edfe440da7475960601bcdf51bd7dd0f2d352066d6e8d2

Observation 3f2fe75e-3a1f-46bd-b133-910c85dec5d8 · outbound

This paper cites an unresolved cited work.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:02:19.257590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:02:00.130253Z digest=sha256:aebc1331c9518aca7b67bce66c31146fee1de39fb532c68f9705af9dc6656f91

Observation de02a3ab-32af-4735-8346-d775dbd41375 · outbound

This paper cites Introduction to optimization.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Introduction to optimization

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:02:19.239432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:02:00.135610Z digest=sha256:f9605f1438ce687a43598c4c0e1b4ea35829e8d8e19bb6e497e80cb0ea6c9c96

Observation 6c8cb3a1-5131-4ffd-8246-c252b27e2fa2 · outbound

This paper cites Improving language understanding by generative pre-training.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Improving language understanding by generative pre-training

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:02:19.223159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:02:00.140515Z digest=sha256:30dd08908c615dc91a7d8ff9a6e0d1f1d26ad7e53bfb966b416f048dfe1cf9ca

Observation a79463ad-6153-432c-89b1-e5ee8fbba693 · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Zero: Memory optimizations toward training trillion parameter models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-16T04:02:00.144822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:02:00.144822Z digest=sha256:f0fa22fff947b8570662ec77e2c4e1d9fe6c06242d05f6f59746ec3ff4f59c1a

Observation 9cefb47f-7977-49d0-90be-0bc73043e303 · outbound

This paper cites Local Curvature Descent: Squeezing More Curvature out of Standard and Polyak Gradient Descent.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Local Curvature Descent: Squeezing More Curvature out of Standard and Polyak Gradient Descent

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-08-16T04:02:00.583192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:02:00.148860Z digest=sha256:425015d1ba0a6d2cb5e548ae0cc052921e8b4c2f2daf81ed55931bc5a4531f29

Observation 294d1134-d00a-42eb-8cf1-06c63e1801f2 · outbound

This paper cites An adaptive polyak heavy-ball method.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size An adaptive polyak heavy-ball method

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:02:19.193234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:02:00.153882Z digest=sha256:e69150eddfd44fda24b23e05da27a37eb55dbb60e5d3a6a59f6ca31ad506f23f

Observation fbeef074-4663-4bf4-8d8b-6c09bd2add1b · outbound

This paper cites Stochastic sign descent methods: New algorithms and better theory.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Stochastic sign descent methods: New algorithms and better theory

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:02:19.177317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:02:00.158401Z digest=sha256:8d718a933bdcbbce868ae663da80aab94d45b486698facd5742659b5d97d53bf

Observation f4c38cc6-add0-4738-ab94-cc84ddbae6c5 · outbound

This paper cites an unresolved cited work.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:02:19.159324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:02:00.163337Z digest=sha256:ee2731e687887821c4f362e4b976aba6aad5041b8f537bd35a9b45c1e0f8f52d

Observation 613a344c-4907-4ab3-acb8-7e234dc9ae37 · outbound

This paper cites The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-16T04:02:00.168680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:02:00.168680Z digest=sha256:c5a98acec1d61cd531667be7267436d86adead861a360f38a7da556d47105f11

Observation 8af1993d-2b04-4556-aad5-2267ae84df49 · outbound

This paper cites Almost sure convergence rates for stochastic gradient descent and stochastic heavy ball.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Almost sure convergence rates for stochastic gradient descent and stochastic heavy ball

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:02:19.141902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:02:00.173470Z digest=sha256:69933a411adda2b70fc649bb0b9564d55bf2a479ebef485586e5b0d2664a687f

Observation 4ce88943-798c-49e1-aede-e36e7f0d2fac · outbound

This paper cites GLU Variants Improve Transformer.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size GLU Variants Improve Transformer

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-16T04:02:00.178104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:02:00.178104Z digest=sha256:01fb48f771e8cd48e7166390cca04dbdc22ede10bb98c6c8097f05cdf197a4fa

Observation 3bf10510-5577-434b-90b2-4d40c2d6a755 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-16T04:02:00.183479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:02:00.183479Z digest=sha256:c2b3925f2a2a43a2eefe7923103f32009bef4655890a054fec9cdaea73274227

Observation 7a7bf361-e9b0-4127-aab4-ec9ea8c62681 · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-16T04:02:00.188602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:02:00.188602Z digest=sha256:c378bada51434678b117b46d60f24553131969f42b6e5544e0d8327154a132be

Observation fefdfb3f-7b4d-4d94-9663-be4e5b230cb6 · outbound

This paper cites an unresolved cited work.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:02:19.124568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:02:00.194360Z digest=sha256:8b7f1998d5f32a0f0c6d9b4bdef5755b82e1abb5dbefc218019aab2227fb9628

Observation 1465ed5f-fb5d-4e46-96a4-80e45f355fe5 · outbound

This paper cites SlimPajama: A 627B token cleaned and deduplicated version of RedPajama , 2023.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size SlimPajama: A 627B token cleaned and deduplicated version of RedPajama , 2023

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-16T04:02:00.199994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:02:00.199994Z digest=sha256:80fc56f20fc9ca6c88bab31c711d128105e5ca3db4b11142f75d164d00c9c22b

Observation 872d84e4-024e-44f0-b87c-f628df32b0d0 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding, 2023.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Roformer: Enhanced transformer with rotary position embedding, 2023

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-16T04:02:00.204867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:02:00.204867Z digest=sha256:41131d10d6fc06ee76484344ac98ece7c898355498f71c10d77191088af5a11d

Observation d60fc248-cc9c-43d2-b4e6-15152672532b · outbound

This paper cites Training data-efficient image transformers & distillation through attention.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Training data-efficient image transformers & distillation through attention

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:02:19.085412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:02:00.209596Z digest=sha256:63fbea341b9c9815db55cb42dfde7b12143e81e52e0a3ad917409050f86705b8

Observation 502f7226-7a91-4b88-b7a3-9667ae9cce9a · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size LLaMA: Open and Efficient Foundation Language Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-16T04:02:00.214360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:02:00.214360Z digest=sha256:ad6a51daacc64cc982a1b7e3b8eb970bc044507b78165bf543fcc32bca7baad9

Observation 21a874d3-f111-41f0-a194-a47ed337656a · outbound

This paper cites Closing the gap between the upper bound and lower bound of adam's iteration complexity.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Closing the gap between the upper bound and lower bound of adam's iteration complexity

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:02:19.066751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:02:00.219873Z digest=sha256:99617b1770201143a29e1c8e8f04fbb097b197fd5bb40d49059f97c178fc1655

Observation 35b6f9a7-1272-4994-9662-a92a3a9fa478 · outbound

This paper cites A modular analysis of provable acceleration via polyak’s momentum: Training a wide relu network and a deep linear network.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size A modular analysis of provable acceleration via polyak’s momentum: Training a wide relu network and a deep linear network

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:02:19.050082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:02:00.225147Z digest=sha256:d9c3cc6c822aff45c22e059897645ced792050b50336b2e13dd302183342f95d

Observation fb67fc94-be09-40f1-b031-aca460d79cd4 · outbound

This paper cites Provable acceleration of heavy ball beyond quadratics for a class of polyak-lojasiewicz functions when the non-convexity is averaged-out.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Provable acceleration of heavy ball beyond quadratics for a class of polyak-lojasiewicz functions when the non-convexity is averaged-out

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:02:19.032376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:02:00.230109Z digest=sha256:a4d11cdb1c0de1b732ec6f8ae2d3909c7dcc373f32303e306982e8e2ce977351

Observation 2d7477f5-7cf3-4026-839c-6048edc0cafb · outbound

This paper cites Generalized polyak step size for first order optimization with momentum.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Generalized polyak step size for first order optimization with momentum

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:02:19.015658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:02:00.234565Z digest=sha256:8767fd7bc75996e4de05d6a8cab0c93873c70bcc8a372530c17958c24493aa04

Observation d3737ef1-bb2b-43d3-824c-6aabdb707b4c · outbound

This paper cites Adagrad stepsizes: Sharp convergence over nonconvex landscapes.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Adagrad stepsizes: Sharp convergence over nonconvex landscapes

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:02:18.998539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:02:00.239265Z digest=sha256:a39d38dabd7d43b1eb132d8901218b86e5d8461ebd8ed2a054f08a7c02325108

Observation 4f763232-ffc2-4fab-9684-251170d20d3f · outbound

This paper cites Pytorch image models.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Pytorch image models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-16T04:02:00.244044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:02:00.244044Z digest=sha256:643d59b80cf2bd2e6c7c860837ffda39f515af4afa1c99ad3e1207d5cd264bf5

Observation d35b28d0-6ac4-41d7-9e19-724aea4601db · outbound

This paper cites The marginal value of adaptive gradient methods in machine learning.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size The marginal value of adaptive gradient methods in machine learning

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:02:18.969760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:02:00.248729Z digest=sha256:64b6b7f080b36e01b8c44f879a1aa3463066afed0dd450c1c27283c766957a40

Observation 74b509a2-58cb-4f15-8fb7-b0ae264c1282 · outbound

This paper cites Small-scale proxies for large-scale Transformer training instabilities.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Small-scale proxies for large-scale Transformer training instabilities

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-16T04:02:00.253488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:02:00.253488Z digest=sha256:4904ff9311a2341b5cdc0c840c1190dd6e979f36c90538546f5693d946bb845e

Observation 80352125-483b-4608-bb84-f07ffe1bfff5 · outbound

This paper cites Rethinking Conventional Wisdom in Machine Learning: From Generalization to Scaling.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Rethinking Conventional Wisdom in Machine Learning: From Generalization to Scaling

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-16T04:02:00.258469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:02:00.258469Z digest=sha256:9ed4e7c827119bdf1850a03cb72dccc9f8d22974b73484d73e63604bfd976680

Observation 0497381e-c882-4f61-9917-006aa4687d02 · outbound

This paper cites Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-16T04:02:00.263330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:02:00.263330Z digest=sha256:5bcec67441b75e33243bca853ba6d44f28350c432efaf2e39a1a8a998ef07ea1

Observation e2fadc0d-6e97-4015-b734-024a70326937 · outbound

This paper cites A Spectral Condition for Feature Learning.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size A Spectral Condition for Feature Learning

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-16T04:02:00.267658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:02:00.267658Z digest=sha256:ca0ba77d06984132cc9ab14265b0a451f0acb1c98d53f375331a6b14cfbaec85

Observation 79910aa5-df7b-4bc9-8d1e-aa6fc5ede60b · outbound

This paper cites Root mean square layer normalization, 2019.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Root mean square layer normalization, 2019

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-16T04:02:00.272036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:02:00.272036Z digest=sha256:0ed2c787bc5845ded3200ede7bfaa953c113ac5d2c4d249a1a3e4a8abc72f208

Observation 9dbe6009-2284-446e-8162-bb8d8e4b7ae2 · outbound

This paper cites Three Mechanisms of Weight Decay Regularization.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Three Mechanisms of Weight Decay Regularization

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-16T04:02:00.276230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:02:00.276230Z digest=sha256:5765eb5f539b6954fbb7758f2ece377753297614cd58d022fa5d78209539b283

Observation 275d864c-e90d-4c78-b56f-a8aebcf0ad94 · outbound

This paper cites Why gradient clipping accelerates training: A theoretical justification for adaptivity.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Why gradient clipping accelerates training: A theoretical justification for adaptivity

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-16T04:02:00.280462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:02:00.280462Z digest=sha256:f0dc8b16b0b78b426c337b1f7a620ee3e1b4247600f6860d90ced593c8e3eb72

Observation e5677c9c-4358-472b-8d5c-8b4a8209a0c5 · outbound

This paper cites Why are adaptive methods good for attention models? Advances in Neural Information Processing Systems, 2020.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Why are adaptive methods good for attention models? Advances in Neural Information Processing Systems, 2020

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:02:18.940280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:02:00.285513Z digest=sha256:95619d5bc673048f5a8b30c7118b1c6f4b95be192adb73205122bc35d312bcb0

Observation be2b43dd-a363-4388-aee0-e9b00845ae87 · outbound

This paper cites Convergence Guarantees for RMSProp and Adam in Generalized-smooth Non-convex Optimization with Affine Noise Variance.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Convergence Guarantees for RMSProp and Adam in Generalized-smooth Non-convex Optimization with Affine Noise Variance

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-16T04:02:00.290034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:02:00.290034Z digest=sha256:ed80093ed5577f6efb521756fef8e8a8a6a2e9118b8391f46df6f7b9717aafe7

Observation ba293250-776b-4161-8682-bd24ed140fbb · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size OPT: Open Pre-trained Transformer Language Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-16T04:02:00.294861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:02:00.294861Z digest=sha256:6ac74a38cde84d7ef5a4a1ad0e6bd01d743dd18b754b156c966bb3bec170d1a5

Observation 0b1bdd6f-6eda-4668-93dd-80e1fe93e523 · outbound

This paper cites Why Transformers Need Adam: A Hessian Perspective.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Why Transformers Need Adam: A Hessian Perspective

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-16T04:02:00.300018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:02:00.300018Z digest=sha256:fee5e5a150ea5c4ba342c36bcc641f83b4845f7c4997ce5c1619a4971702969d

Observation 59428661-2869-4a47-8111-81f1d806528a · outbound

This paper cites Adabelief optimizer: Adapting stepsizes by the belief in observed gradients.

Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size Adabelief optimizer: Adapting stepsizes by the belief in observed gradients

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:02:18.923531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T04:02:00.305004Z digest=sha256:d6e8c618315ee8e70d7c4d5f5f4f863a9c7aac5e2a52bbedceaa27c5b2cddb01

Pith citing papers

No inbound Pith citation observations are available.