Pith. sign in

Paper Citation Record · LEDGER

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization

As of 18 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 1 inbound Pith citation observation for arXiv:2412.02153.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.02153 v2

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:55:18.298267Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-07T10:57:04.360455Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T09:26:25.980990Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved21
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d27cd9ed-7dde-4c26-8e3d-246a2fdad191 · outbound

This paper cites SIAM review, 60(2):223–311, 2018.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization SIAM review, 60(2):223–311, 2018

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.875502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.090406Z digest=sha256:b13144c473f21b1a1c4e6c1a39a794439acd2136796dd0310f7010ac848e07a8

Observation fbc0659b-0346-4dc8-bac3-3b6acd5abaa7 · outbound

This paper cites Adaptivesubgradientmethodsforonlinelearning and stochastic optimization.Journal of machine learning research, 12(7), 2011.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Adaptivesubgradientmethodsforonlinelearning and stochastic optimization.Journal of machine learning research, 12(7), 2011

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.863774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.094726Z digest=sha256:7fc47b75971a5ec6120601b2f2ba4a49938eeeaffa7a0593c75096911c850c3f

Observation 0b55d37c-ec4e-489f-a94f-f40e7caf39ea · outbound

This paper cites Neuralnetworksformachinelearning lecture 6a overview of mini-batch gradient descent.Cited on, 14(8):2, 2012.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Neuralnetworksformachinelearning lecture 6a overview of mini-batch gradient descent.Cited on, 14(8):2, 2012

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.851195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.098513Z digest=sha256:f378f3f95d5b22859d3c1c841fb381fa9d030dcdd862818f42bde868aee803a8

Observation 7154ba05-afe1-4f39-92c0-adf66094f4ee · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Adam: A Method for Stochastic Optimization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.102499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.102499Z digest=sha256:9988ca2463bddaeaf2c715051aaaf6ed8cc467aec75c234e08d15a5d09b92d80

Observation 791857b2-973e-4f78-a31b-d11c903f21e4 · outbound

This paper cites Why adam outperforms gradient descent on language models: A heavy-tailed class imbalance problem.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Why adam outperforms gradient descent on language models: A heavy-tailed class imbalance problem

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.839261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.106926Z digest=sha256:5d6a3556c350bac28b761e7dbc8bef6da175e4febae89efc567568d8668c392b

Observation 5d3d1554-d265-4b76-8d1e-8be221b1f8a0 · outbound

This paper cites Toward Understanding Why Adam Converges Faster Than SGD for Transformers.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Toward Understanding Why Adam Converges Faster Than SGD for Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.115103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.115103Z digest=sha256:bb6d07e896f0f0f34b70c2633e59c74655876ef7f56b4c5cf8ecf1b393b66757

Observation 74ba7857-1527-425f-ab0b-7398f40742e2 · outbound

This paper cites an unresolved cited work.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-11T23:55:18.826280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.119808Z digest=sha256:2e0dd411353883c75815b003ae5107c86ef34412acbfca518c35050944ed3371

Observation c5191b5a-a2da-49eb-80b0-249757cd9174 · outbound

This paper cites Why Transformers Need Adam: A Hessian Perspective.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Why Transformers Need Adam: A Hessian Perspective

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.123604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.123604Z digest=sha256:bb8582fe5423863e45d100f051a0d0112a750bd1a1d0da17078ca1ea8310db18

Observation 92d40876-d56f-4502-a5ef-fc0265e9f73d · outbound

This paper cites Adam can converge without any modification on update rules.Advances in neural information processing systems, 35:28386–28399, 2022.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Adam can converge without any modification on update rules.Advances in neural information processing systems, 35:28386–28399, 2022

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.815678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.127601Z digest=sha256:957fe59a78d61dff51f0c058cc13c2ce768ad3cd0685f9314c52f233db631f62

Observation 0e254062-20e2-40e9-8364-d3a6bf27174b · outbound

This paper cites Ontheconvergenceofadamandbeyond.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Ontheconvergenceofadamandbeyond

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.804212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.131423Z digest=sha256:75d4ec8e15cd6107f2dab23d58eb9ae7658d07672eca7a0c658d0838e383a266

Observation b531cc5b-adf0-4d4e-9c55-05dde3af477c · outbound

This paper cites Onthevarianceoftheadaptivelearningrateandbeyond.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Onthevarianceoftheadaptivelearningrateandbeyond

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.793186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.135397Z digest=sha256:f2f0a7f99064821b99c36e18aaa7900d3e266aa4c3e54fcb788e33d21b4aeda3

Observation eeb74156-0b34-4583-9847-c698e58a1d5a · outbound

This paper cites Adaptive gradient methods with dy- namic bound of learning rate.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Adaptive gradient methods with dy- namic bound of learning rate

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.781935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.140039Z digest=sha256:3192385902f1f513c25681fa18bc6eaaed287916daeb8cb8eb6ae6915a34d370

Observation 2a2d0d84-fbfa-4931-a3ae-5280e971826b · outbound

This paper cites Adabelief optimizer: Adapting stepsizes by the belief in observed gradients.Advances in neural information processing systems, 33:18795–18806, 2020.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Adabelief optimizer: Adapting stepsizes by the belief in observed gradients.Advances in neural information processing systems, 33:18795–18806, 2020

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.143664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.143664Z digest=sha256:29e3c97970f772e88831f9fe5f9c31e23464641c693fe25961f92eaa63d6647d

Observation 60c21456-6732-4846-908e-81c9634db9b2 · outbound

This paper cites Momentum is all you need for data-driven adaptive optimization.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Momentum is all you need for data-driven adaptive optimization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.762047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.147621Z digest=sha256:0f8a220ca078908900f0184aad5459b31ec361170c9c346b05e2b7802387e5b3

Observation fd2b5e40-05e0-4b00-bfc2-808a81e52afa · outbound

This paper cites Survey of optimization algo- rithms in modern neural networks.Mathematics, 11(11):2466, 2023.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Survey of optimization algo- rithms in modern neural networks.Mathematics, 11(11):2466, 2023

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.749731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.151424Z digest=sha256:05c898eff8d6fcf17990b9425250617ba01181a2d37c999ea3cfc86f7f255853

Observation 4e825763-1ff0-458c-bf6f-5ebb7ecc4d23 · outbound

This paper cites No train no gain: Revisitingefficienttrainingalgorithmsfortransformer-basedlanguagemodels.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization No train no gain: Revisitingefficienttrainingalgorithmsfortransformer-basedlanguagemodels

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.733754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.155290Z digest=sha256:a2459bd7fa8bdfebc309cc78f2ce2a63aed6632d33b5f550df3f36b040a1d8b4

Observation 9503e7e6-637f-4e50-8912-3ff3dea5a0b7 · outbound

This paper cites Attention is all you need.Advances in Neural Information Processing Systems, 2017.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Attention is all you need.Advances in Neural Information Processing Systems, 2017

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.159144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.159144Z digest=sha256:9e6ced5ecf94b79b71b2b11776463fab5fbc8a45af2952187d8d4b3cd8eb6ae8

Observation 616b0594-0698-463c-b76b-efe36d10d86f · outbound

This paper cites Dissecting adam: The sign, magnitude and variance of stochastic gradients.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Dissecting adam: The sign, magnitude and variance of stochastic gradients

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.714484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.162590Z digest=sha256:542b262729aa56ab8c3a509d28244829ab9f081fba603442bbc7d4b6dd3d8ed7

Observation e3e569c9-d134-4e9a-91e2-950098a2c5d8 · outbound

This paper cites Noise is not the main factor behind the gap between sgd and adam on transformers, but sign descent might be.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Noise is not the main factor behind the gap between sgd and adam on transformers, but sign descent might be

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.700866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.166576Z digest=sha256:37df2d2482b1df2fc812723fcd72f59f93e5744746f8b39d851af53347146b6f

Observation 483bd74d-39e7-4e6f-ab3b-bc86c1cfd077 · outbound

This paper cites Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.171126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.171126Z digest=sha256:e5b8ba427a8f0d12d23f30e9cef2f966a51bf8488e57768e960316fa20fabcf9

Observation fedc59b5-eaae-43be-8a59-e702ccc86ad0 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.174774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.174774Z digest=sha256:aecf8c3ad8b9f1616dd062d54af635a6dc79f3fbc595ad63d597f37f21990e62

Observation c18c163d-93a1-47ce-b256-7de82ad31881 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization An image is worth 16x16 words: Transformers for image recognition at scale

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.178936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.178936Z digest=sha256:9276cf3d026f9dac2d5ce0d428a92c8d1c8db4a40f41f03bdc063039756957be

Observation bee428aa-ab62-43fa-afe1-524ba5a3c689 · outbound

This paper cites Escaping the Big Data Paradigm with Compact Transformers.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Escaping the Big Data Paradigm with Compact Transformers

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.183017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.183017Z digest=sha256:ee2145acd0b79c2ca0e38b6ebd81fc666a93d9a9c4b63250fffd2e4fd2bc7fcc

Observation ee7bd2b8-1511-4123-9a25-16db88075725 · outbound

This paper cites Improving transformer op- timization through better initialization.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Improving transformer op- timization through better initialization

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.681565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.187370Z digest=sha256:61f5ae1ddc0c817f80ed8a76d4ed77788e4bd21637dc5c237df93c2b0a28c19b

Observation d2bc7fba-6729-4177-b206-a07a6563fe18 · outbound

This paper cites Scan and snap: Understanding trainingdynamicsandtokencompositionin1-layertransformer.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Scan and snap: Understanding trainingdynamicsandtokencompositionin1-layertransformer

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.669067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.190823Z digest=sha256:bb9862f1cc0e7cfbc678a7a2ca9a9e42813c3238edba8827f81038a0af0a95ca

Observation 3dea8294-5312-4dad-8477-2a38c972e11a · outbound

This paper cites On the difficulty of training Recurrent Neural Networks.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization On the difficulty of training Recurrent Neural Networks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.194230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.194230Z digest=sha256:7c6c6f24f86fcc38e296cd675f091264e38928eeb5ec050d353d49fa900f80ba

Observation a1632498-320c-4c21-ac5a-f62c91ea4c9d · outbound

This paper cites Visualizing the loss landscape of neural nets.Advances in neural information processing systems, 31, 2018.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Visualizing the loss landscape of neural nets.Advances in neural information processing systems, 31, 2018

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.198263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.198263Z digest=sha256:f8ce9b875873a59e605296e595daf0c6f765f68a4f4548cede676a179c53d10b

Observation f9e78a07-1177-471e-9b03-3a0ee96d91df · outbound

This paper cites On the sdes and scal- ing rules for adaptive gradient algorithms.Advances in Neural Information Processing Systems, 35:7697–7711, 2022.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization On the sdes and scal- ing rules for adaptive gradient algorithms.Advances in Neural Information Processing Systems, 35:7697–7711, 2022

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.650192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.202867Z digest=sha256:bf18b9a7c5db4afcf75bdecb0ed4e7c969bd1668bf7606168f7f81ff2438f098

Observation e2030e20-49c8-4dec-9706-1db481a493ab · outbound

This paper cites an unresolved cited work.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-11T23:55:18.638428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.210430Z digest=sha256:2a05e11f0521a7f843094d03146454a84d663727711502c870dd4917827b7d00

Observation b306e551-bd51-4afc-b145-a2dc601c30f5 · outbound

This paper cites Understanding the difficulty of training deep feedforward neural networks.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Understanding the difficulty of training deep feedforward neural networks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.214398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.214398Z digest=sha256:7323f9d85419e7085bac2d3515b049dbddd9427d9c89b5d97497cea77c33ca65

Observation aa438afe-3957-41d3-a43e-215acaba8e28 · outbound

This paper cites On the Convergence of Adaptive Gradient Methods for Nonconvex Optimization.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization On the Convergence of Adaptive Gradient Methods for Nonconvex Optimization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.218036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.218036Z digest=sha256:5c8d6a635615b1f5967b43de0cb54e595b05705dcfcfaed4db68d66b7b52eb19

Observation a3239ca3-1b01-4c0f-bc14-0c8ed465182f · outbound

This paper cites Deep residual learning for image recognition.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Deep residual learning for image recognition

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.620554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.221996Z digest=sha256:bdfdcc2860965df13146b2d48358463961e281add5f6c54bb9f0ff42bd2ce3dc

Observation 6e4a4f91-c95a-47b0-8350-dd3f958c4894 · outbound

This paper cites Generative adversarial networks.Communications of the ACM, 63(11):139–144, 2020.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Generative adversarial networks.Communications of the ACM, 63(11):139–144, 2020

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.609138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.226209Z digest=sha256:fa29cf3abd9aa04de163855d345c97ca6b85dc38977893a938856ebbd9fe0199

Observation aa2311f3-a0de-4ec4-9659-49bacba4de32 · outbound

This paper cites Long short-term memory.Neural Computation, 9(8):1735–1780, 1997.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Long short-term memory.Neural Computation, 9(8):1735–1780, 1997

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.597359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.230462Z digest=sha256:ef1aebf87b3226182b2bc68136d69dfac0e4a89b3942bfa558a4e6b4ac65d719

Observation 9ec2b5c5-28f0-4b26-8b76-093774850168 · outbound

This paper cites A stochastic approximation method.The annals of mathe- matical statistics, pages 400–407, 1951.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization A stochastic approximation method.The annals of mathe- matical statistics, pages 400–407, 1951

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.585088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.234921Z digest=sha256:81594b238e8dde48209b4807c0197c3681e06c09f36c8a8cf8413cd5a7c55e5c

Observation f29ab8a6-0f0f-4bc1-b165-631fff1c70a3 · outbound

This paper cites Some methods of speeding up the convergence of iteration methods.Ussr computational mathematics and mathematical physics, 4(5):1–17, 1964.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Some methods of speeding up the convergence of iteration methods.Ussr computational mathematics and mathematical physics, 4(5):1–17, 1964

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.239074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.239074Z digest=sha256:061a87004091c5225c76251c255dcdd4fee02afc741d7d49713b83ccf0318728

Observation ae64e145-0cc5-4fda-b4f1-9bc6ba3983c6 · outbound

This paper cites Decoupled weight decay regularization.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Decoupled weight decay regularization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.243371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.243371Z digest=sha256:525e923bd7f30b13ba11660455a5f57f5e33c84ac99716536ad317b6bd135e4b

Observation 8b053159-668b-4b04-baf2-91b25ac8acf6 · outbound

This paper cites Learningmultiplelayersoffeaturesfromtinyimages.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Learningmultiplelayersoffeaturesfromtinyimages

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.557433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.246993Z digest=sha256:0fa0cd163324add558237d6c321d8384e1a71a9c52b83e550abdb10cf37624bc

Observation 767f1587-9603-4c95-a5d5-ead3f3c3b69e · outbound

This paper cites Imagenet large scale visual recognition challenge.International journal of computer vision, 115:211–252, 2015.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Imagenet large scale visual recognition challenge.International journal of computer vision, 115:211–252, 2015

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.250663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.250663Z digest=sha256:dc7757e0c0e9151fedfd77e3f62cc7c28947b36d0e453bde941f1574addfaf80

Observation 8485ca84-2abe-4b3b-848f-a44ac385636e · outbound

This paper cites Building a large annotated corpus of english: The penn treebank.Computational linguistics, 19(2):313–330, 1993.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Building a large annotated corpus of english: The penn treebank.Computational linguistics, 19(2):313–330, 1993

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.254472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.254472Z digest=sha256:9d5501eb002ac21b14ce72d3c5f7501c41fa382824a707a198d5d70304afb965

Observation 623c7226-ab06-4426-9dfe-b0543400e1c1 · outbound

This paper cites fairseq: A fast, extensible toolkit for sequence modeling.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization fairseq: A fast, extensible toolkit for sequence modeling

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.533290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.258037Z digest=sha256:69f7458564da7c2e3bbd20a403d2e4025aca0e830391f08e0657939658a58b37

Observation fe6288ff-e5e5-494e-b28a-d770ca723e25 · outbound

This paper cites Bleu: amethodforautomatic evaluation of machine translation.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Bleu: amethodforautomatic evaluation of machine translation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.521952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.261515Z digest=sha256:4000aeea2749f36dfa00487e7f7799f6bbfdc4209ded75906d8c41412fa4b20a

Observation b8bfc735-fe4c-4d20-b7fd-e98bc15aedaf · outbound

This paper cites Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.265028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.265028Z digest=sha256:4d6e9df73c60d2e6251ab51379e448d37458f62fd6f6665ed1301dfe6aee3fb0

Observation ae5cb523-8896-457c-bfa7-14d0dbffc699 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.Ad- vances in neural information processing systems, 30, 2017.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Gans trained by a two time-scale update rule converge to a local nash equilibrium.Ad- vances in neural information processing systems, 30, 2017

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.510175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.269334Z digest=sha256:5f64b470902742ad1ecf99b2c263a9d006395ead8277182d26be6ba68a03dfa5

Observation 3873b934-4b51-4775-aa0d-001588530268 · outbound

This paper cites Asymmetric valleys: Beyond sharp and flat local minima.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Asymmetric valleys: Beyond sharp and flat local minima

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.499600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.272984Z digest=sha256:97b8e1dcce6f181633cd2ead9a15b69869dc336d9a4887983dc1539c3b214e6d

Observation 36084f00-71b3-4382-ba2a-97d75361106a · outbound

This paper cites Towards the- oretically understanding why sgd generalizes better than adam in deep learning.Advances in Neural Information Processing Systems, 33:21285–21296, 2020.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Towards the- oretically understanding why sgd generalizes better than adam in deep learning.Advances in Neural Information Processing Systems, 33:21285–21296, 2020

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.487962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.276768Z digest=sha256:5b956481b9a6da87831ad58332d37c1aede2848edaba480b3f8816f70163c1a9

Observation 9b683fd7-1aef-43d0-bb67-1032c1c28dfc · outbound

This paper cites On the adequacy of untuned warmup for adaptive optimization.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization On the adequacy of untuned warmup for adaptive optimization

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.475374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.280225Z digest=sha256:81e86f9ee506dcd7e3bd368d0969ec37ef6ee18d2e51ad77767168ff9eb2861c

Observation 4de66683-7cb4-4aa8-948d-f2f6eb1234f1 · outbound

This paper cites Re- thinkingtheinceptionarchitectureforcomputervision.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Re- thinkingtheinceptionarchitectureforcomputervision

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.464169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.283631Z digest=sha256:0179934ed5791ab45c2e5d74a76392d60783f2afd8fd3dd6725cf060a4a666bb

Observation 7b7fe0e2-1d91-49d4-9b38-9a320a418488 · outbound

This paper cites SGDR: Stochastic gradient descent with warm restarts.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization SGDR: Stochastic gradient descent with warm restarts

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.287122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.287122Z digest=sha256:1b4956d0cdd6931dd3f540c178c048726f4f0dbb73cdeaa1c075f0248fe78573

Observation 018f5842-299b-4eae-9179-df7b4a65e0df · outbound

This paper cites Linear mode connectivity and the lottery ticket hypothesis.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Linear mode connectivity and the lottery ticket hypothesis

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.445825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.290543Z digest=sha256:00ed6e4c8889dd62cee4642003ebab0e9b60e48fd62a0f2af8bb0f91d9bd3e74

Observation 478ed16b-4bee-4527-bf8a-ceb2170032bf · outbound

This paper cites The small denominator leads to excessively large initial updates, particularly when¯g is small orσ2 is large.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization The small denominator leads to excessively large initial updates, particularly when¯g is small orσ2 is large

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:55:18.434649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.294054Z digest=sha256:f8bf7378de2d484a50f32ac4543e4731c2e5242bdfc8a1a382e7a9e778ddfcab

Observation 9c013d2f-01a8-4f6a-98c3-74560a32b667 · outbound

This paper cites The model is trained with a length penalty of 1.0, a beam size of 5, and an initial warmup step size of10−7.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization The model is trained with a length penalty of 1.0, a beam size of 5, and an initial warmup step size of10−7

Reference 52

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T23:55:18.423003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:55:18.298267Z digest=sha256:d15c3bd82dcbc2033140af751754674a7839e8138460b6457904a1d8a7e6ea69

Pith citing papers

Observation c147f485-ea7b-4125-8d91-2a028299233c · inbound

A Provably Robust Multi-Jet Framework applied to Active Flow Control of an Airfoil in Weakly Compressible Flow cites this paper.

A Provably Robust Multi-Jet Framework applied to Active Flow Control of an Airfoil in Weakly Compressible Flow Revisiting the Initial Steps in Adaptive Gradient Descent Optimization

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:26:25.982708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-07T10:57:04.360455Z digest=sha256:49d38edd92bc597ed9072feb4f769bf042ec443685b8c0b30f4e4e508c21b2da