Pith. sign in

Paper Citation Record · LEDGER

Linear Convergence of Adaptive Stochastic Gradient Descent

As of 17 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:1908.10525.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.10525 v2

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T10:47:31.432012Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6ab720da-0006-41a2-8b13-bdf23c017ba5 · outbound

This paper cites (2018): F (xk0−1) ≤ F (x0) + η2L 2 (1 + log( b2 k0−1 b2 0 )).

Linear Convergence of Adaptive Stochastic Gradient Descent (2018): F (xk0−1) ≤ F (x0) + η2L 2 (1 + log( b2 k0−1 b2 0 ))

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:47:31.747751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T10:47:31.421243Z digest=sha256:1eeb9e30d99d446ea10a135ec6123bb23e903af5eb27d910a07bdfd9b8fa3774

Observation 8eb866e5-a1fe-4b43-9a4c-f99af31a8853 · outbound

This paper cites However, after tuning η = 10000 in stochastic setting and η = 100 in batch setting, the convergence rate of AdaGrad-Norm is better again.

Linear Convergence of Adaptive Stochastic Gradient Descent However, after tuning η = 10000 in stochastic setting and η = 100 in batch setting, the convergence rate of AdaGrad-Norm is better again

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:47:31.731828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T10:47:31.426103Z digest=sha256:5294da20fb76ba23f20159bc05defccc16a57ee83c6388eb5f53807b2d23d0c4

Observation 4b273839-6338-4422-83f8-afdb1f0a4925 · outbound

This paper cites an unresolved cited work.

Linear Convergence of Adaptive Stochastic Gradient Descent Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-14T10:47:31.715509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T10:47:31.432012Z digest=sha256:050d882cb046dc03ef55fa68d4dd5985c5ff0f5099f3c014878946474dddeb94

Observation 4c03a4da-f398-43b0-a21e-c62c94452a9d · outbound

This paper cites A Sufficient Condition for Convergences of Adam and RMSProp.

Linear Convergence of Adaptive Stochastic Gradient Descent A Sufficient Condition for Convergences of Adam and RMSProp

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.375228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.375228Z digest=sha256:3e47808997dfb238f901042a12476d1736a332bfdb2c1ca3759a7b7da0ed9835

Observation 904dc547-dc3b-4335-a6f3-35741be8a6b3 · outbound

This paper cites AdaGrad stepsizes: Sharp convergence over nonconvex landscapes.

Linear Convergence of Adaptive Stochastic Gradient Descent AdaGrad stepsizes: Sharp convergence over nonconvex landscapes

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.380155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.380155Z digest=sha256:9c83b9b7df72a2fe056825ed2112c42e51b37b6e6c2a5ea7c0c717e76ba23735

Observation be367608-7198-49ba-b812-13d0d7321d36 · outbound

This paper cites WNGrad: Learn the Learning Rate in Gradient Descent.

Linear Convergence of Adaptive Stochastic Gradient Descent WNGrad: Learn the Learning Rate in Gradient Descent

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.385502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.385502Z digest=sha256:771c33f7df501d956e4796b5c06f1aeb27b1a8b61b2463a6db21ee64ba4c2feb

Observation f6db588b-1399-4e02-9f65-f0dcd41ac1df · outbound

This paper cites On the Convergence of Stochastic Gradient Descent with Adaptive Stepsizes.

Linear Convergence of Adaptive Stochastic Gradient Descent On the Convergence of Stochastic Gradient Descent with Adaptive Stepsizes

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.390275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.390275Z digest=sha256:3a18002bdfece6a7d1c5b8494bf22ca0b848f99ee71e35b41e686de90465176d

Observation ca916c6b-dc1e-4797-849b-39e7af5b864d · outbound

This paper cites Fast and Faster Convergence of SGD for Over-Parameterized Models and an Accelerated Perceptron.

Linear Convergence of Adaptive Stochastic Gradient Descent Fast and Faster Convergence of SGD for Over-Parameterized Models and an Accelerated Perceptron

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.394878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.394878Z digest=sha256:c05238ac21edcfd9f33dc6c922a96f87b72201938f73bb0c3f769169b24f854a

Observation d50a180f-be27-4024-9110-15c9ec01e24b · outbound

This paper cites Understanding deep learning requires rethinking generalization.

Linear Convergence of Adaptive Stochastic Gradient Descent Understanding deep learning requires rethinking generalization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.401412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.401412Z digest=sha256:eb7fe3293b843684912abb6ee5e4d04b8df225dd2313a7eb1b94e081293a4d3d

Observation 4fb55d86-6a6a-4bb4-af75-616871783ff7 · outbound

This paper cites Fast Convergence of Stochastic Gradient Descent under a Strong Growth Condition.

Linear Convergence of Adaptive Stochastic Gradient Descent Fast Convergence of Stochastic Gradient Descent under a Strong Growth Condition

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.411341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.411341Z digest=sha256:a271fa35abc47b3b93aa06b86d3d69380ed92f18d5a337da07fc1dfe9332fdb4

Observation 77a35abd-4fcb-4c89-b5a7-2933697ae1f4 · outbound

This paper cites Global Convergence of Adaptive Gradient Methods for An Over-parameterized Neural Network.

Linear Convergence of Adaptive Stochastic Gradient Descent Global Convergence of Adaptive Gradient Methods for An Over-parameterized Neural Network

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.416037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.416037Z digest=sha256:392ffc45e5e0986031e3d0e5f11a436ab9aa93296f43d46ec1ddaf3a70b2476d

Observation a0b0faa3-ae45-40ad-9c3b-1769a56a0136 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Linear Convergence of Adaptive Stochastic Gradient Descent Adam: A Method for Stochastic Optimization

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.359712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.359712Z digest=sha256:25c30295e0de47bb9f763dbef8e4b5e8f2512e29ad070c7cc78def07baeaa7c2

Observation 0f4a679c-2d5e-4bbb-a109-ee8c53aa91ff · outbound

This paper cites Adaptive Bound Optimization for Online Convex Optimization.

Linear Convergence of Adaptive Stochastic Gradient Descent Adaptive Bound Optimization for Online Convex Optimization

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.354327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.354327Z digest=sha256:e28d3d7f47933cc9453d137fdd6b86c7d02688f23422f2922912d6edd3f24f27

Observation a279fb87-f15b-4112-b7d2-7bc0a0e96f4e · outbound

This paper cites Diagonal Rescaling For Neural Networks.

Linear Convergence of Adaptive Stochastic Gradient Descent Diagonal Rescaling For Neural Networks

Reference 2014

Resolution
verified exact
local_arxiv, observed 2026-08-14T10:47:31.631995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T10:47:31.364853Z digest=sha256:7bdf61aa93ee4a8133650bb1163f20d13be55bd6c19613517d5a4894d3059a86

Observation e54fa5fe-e7ce-4bf7-b501-f56eb3228447 · outbound

This paper cites A Convergence Theory for Deep Learning via Over-Parameterization.

Linear Convergence of Adaptive Stochastic Gradient Descent A Convergence Theory for Deep Learning via Over-Parameterization

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.343819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.343819Z digest=sha256:52a5fd403895c5f44a132194094ecd33b047c7f3e6db44ace5aede355922ba39

Observation 9c02d46b-dbf4-4e69-9580-1d8f4f47814f · outbound

This paper cites On the Convergence of Adam and Beyond.

Linear Convergence of Adaptive Stochastic Gradient Descent On the Convergence of Adam and Beyond

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.370129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.370129Z digest=sha256:7ed55114146876390baed5d185806adb371d282fe7e0966ec949a9303e543fa8

Observation 2c219bbd-55e3-4a38-bbd7-6e3ab40b0eda · outbound

This paper cites Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks.

Linear Convergence of Adaptive Stochastic Gradient Descent Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.349191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.349191Z digest=sha256:1c13dab61c54ce5d98fd4768bb516633fed3c1aab4ccbc1c9b5fef11d60d03b4

Observation e882a8e9-4bb5-4bd9-b6c9-6a8e06f19d78 · outbound

This paper cites On exponential convergence of SGD in non-convex over-parametrized learning.

Linear Convergence of Adaptive Stochastic Gradient Descent On exponential convergence of SGD in non-convex over-parametrized learning

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-14T10:47:31.406356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:47:31.406356Z digest=sha256:b47bc39ec6729de3c15c641ea70644abcee53e0e30c1b98a1f92d66891d298ea

Pith citing papers

No inbound Pith citation observations are available.