Pith. sign in

Paper Citation Record · LEDGER

Toward a Unified Theory of Gradient Descent under Generalized Smoothness

As of 11 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2412.11773.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11773 v2

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:47:46.392025Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved9
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 28a837b3-8d0f-4e12-9adc-e796e768e00e · outbound

This paper cites For any M ≥ 0, taking ¯T (M ) such that ∥∇f (x ¯T )∥ ≤M, we get f (xT ) − f (x∗) ≤ ℓ(2M ) ∥x0 − x∗∥2 2(T − ¯T (M ) + 1).

Toward a Unified Theory of Gradient Descent under Generalized Smoothness For any M ≥ 0, taking ¯T (M ) such that ∥∇f (x ¯T )∥ ≤M, we get f (xT ) − f (x∗) ≤ ℓ(2M ) ∥x0 − x∗∥2 2(T − ¯T (M ) + 1)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:47:46.527885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:47:46.392025Z digest=sha256:88b2beda8c8a29e457ce2e064c733230fddea7602fb9c91cfe04a281a767d1bc

Observation 24cac4c9-6c86-493e-abd1-81a08474564f · outbound

This paper cites This function is (3.3, 1)–smooth, meaning we can run Algorithm 1 with ℓ(s) = 3.3 + s.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness This function is (3.3, 1)–smooth, meaning we can run Algorithm 1 with ℓ(s) = 3.3 + s

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T14:47:46.583136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:47:46.374599Z digest=sha256:6f5cff654a2659e8d7feace46ddfc090340836d09e1f19adecba91e96809fa11

Observation 94b02397-078e-4add-bcde-fa63625f7551 · outbound

This paper cites Large Deviations of Vector-valued Martingales in 2-Smooth Normed Spaces.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Large Deviations of Vector-valued Martingales in 2-Smooth Normed Spaces

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T14:47:46.340297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:47:46.340297Z digest=sha256:71015348fa2467d120ac7ac440c2fd6b54acddd2d5014e49948769e79257eb8e

Observation 798fd856-6e30-4944-bdbe-37cc476a9141 · outbound

This paper cites Federated Learning: Strategies for Improving Communication Efficiency.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Federated Learning: Strategies for Improving Communication Efficiency

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T14:47:46.344761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:47:46.344761Z digest=sha256:8a022ff5da7009b311bdef082e688cfca85e8d3320ff754dfec47ae0f995ace5

Observation e65d4d06-6939-4437-879b-52de45629bef · outbound

This paper cites Optimizing $(L_0, L_1)$-Smooth Functions by Gradient Methods.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Optimizing $(L_0, L_1)$-Smooth Functions by Gradient Methods

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T14:47:46.353647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:47:46.353647Z digest=sha256:76c5e88b29b6426a0c9647ccd49d663abe2f748395351ea7b9ba60fbb64a4801

Observation 396e8e86-8831-423d-8707-c84d8eb55bd0 · outbound

This paper cites Gradient-Variation Online Learning under Generalized Smoothness.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Gradient-Variation Online Learning under Generalized Smoothness

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-11T14:47:46.446044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:47:46.358018Z digest=sha256:7f5cfeb6b0bee37819c03d4c3173beefc3e71282777b3dd8072d927c86a5b4fa

Observation 96fe0dda-9a44-4369-a088-4427ece666d3 · outbound

This paper cites Why gradient clipping accelerates training: A theoretical justification for adaptivity.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Why gradient clipping accelerates training: A theoretical justification for adaptivity

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T14:47:46.361895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:47:46.361895Z digest=sha256:780424fff11d49428d5ec35a6d585e20d12a0250c432140a56e2e4db33625c49

Observation 4eb5ac19-5ad3-4592-93e8-d45fd3bd0718 · outbound

This paper cites Next, we take the step size γk = 1/(800 + 2(2f ′(x0))2) from (Li et al., 2024a) and observe that GD requires at least 20.000 iterations because f ′(x0) is huge.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Next, we take the step size γk = 1/(800 + 2(2f ′(x0))2) from (Li et al., 2024a) and observe that GD requires at least 20.000 iterations because f ′(x0) is huge

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:47:46.595659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:47:46.370237Z digest=sha256:72bb9d69ffa5455f7d9f5e7e26747c0edca901e33d7abf5b64c03d402b2bc5a5

Observation de1bb650-520f-4772-b690-52acced1c794 · outbound

This paper cites Using the standard differential algebra, we can solve it: dg(t) ℓ(∥∇f (x)∥ + g(t)) = dt ⇒ Z t 0 dg(v) ℓ(∥∇f (x)∥ + g(v)) = t ⇒ Z g(t) 0 dv ℓ(∥∇f (x)∥ + v) = t.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Using the standard differential algebra, we can solve it: dg(t) ℓ(∥∇f (x)∥ + g(t)) = dt ⇒ Z t 0 dg(v) ℓ(∥∇f (x)∥ + g(v)) = t ⇒ Z g(t) 0 dv ℓ(∥∇f (x)∥ + v) = t

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:47:46.569462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:47:46.378443Z digest=sha256:f3976e97e7aadf5a6082ab8d1a1eff8eb7be2ba0d9c179b303018d410f1faee1

Observation 43edb0f4-f171-4a37-9873-2c0ff80b75b9 · outbound

This paper cites an unresolved cited work.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:47:46.555804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:47:46.383119Z digest=sha256:3f64fb11131e2fa474d033dad05e3d821971700df7cd7c6f6e4b54bbd35f8491

Observation 993fdb7d-0f8e-4e28-a9d8-7bc75035d8a6 · outbound

This paper cites Due to the strategy from Alg.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Due to the strategy from Alg

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:47:46.542435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:47:46.387503Z digest=sha256:bcc059b90755912817e5f37705381b1905f53c301a467ff97d8458fdd1a32292

Observation 0892a7b6-1813-4d05-ad69-5c4ba1b58b0a · outbound

This paper cites Parameter-free Clipped Gradient Descent Meets Polyak.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Parameter-free Clipped Gradient Descent Meets Polyak

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T14:47:46.349125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:47:46.349125Z digest=sha256:128939af7d631f8125f571a0a8f90c12254bc2425a2faa68e8ab76be169c0f03

Observation 9dd21d60-127c-4ee0-b65e-54bccbb5aa1a · outbound

This paper cites an unresolved cited work.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Unresolved cited work

Reference 2019

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:47:46.608869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:47:46.366168Z digest=sha256:b73bdeb687274bb4938dbc2bc7a9f0f3c691eac11c5b07e7abc41dfc628cafcd

Observation 8dd84dcc-81f2-49d3-a939-28321be14537 · outbound

This paper cites Methods for Convex $(L_0,L_1)$-Smooth Optimization: Clipping, Acceleration, and Adaptivity.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Methods for Convex $(L_0,L_1)$-Smooth Optimization: Clipping, Acceleration, and Adaptivity

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T14:47:46.330224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:47:46.330224Z digest=sha256:18ea5fcbfc1177b44aba681f93b26bba320bb3552edd98974014b8d4cd422916

Observation d5a437c0-4c84-4287-b832-013d350911c3 · outbound

This paper cites A theoretical study of the(l 0, l1)-smoothness condition in deep learning.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness A theoretical study of the(l 0, l1)-smoothness condition in deep learning

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:47:46.621303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:47:46.324863Z digest=sha256:d9a9f8764b60b04cf58f5c8b7b6e806ed4a80f62a67c56aedd9e58d261e64ad4

Observation dac0dda4-a466-4f9c-9efa-70e19b249496 · outbound

This paper cites Accelerated Objective Gap and Gradient Norm Convergence for Gradient Descent via Long Steps.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Accelerated Objective Gap and Gradient Norm Convergence for Gradient Descent via Long Steps

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T14:47:46.334977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:47:46.334977Z digest=sha256:8bda5c631247273f06ff3ea5c10143e2ee6b95484dd7fc6a213f3a24bc89fcb5

Pith citing papers

No inbound Pith citation observations are available.