Pith. sign in

Paper Citation Record · LEDGER

Toward a Unified Theory of Gradient Descent under Generalized Smoothness

As of 12 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2412.11773.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11773 v2

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:47:46.392025Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved9
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 28a837b3-8d0f-4e12-9adc-e796e768e00e · outbound

This paper cites For any M ≥ 0, taking ¯T (M ) such that ∥∇f (x ¯T )∥ ≤M, we get f (xT ) − f (x∗) ≤ ℓ(2M ) ∥x0 − x∗∥2 2(T − ¯T (M ) + 1).

Toward a Unified Theory of Gradient Descent under Generalized Smoothness For any M ≥ 0, taking ¯T (M ) such that ∥∇f (x ¯T )∥ ≤M, we get f (xT ) − f (x∗) ≤ ℓ(2M ) ∥x0 − x∗∥2 2(T − ¯T (M ) + 1)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:47:46.527885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:47:46.392025Z digest=sha256:9289cbdff27de79644d8ad286b03042769fa756fe7e96d508167b48efef43658

Observation 24cac4c9-6c86-493e-abd1-81a08474564f · outbound

This paper cites This function is (3.3, 1)–smooth, meaning we can run Algorithm 1 with ℓ(s) = 3.3 + s.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness This function is (3.3, 1)–smooth, meaning we can run Algorithm 1 with ℓ(s) = 3.3 + s

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T14:47:46.583136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:47:46.374599Z digest=sha256:657c1c63b4d2ab5a7ae9d7981f773760cc750c67d91470f4f595fe2e5ce71b6e

Observation 94b02397-078e-4add-bcde-fa63625f7551 · outbound

This paper cites Large Deviations of Vector-valued Martingales in 2-Smooth Normed Spaces.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Large Deviations of Vector-valued Martingales in 2-Smooth Normed Spaces

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T14:47:46.340297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:47:46.340297Z digest=sha256:fd925766cbf791f0e9add0ae7f97873087dd4132f96e7cf9078b60e8d0838461

Observation 798fd856-6e30-4944-bdbe-37cc476a9141 · outbound

This paper cites Federated Learning: Strategies for Improving Communication Efficiency.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Federated Learning: Strategies for Improving Communication Efficiency

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T14:47:46.344761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:47:46.344761Z digest=sha256:9c0a911b1241dc75e4030dfb9d0fa3f1d655583dc7c944f721fbc3e6a2c29321

Observation e65d4d06-6939-4437-879b-52de45629bef · outbound

This paper cites Optimizing $(L_0, L_1)$-Smooth Functions by Gradient Methods.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Optimizing $(L_0, L_1)$-Smooth Functions by Gradient Methods

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T14:47:46.353647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:47:46.353647Z digest=sha256:eadeccb5127640a833d18fbb564b22d2809dc2af47fda5623b54747309cd4aad

Observation 396e8e86-8831-423d-8707-c84d8eb55bd0 · outbound

This paper cites Gradient-Variation Online Learning under Generalized Smoothness.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Gradient-Variation Online Learning under Generalized Smoothness

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-11T14:47:46.446044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:47:46.358018Z digest=sha256:4c3a32b702ffe885d57df18a0695e58cd148ba8a1389b9b6787de91ca5a8ac61

Observation 96fe0dda-9a44-4369-a088-4427ece666d3 · outbound

This paper cites Why gradient clipping accelerates training: A theoretical justification for adaptivity.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Why gradient clipping accelerates training: A theoretical justification for adaptivity

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T14:47:46.361895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:47:46.361895Z digest=sha256:c8e507e3e49a683f713980ef6ca0edfd35a7440a5fe16c1c9d9612127c874ecb

Observation 4eb5ac19-5ad3-4592-93e8-d45fd3bd0718 · outbound

This paper cites Next, we take the step size γk = 1/(800 + 2(2f ′(x0))2) from (Li et al., 2024a) and observe that GD requires at least 20.000 iterations because f ′(x0) is huge.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Next, we take the step size γk = 1/(800 + 2(2f ′(x0))2) from (Li et al., 2024a) and observe that GD requires at least 20.000 iterations because f ′(x0) is huge

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:47:46.595659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:47:46.370237Z digest=sha256:d0620380ba3b49bfe087b8f75bd5300d2238fec09edbea98fea764f87c3ec51a

Observation de1bb650-520f-4772-b690-52acced1c794 · outbound

This paper cites Using the standard differential algebra, we can solve it: dg(t) ℓ(∥∇f (x)∥ + g(t)) = dt ⇒ Z t 0 dg(v) ℓ(∥∇f (x)∥ + g(v)) = t ⇒ Z g(t) 0 dv ℓ(∥∇f (x)∥ + v) = t.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Using the standard differential algebra, we can solve it: dg(t) ℓ(∥∇f (x)∥ + g(t)) = dt ⇒ Z t 0 dg(v) ℓ(∥∇f (x)∥ + g(v)) = t ⇒ Z g(t) 0 dv ℓ(∥∇f (x)∥ + v) = t

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:47:46.569462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:47:46.378443Z digest=sha256:20140445a9b324a2463f40ebc802a370742f2292d5df9cd8f4bb049de3f3748c

Observation 43edb0f4-f171-4a37-9873-2c0ff80b75b9 · outbound

This paper cites an unresolved cited work.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:47:46.555804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:47:46.383119Z digest=sha256:f8e957f21bc1b40a5ac59100697bd29a028121eaebd69e23874384ad8a26302b

Observation 993fdb7d-0f8e-4e28-a9d8-7bc75035d8a6 · outbound

This paper cites Due to the strategy from Alg.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Due to the strategy from Alg

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:47:46.542435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:47:46.387503Z digest=sha256:2edc411a2866f5d0dd799331308ccfea6f8cd28867a8952cc1c8085ab0ff566a

Observation 0892a7b6-1813-4d05-ad69-5c4ba1b58b0a · outbound

This paper cites Parameter-free Clipped Gradient Descent Meets Polyak.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Parameter-free Clipped Gradient Descent Meets Polyak

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T14:47:46.349125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:47:46.349125Z digest=sha256:34798617a1495ad3a594ce18762179884bcdd8d91433d7998f0a4ba95c9e952e

Observation 9dd21d60-127c-4ee0-b65e-54bccbb5aa1a · outbound

This paper cites an unresolved cited work.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Unresolved cited work

Reference 2019

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:47:46.608869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:47:46.366168Z digest=sha256:7eba69de623c3744815cf72256a3b4fb1ad2ac174683d540506f51915f39470e

Observation 8dd84dcc-81f2-49d3-a939-28321be14537 · outbound

This paper cites Methods for Convex $(L_0,L_1)$-Smooth Optimization: Clipping, Acceleration, and Adaptivity.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Methods for Convex $(L_0,L_1)$-Smooth Optimization: Clipping, Acceleration, and Adaptivity

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T14:47:46.330224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:47:46.330224Z digest=sha256:5e924b16b88707974fa29172c87bca624e17860604062687494a29759cc8305b

Observation d5a437c0-4c84-4287-b832-013d350911c3 · outbound

This paper cites A theoretical study of the(l 0, l1)-smoothness condition in deep learning.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness A theoretical study of the(l 0, l1)-smoothness condition in deep learning

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:47:46.621303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:47:46.324863Z digest=sha256:38dc6e31509e12f03994bb354fde9681b1813b0dde2854efa9f58c56ba2c73b3

Observation dac0dda4-a466-4f9c-9efa-70e19b249496 · outbound

This paper cites Accelerated Objective Gap and Gradient Norm Convergence for Gradient Descent via Long Steps.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Accelerated Objective Gap and Gradient Norm Convergence for Gradient Descent via Long Steps

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T14:47:46.334977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:47:46.334977Z digest=sha256:4a4eb7ea6781ddf55636510ebc81c867f3c966cbde0e70f189838ecfa9e6a643

Pith citing papers

No inbound Pith citation observations are available.