Pith. sign in

Paper Citation Record · LEDGER

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping

As of 16 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 3 inbound Pith citation observations for arXiv:2412.19529.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.19529 v4

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:23:05.364651Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:22:55.341832Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T05:37:19.252093Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact1
  • verified fuzzy29
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0fc28118-baf3-49a6-b1eb-9b8ff1e4a8eb · outbound

This paper cites write newline.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.160353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.160353Z digest=sha256:4a3590f217b45725c515014ed18523d2fda6f601c705c9d883655df51d4d17e5

Observation 24bcfccf-0945-4922-83e1-30dc3419e7ef · outbound

This paper cites Lower bounds for non-convex stochastic optimization.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Lower bounds for non-convex stochastic optimization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.165327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.165327Z digest=sha256:2e874aee447501d7e5314b3ab2a03a5b89cbed0141afce80679d6795eb3a3a69

Observation d3ad11f4-6b02-432f-9a2e-524cc9b75066 · outbound

This paper cites Optimization methods for large-scale machine learning.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Optimization methods for large-scale machine learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.170048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.170048Z digest=sha256:6a9ef4ded806b0dd49acf5515ce7436c8b9c4f5389c1edec006b26f4ab55e40d

Observation fc9be62a-f9e0-4c6e-9034-0fac49ad27a5 · outbound

This paper cites Extrapolation and interpolation of quasi-linear operators on martingales.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Extrapolation and interpolation of quasi-linear operators on martingales

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:06.005562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.174249Z digest=sha256:7d1731f5cc85ef3e06e37db40081c5f21cceab8c1e6b65b02808b70d0da61b91

Observation a04d6d5b-be21-48a0-9ad1-7c5435dcad12 · outbound

This paper cites Martingale transforms.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Martingale transforms

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.994834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.178054Z digest=sha256:a270627a367af409bc867d56d115dccfe060e1871f38e7cd7cb5f6154d3d1686

Observation ad4242fa-2a07-4bc1-be10-f72ca3ff91d1 · outbound

This paper cites Lower bounds for finding stationary points i.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Lower bounds for finding stationary points i

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.983749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.181956Z digest=sha256:c239dc665a92cb48f8b4e41f48d9703b87e81719797831333c3718b3eab01463

Observation 2b4ca97b-f132-46bd-9581-fe7ede0f56e8 · outbound

This paper cites Generalized-smooth nonconvex optimization is as efficient as smooth nonconvex optimization.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Generalized-smooth nonconvex optimization is as efficient as smooth nonconvex optimization

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.971418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.185848Z digest=sha256:ea924f4c1d4b9e65b828dbda026b0be0d24034de74508fa84a326cfa84ce7e8b

Observation ad5105ba-d011-475e-91a7-7c4bba9741f7 · outbound

This paper cites Robustness to unbounded smoothness of generalized signsgd.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Robustness to unbounded smoothness of generalized signsgd

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.960483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.190382Z digest=sha256:fc3529e692166d0d8c0ad0c56133eff2f49c836d13f3a43e62e71a054219bf6a

Observation 2c750b12-965d-46cb-9e03-0e660ef42e10 · outbound

This paper cites Momentum improves normalized SGD.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Momentum improves normalized SGD

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.949463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.195061Z digest=sha256:891fc5a5a28ab1ab22f21fff4d16b5fa5caffff81043cd16951b1e532556a012

Observation a16ab0f0-9084-4c36-aa29-4531ae374c28 · outbound

This paper cites High-probability bounds for non-convex stochastic optimization with heavy tails.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping High-probability bounds for non-convex stochastic optimization with heavy tails

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.938843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.198729Z digest=sha256:ffd0b8e6c6e8f7b624d6a9b59a5fdeaee7676d6720a940b38640f31cc2ac745c

Observation 80b17b9a-bb29-414c-9873-9c2b50e97c26 · outbound

This paper cites On the intergrability of the martingale square function.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping On the intergrability of the martingale square function

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.928488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.202455Z digest=sha256:6c366a9dc3e9fe5d0a70ac744b7eb41b9d6f232e0eecaef713c2c2708d7ed397

Observation afd19ac1-36ca-4c2c-8f98-181502c26e4b · outbound

This paper cites Adaptive subgradient methods for online learning and stochastic optimization.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Adaptive subgradient methods for online learning and stochastic optimization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.206190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.206190Z digest=sha256:2b0572461f727bbbcacea974e308a3b4fecfec6a1e987d9d9abcbc0743213319

Observation 58aea4d7-3ec7-4d56-8a4e-71a109f52808 · outbound

This paper cites Beyond uniform smoothness: A stopped analysis of adaptive sgd.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Beyond uniform smoothness: A stopped analysis of adaptive sgd

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.911746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.209701Z digest=sha256:19aa0cc28bfb1826b5aead79561a2ffe6d5e1420858d10d4e15e8c3285adcea7

Observation 35ad4cc9-880d-4bf3-8011-0a57d2d20282 · outbound

This paper cites Stochastic first-and zeroth-order methods for nonconvex stochastic programming.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Stochastic first-and zeroth-order methods for nonconvex stochastic programming

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.213244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.213244Z digest=sha256:5fb419d9c03f97f5da72a3a681094f57ec467e0b6b6d5c276feba8d981002cf4

Observation f6c1cbb4-d6d2-4e45-a082-e09f526ac253 · outbound

This paper cites High-probability convergence for composite and distributed stochastic minimization and variational inequalities with heavy-tailed noise.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping High-probability convergence for composite and distributed stochastic minimization and variational inequalities with heavy-tailed noise

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.895550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.216626Z digest=sha256:daad87852e26ccdba7003e0936d29ca33ccfcf28c7719e1443704683112ef7bd

Observation 75b954ad-2c0f-4a4a-a7f2-e71bf02b0ded · outbound

This paper cites Beyond convexity: Stochastic quasi-convex optimization.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Beyond convexity: Stochastic quasi-convex optimization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.220660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.220660Z digest=sha256:7e2472de2c33f5637cef1a65e09c6b140b391dbbf065cd25378e45f97a95abae

Observation ee49d851-ff89-406e-a56a-2afb4caed0c7 · outbound

This paper cites Revisiting Convergence of AdaGrad with Relaxed Assumptions.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Revisiting Convergence of AdaGrad with Relaxed Assumptions

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.225264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.225264Z digest=sha256:7171215d9aba32793fc7e3ae3e3ff2409052130897cbadd5504e4024723ae70a

Observation 1a407fcb-ee45-407f-bae5-678c604665f2 · outbound

This paper cites From Gradient Clipping to Normalization for Heavy Tailed SGD.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping From Gradient Clipping to Normalization for Heavy Tailed SGD

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.229100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.229100Z digest=sha256:48daf216cc28021959954e9ad02fb646311c30f62ce8830f0d253b072a906238

Observation 8ee79d4f-0f3a-4261-827b-e32fd7b80bba · outbound

This paper cites Parameter-agnostic optimization under relaxed smoothness.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Parameter-agnostic optimization under relaxed smoothness

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.879086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.233151Z digest=sha256:0af73757c8d4e79eec6544f112ba41bc7b44ba68442cfd7548abc3b747873c93

Observation 77125654-26d1-4b33-8fa1-ae009d22de61 · outbound

This paper cites Non-convex distributionally robust optimization: Non-asymptotic analysis.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Non-convex distributionally robust optimization: Non-asymptotic analysis

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.868128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.236891Z digest=sha256:85b5a70ab2a733e323843562ef950fad03a0f90d8643bbfa393be1a5ca608180

Observation 46a386e5-9403-452e-a034-af1360ef99b3 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Adam: A Method for Stochastic Optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.240285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.240285Z digest=sha256:63c1bc26be8d17a2b544ccccb9cf37cb8549d355316ad9fe8e2e6fad2f044e13

Observation 638ce608-afd1-4407-9a37-5c6c522e82ea · outbound

This paper cites Convergence and efficiency of subgradient methods for quasiconvex minimization.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Convergence and efficiency of subgradient methods for quasiconvex minimization

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.857095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.243723Z digest=sha256:d2e09deb27d211c2e554c75df5358598fb5ae7c41b36ec07a27f9fa8bc92de37

Observation 5fa88c4b-f563-4aae-bc01-ea316aca6af1 · outbound

This paper cites Accelerated zeroth-order method for non-smooth stochastic convex optimization problem with infinite variance.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Accelerated zeroth-order method for non-smooth stochastic convex optimization problem with infinite variance

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.845941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.247430Z digest=sha256:63fbddd70827917feb1a5f52119c296696d5bae7f60442fe2dbe0d14e2761d9f

Observation ebd65b1a-f0c5-4cf6-a42e-a3ce8600a10e · outbound

This paper cites First-order and stochastic optimization methods for machine learning.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping First-order and stochastic optimization methods for machine learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.250868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.250868Z digest=sha256:38c034ff469c4a6928bd8f3953ca8d7bd53ee5343ef2550aa8999e720fa2cf7e

Observation a0ca87c4-de89-4b4b-af29-4e7be5e90b01 · outbound

This paper cites The Power of Normalization: Faster Evasion of Saddle Points.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping The Power of Normalization: Faster Evasion of Saddle Points

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.255061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.255061Z digest=sha256:d2c888f9900d011710fb1284cc36126164b1a6aaf0125e98fa4a4e3b62945f4e

Observation 132f5c5f-bb67-4e2a-80ea-35306b4bf7d1 · outbound

This paper cites Convex and non-convex optimization under generalized smoothness.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Convex and non-convex optimization under generalized smoothness

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.828707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.258995Z digest=sha256:b68ae15f93ae2971937c76e7f8c506b7cfd6d9715b52015a316e4338870f570f

Observation e3645a8f-7a53-4e4c-978f-e0e5c03231de · outbound

This paper cites Convergence of adam under relaxed assumptions.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Convergence of adam under relaxed assumptions

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.817103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.262381Z digest=sha256:261f9ded9b212bd82d939e81419cbb836c1cf0f55b720d85d354fdc3bb77c8d5

Observation 81e46ea8-f965-42e2-9c13-137eec65ac64 · outbound

This paper cites High-probability bound for non-smooth non-convex stochastic optimization with heavy tails.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping High-probability bound for non-smooth non-convex stochastic optimization with heavy tails

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.806122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.266084Z digest=sha256:d6770665521a03879108d2a8f6fceb240efd1752617da05add5d40e182c4dfb6

Observation 9b9e4e8e-941d-4737-9bfa-7d0161702bc8 · outbound

This paper cites Stochastic Nonsmooth Convex Optimization with Heavy-Tailed Noises: High-Probability Bound, In-Expectation Rate and Initial Distance Adaptation.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Stochastic Nonsmooth Convex Optimization with Heavy-Tailed Noises: High-Probability Bound, In-Expectation Rate and Initial Distance Adaptation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.269460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.269460Z digest=sha256:6646502481be8728d7598aaa097668863d888ac96317b5997dac5db38373e039

Observation 011f68f0-915a-4656-ba06-fdcf2690a83b · outbound

This paper cites Near-Optimal Non-Convex Stochastic Optimization under Generalized Smoothness.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Near-Optimal Non-Convex Stochastic Optimization under Generalized Smoothness

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-11T00:23:05.562657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.273144Z digest=sha256:a101495ea3a32987528b723032d888e4b3c74637ecc8a240346e7e64fa9a3bcd

Observation 80e220e5-3bbd-4626-8174-1ce74974eca4 · outbound

This paper cites Breaking the lower bound with (little) structure: Acceleration in non-convex stochastic optimization with heavy-tailed noise.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Breaking the lower bound with (little) structure: Acceleration in non-convex stochastic optimization with heavy-tailed noise

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.795062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.276793Z digest=sha256:9579de94df88afed8f436e4b0cf5ac09df5681cecc842af09d2961527190846c

Observation 8c848301-a694-4ceb-a6c3-bd1c76924c35 · outbound

This paper cites Brendan McMahan and Matthew J.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Brendan McMahan and Matthew J

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.783834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.280377Z digest=sha256:0e688feec9a58712edd9f690268832e730a1d5ab895004b40e537fca5b3e7705

Observation 95a200f2-69dc-4234-98d9-d6fb1b0aa4cd · outbound

This paper cites Convergence of gradient descent on separable data.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Convergence of gradient descent on separable data

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.773141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.283937Z digest=sha256:4e4b62ea2c6daa3219ce2d4872dd99525fae3edbfc4f8c76fdc0fd654c1dea19

Observation 339e4444-5c6e-4323-ac52-d31237d0a8d7 · outbound

This paper cites Lectures on convex optimization, volume 137.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Lectures on convex optimization, volume 137

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.287926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.287926Z digest=sha256:fc19165b1bc3bb6a57b436ed2bf3d487dddb11e4e30295636a937adec80c6c61

Observation fd650228-1fa1-4104-994f-d706340f71d2 · outbound

This paper cites Minimization methods for nonsmooth convex and quasiconvex functions.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Minimization methods for nonsmooth convex and quasiconvex functions

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.291488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.291488Z digest=sha256:b4c494b108cc35323d7075d964ebb4a1b65d8e2567a031e6480b9ab4c7264954

Observation e66b33ec-86c6-4d0d-a181-cb0598653329 · outbound

This paper cites Improved convergence in high probability of clipped gradient methods with heavy tailed noise.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Improved convergence in high probability of clipped gradient methods with heavy tailed noise

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.747536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.296096Z digest=sha256:6aee5b4d1bc08a60d733fff0dfc5715213c8318945e5402333f0f9a1781e2f93

Observation 62e710e2-5c40-4598-b580-5d1b39161451 · outbound

This paper cites Breaking the heavy-tailed noise barrier in stochastic optimization problems.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Breaking the heavy-tailed noise barrier in stochastic optimization problems

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.735972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.299813Z digest=sha256:b577d363aad9da7cba5ea59f553c5049709427dd77d84e9df54f8a8b3baac331

Observation 5b44bcd0-0498-4aee-ab37-f4294002ddf1 · outbound

This paper cites On equivalence of martingale tail bounds and deterministic regret inequalities.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping On equivalence of martingale tail bounds and deterministic regret inequalities

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.725144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.303409Z digest=sha256:756dfa7ffb12de78d2a4a94f1cd78d594f46d7b2d07cfb7d7721b5abb2a7661a

Observation 1f43cce4-4eeb-450d-bbb9-3781d43cb2a0 · outbound

This paper cites A stochastic approximation method.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping A stochastic approximation method

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.307200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.307200Z digest=sha256:a7e2eb8230360e49098f31c7053026e62d034af691c3a9e2475f68e296e847b9

Observation 262ab119-b5b0-4809-9591-85ae4bcf5d8f · outbound

This paper cites High-probability bounds for stochastic optimization and variational inequalities: the case of unbounded variance.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping High-probability bounds for stochastic optimization and variational inequalities: the case of unbounded variance

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.707260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.310737Z digest=sha256:e13872ffab044fc2539f159aa280362d0cbc0661cd7af461de563927f91871d4

Observation 13d035c6-4109-4359-8599-e2fe27fa8e5d · outbound

This paper cites On the Heavy-Tailed Theory of Stochastic Gradient Descent for Deep Neural Networks.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping On the Heavy-Tailed Theory of Stochastic Gradient Descent for Deep Neural Networks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.314264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.314264Z digest=sha256:616be9973ebc28e3b1d12713daaaeadba8b5e2ca2749b5e9decd424cde712190

Observation d6f61eb8-111b-476c-acda-1c1211037043 · outbound

This paper cites A tail-index analysis of stochastic gradient noise in deep neural networks.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping A tail-index analysis of stochastic gradient noise in deep neural networks

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.696259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.318112Z digest=sha256:b85abf8daa00c2ade8791735476a0f45870b78d496bd992318c212ee5ce04f5d

Observation 0afb4bcc-8ae6-46d0-b042-7fc16b8c506d · outbound

This paper cites Gradient normalization provably benefits nonconvex sgd under heavy-tailed noise.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Gradient normalization provably benefits nonconvex sgd under heavy-tailed noise

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.321472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.321472Z digest=sha256:b9a90c3ea9f925a564ca22b432b1bd9ed9d047e476d7014d94eb540b87bb2e2b

Observation d1a17774-fafa-4920-8f77-1e7301bea51e · outbound

This paper cites Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.325034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.325034Z digest=sha256:ad7a72c883e760b4c3bcf9504c3c27c4ef59439180aede44165c140fc424b478

Observation 30285a24-ec49-4c23-acc5-adf26fa2ae50 · outbound

This paper cites Convergence of adagrad for non-convex objectives: Simple proofs and relaxed assumptions.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Convergence of adagrad for non-convex objectives: Simple proofs and relaxed assumptions

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.679329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.328428Z digest=sha256:9088ace243345eb07c6348b5759fc2f6b4041fdd56295c4a96bfb352f1faaa74

Observation 3ea02deb-cf8d-4a4e-acb3-1c5831536187 · outbound

This paper cites On the Convergence of Adam under Non-uniform Smoothness: Separability from SGDM and Beyond.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping On the Convergence of Adam under Non-uniform Smoothness: Separability from SGDM and Beyond

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.332066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.332066Z digest=sha256:93754c07689adc1dbb968ba37223963cfe9a768d8db4ef0154df2455b42fa4ab

Observation 807b14e1-fd80-477c-832e-25c3c34487bf · outbound

This paper cites Large Batch Training of Convolutional Networks.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Large Batch Training of Convolutional Networks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.335736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.335736Z digest=sha256:6a245e8078c0324ec8a741377f61db743043c37686c2384d6b6a42fd7fff1a63

Observation 3b9e8fe4-1f53-4c19-af55-00a91c1fa4b4 · outbound

This paper cites Large Batch Optimization for Deep Learning: Training BERT in 76 minutes.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Large Batch Optimization for Deep Learning: Training BERT in 76 minutes

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.339526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.339526Z digest=sha256:41c00f5112a68a816ddb43bd55ddea74a62f3d9f63e4c429a80762c0a252b736

Observation dc5a44f9-b838-4056-9a91-5a5472e3f179 · outbound

This paper cites Improved analysis of clipping algorithms for non-convex optimization.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Improved analysis of clipping algorithms for non-convex optimization

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.343597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.343597Z digest=sha256:e9bde05ea9183b2ef89a34b95b2f41d93d76eb7ccba487872188b9a47cdfb9f4

Observation 41ec6a0c-d5c0-421d-b1dd-170400ff3825 · outbound

This paper cites Why gradient clipping accelerates training: A theoretical justification for adaptivity.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Why gradient clipping accelerates training: A theoretical justification for adaptivity

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.662393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.346964Z digest=sha256:f02a572c3345b25830a5d2f51178fbac2aa1e754a866c68628cae9f156e53e24

Observation 1e9764ac-5c63-4251-b2a9-30925eb9bfb2 · outbound

This paper cites Why are adaptive methods good for attention models? Advances in Neural Information Processing Systems, 33: 0 15383--15393, 2020 c.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Why are adaptive methods good for attention models? Advances in Neural Information Processing Systems, 33: 0 15383--15393, 2020 c

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.350245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.350245Z digest=sha256:7a94cc075b1591f4a1264d5095a50aa6f50b6f713e9e06d91e62c7bbfcb1a155

Observation b318596c-0bea-45b7-889a-c199c38b5511 · outbound

This paper cites Parameter-free regret in high probability with heavy tails.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Parameter-free regret in high probability with heavy tails

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.645301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.353870Z digest=sha256:582bb4da96edc2a3d1e1e78ddce7c9ccf40019c98fd668cc55c5e2edc3642882

Observation 60e2cdf6-d79d-4105-89e5-052ddd30fd9d · outbound

This paper cites @esa (Ref.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping @esa (Ref

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.357279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.357279Z digest=sha256:a4bcb6c66058e34f80fbb981a7bc7497ecd6e4748614c6234e37ecb16a968ab6

Observation e5e25932-8001-4f9d-a7f1-8691e04d1534 · outbound

This paper cites an unresolved cited work.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.360942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.360942Z digest=sha256:45ffce3f5f82a10a7440fb6c0ae083161f74be35ba1f5695c9c1e80a87002182

Observation e6d7eb5e-8a71-4dee-8029-dab6d8102179 · outbound

This paper cites denotes the set of natural numbers (excluding 0 ).

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping denotes the set of natural numbers (excluding 0 )

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.621655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.364651Z digest=sha256:6533a56e1071b66ca89d0190403c423c6928a8981c2fa247677d301fe304dc34

Pith citing papers

Observation 587bb326-e20d-4e63-b705-90653988a7e6 · inbound

Revisiting Randomized Smoothing: Nonsmooth Nonconvex Optimization Beyond Global Lipschitz Continuity cites this paper.

Revisiting Randomized Smoothing: Nonsmooth Nonconvex Optimization Beyond Global Lipschitz Continuity Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.341832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.341832Z digest=sha256:ace9c131e851bc0533ca6cffe06808348228b9c0d7e4badb9f4747e8beec6308

Observation bc3e5b1d-5c10-43f1-ba8d-8d9dafd4b9a5 · inbound

Decentralized Stochastic Nonconvex Optimization under the $(L_0,L_1)$-Smoothness cites this paper.

Decentralized Stochastic Nonconvex Optimization under the $(L_0,L_1)$-Smoothness Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T16:18:52.155024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:18:52.155024Z digest=sha256:854431dee7e95f636af0193c3f53ca96687f7440106b6d3ef30b951146a4163e

Observation 411c530c-4a62-4b55-a653-5c317d1eed7e · inbound

Constrained Stochastic Spectral Preconditioning Converges for Nonconvex Objectives cites this paper.

Constrained Stochastic Spectral Preconditioning Converges for Nonconvex Objectives Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:37:19.253844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T05:34:48.195468Z digest=sha256:19c6d414935b2d0ba02820cc8675a1214d0c5782a5e358fb71a85c27300c6091