Pith. sign in

REVIEW 3 cited by

AdaGrad under Anisotropic Smoothness

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.15244 v2 pith:E46AT2N3 submitted 2024-06-21 cs.LG math.OC

classification cs.LGmath.OC
keywords adagradanisotropiclargepracticesmoothnesstheoreticalacrossalgorithms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Adaptive gradient methods have been widely adopted in training large-scale deep neural networks, especially large foundation models. Despite the huge success in practice, their theoretical advantages over classical gradient methods with uniform step sizes across all coordinates (e.g. SGD) have not been fully understood, especially in the large batch-size setting commonly used in practice. This is because the only theoretical result that can demonstrate this benefit was obtained in the original paper of Adagrad for convex nonsmooth objective functions, which is insufficient for large batch algorithms. In this work, we attempt to resolve this gap between theory and practice by proposing a novel anisotropic generalized smoothness assumption and providing corresponding analyses of Adagrad. It is shown that under anisotropic smoothness and noise conditions, AdaGrad can achieve faster convergence guarantees in terms of better dimensional dependence than algorithms with uniform step sizes across all coordinates. Experiments in logistic regression and instruction following fine-tuning tasks provide strong evidence to support our novel assumption and theoretical analysis.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SGD with Adaptive Preconditioning: Unified Analysis and Momentum Acceleration

    cs.LG 2025-06 conditional novelty 8.0 of 10

    A single proof unifies convergence analyses of AdaGrad-Norm, AdaGrad, ASGO, and DASGO under Hölder smoothness, and shows AdaGrad/DASGO can be accelerated with Nesterov momentum.

  2. Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence

    cs.LG 2025-09 conditional novelty 7.0 of 10

    Under a new smoothness assumption linking local curvature to the loss gap, increasing learning rates provably accelerate GD and SGD convergence, with up to Theta(T) speedup in special cases.

  3. Non-Euclidean SGD for Structured Optimization: Unified Analysis and Improved Rates

    math.OC 2025-11 conditional novelty 5.0 of 10

    Non-Euclidean SGD variants (SignSGD, Muon) provably match adaptive optimizers' convergence rates under structured smoothness and noise assumptions.

Pith tools