REVIEW 4 cited by
An Adaptive Stochastic Gradient Method with Non-negative Gauss-Newton Stepsizes
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We consider the problem of minimizing the average of a large number of smooth but possibly non-convex functions. In the context of most machine learning applications, each loss function is non-negative and thus can be expressed as the composition of a square and its real-valued square root. This reformulation allows us to apply the Gauss-Newton method, or the Levenberg-Marquardt method when adding a quadratic regularization. The resulting algorithm, while being computationally as efficient as the vanilla stochastic gradient method, is highly adaptive and can automatically warmup and decay the effective stepsize while tracking the non-negative loss landscape. We provide a tight convergence analysis, leveraging new techniques, in the stochastic convex and non-convex settings. In particular, in the convex case, the method does not require access to the gradient Lipshitz constant for convergence, and is guaranteed to never diverge. The convergence rates and empirical evaluations compare favorably to the classical (stochastic) gradient method as well as to several other adaptive methods.
Forward citations
Cited by 4 Pith papers
-
Safeguarded Stochastic Polyak Step Sizes for Non-smooth Optimization: Robust Performance Without Small (Sub)Gradients
A safeguarded stochastic Polyak step size, SPS_safe, yields O(1/√T) convergence to a neighborhood for convex non-smooth problems without interpolation or oracle loss values, with a momentum variant.
-
LionVote: Per-Layer Learning Rate Adaptation for Lion
Per-layer LionVote shows Lion's prescribed rate is 2–2.8× too high on ViT layer types and lifts ViT-Tiny/CIFAR-100 accuracy to 69.7% vs Lion's 69.0%.
-
Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size
NGN-M, a momentum variant of the NGN step-size, provably converges at O(1/sqrt(K)) under milder assumptions and shows wider step-size stability than Adam, Momo, and SGDM in vision and language tasks.
-
Learning by solving differential equations
Runge-Kutta optimizers adapted with momentum, preconditioning, or adaptive learning rates can close the large-batch generalization gap and match Adam on small MLP workloads.
Discussion (0). Continue with ORCID to comment.