Pith. sign in

REVIEW 4 cited by

Stochastic Gradient Methods with Layer-wise Adaptive Moments for Training of Deep Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1905.11286 v3 pith:QMIP3DZS submitted 2019-05-27 cs.LG stat.ML

classification cs.LGstat.ML
keywords gradientadamadaptivelayer-wisenetworksnovogradstochasticweight
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose NovoGrad, an adaptive stochastic gradient descent method with layer-wise gradient normalization and decoupled weight decay. In our experiments on neural networks for image classification, speech recognition, machine translation, and language modeling, it performs on par or better than well tuned SGD with momentum and Adam or AdamW. Additionally, NovoGrad (1) is robust to the choice of learning rate and weight initialization, (2) works well in a large batch setting, and (3) has two times smaller memory footprint than Adam.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A meta-pipeline plus LMO four-axis view yields a dual taxonomy of 108 optimizers, and a multi-objective LLM/vision benchmark shows no single family dominates the quality–cost–memory frontier.

  2. Learning state and proposal dynamics in state-space models using differentiable particle filters and neural networks

    cs.LG 2024-11 conditional novelty 6.0 of 10

    StateMixNN learns particle-filter transition and proposal densities as Gaussian mixtures parameterized by neural networks, trained only on the observation likelihood, and reports improved state recovery on Lorenz 96 a...

  3. SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training

    cs.LG 2025-01 conditional novelty 5.0 of 10

    SPAM, an Adam variant with periodic momentum reset and ratio-based spike clipping, reports better validation perplexity and lower memory use than Adam, Adafactor, GaLore, and Adam-mini in LLM pretraining and fine-tuning.

  4. GraphGrad: Efficient Estimation of Sparse Polynomial Representations for General State-Space Models

    stat.CO 2024-11 conditional novelty 5.0 of 10

    A differentiable particle filter with L1 proximal updates estimates sparse polynomial transition functions and interaction graphs for nonlinear state-space models.

Pith tools