REVIEW 4 cited by
Stochastic Gradient Methods with Layer-wise Adaptive Moments for Training of Deep Networks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We propose NovoGrad, an adaptive stochastic gradient descent method with layer-wise gradient normalization and decoupled weight decay. In our experiments on neural networks for image classification, speech recognition, machine translation, and language modeling, it performs on par or better than well tuned SGD with momentum and Adam or AdamW. Additionally, NovoGrad (1) is robust to the choice of learning rate and weight initialization, (2) works well in a large batch setting, and (3) has two times smaller memory footprint than Adam.
Forward citations
Cited by 4 Pith papers
-
OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers
A meta-pipeline plus LMO four-axis view yields a dual taxonomy of 108 optimizers, and a multi-objective LLM/vision benchmark shows no single family dominates the quality–cost–memory frontier.
-
Learning state and proposal dynamics in state-space models using differentiable particle filters and neural networks
StateMixNN learns particle-filter transition and proposal densities as Gaussian mixtures parameterized by neural networks, trained only on the observation likelihood, and reports improved state recovery on Lorenz 96 a...
-
SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training
SPAM, an Adam variant with periodic momentum reset and ratio-based spike clipping, reports better validation perplexity and lower memory use than Adam, Adafactor, GaLore, and Adam-mini in LLM pretraining and fine-tuning.
-
GraphGrad: Efficient Estimation of Sparse Polynomial Representations for General State-Space Models
A differentiable particle filter with L1 proximal updates estimates sparse polynomial transition functions and interaction graphs for nonlinear state-space models.
Discussion (0). Continue with ORCID to comment.