Pith. sign in

REVIEW 15 cited by

On the Variance of the Adaptive Learning Rate and Beyond

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1908.03265 v4 pith:QZOJLYIM submitted 2019-08-08 cs.LG cs.CLstat.ML

classification cs.LGcs.CLstat.ML
keywords adaptivelearningratevariancewarmupadamradamverify
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The learning rate warmup heuristic achieves remarkable success in stabilizing training, accelerating convergence and improving generalization for adaptive stochastic optimization algorithms like RMSprop and Adam. Here, we study its mechanism in details. Pursuing the theory behind warmup, we identify a problem of the adaptive learning rate (i.e., it has problematically large variance in the early stage), suggest warmup works as a variance reduction technique, and provide both empirical and theoretical evidence to verify our hypothesis. We further propose RAdam, a new variant of Adam, by introducing a term to rectify the variance of the adaptive learning rate. Extensive experimental results on image classification, language modeling, and neural machine translation verify our intuition and demonstrate the effectiveness and robustness of our proposed method. All implementations are available at: https://github.com/LiyuanLucasLiu/RAdam.

Discussion (0). Sign in to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 608 citations worldwide. Full citation record

  1. Unified convergence analysis for gradient descent optimization methods in the training of deep neural networks

    math.OC 2026-07 accept novelty 7.0 of 10

    Bounded trajectories of a broad class of GD optimizers (Adam, RMSprop, NAG, Adan, etc.) converge with polynomial rates to critical points of KL objectives with locally Lipschitz gradients, covering analytic-activation...

  2. HealthCAT: An Interpretable Encoder-only Transformer Framework for Health Indicator Prediction and Temporal Interpretation of Wearable Sensor Data

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A transformer with AttentiveCAT yields class-specific, time-step importance scores for wearable health data, beats deep-learning baselines, and beats random time-step selection in masking tests.

  3. Amortized Moment Matching for Visual Generation

    cs.LG 2026-07 accept novelty 6.0 of 10

    Amortized Fréchet Distance uses neural nets to match conditional means and covariances, yielding stronger one-step visual generators than explicit FD-loss or multi-step teachers.

  4. Payne4GAIN: NLTE Corrections for Red Giants in Milky Way Mapper using H-Band Neural Network Emulators

    astro-ph.SR 2026-07 conditional novelty 6.0 of 10

    Applying NLTE physics to APOGEE H-band spectra shifts red-giant abundances by ~0.1 dex for Al, Mn, and Ti; the paper provides a 360k-star correction catalog.

  5. A machine-learned probability distribution in the phase space of turbulent channel flow for synthetic turbulence and flow reconstruction

    physics.flu-dyn 2026-07 conditional novelty 6.0 of 10

    A flow-matching generative model trained on minimal conditional flow units approximates the invariant phase-space distribution of turbulent channel flow at Re_tau=180, enabling synthetic turbulence generation and flow...

  6. OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A meta-pipeline plus LMO four-axis view yields a dual taxonomy of 108 optimizers, and a multi-objective LLM/vision benchmark shows no single family dominates the quality–cost–memory frontier.

  7. Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    Splitting weight matrices into a fixed-norm direction and learnable per-row/column magnitudes improves LLM training over AdamW/Muon, removes weight decay and warmup, and transfers the optimal LR across width.

  8. TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning

    cs.MA 2026-02 unverdicted novelty 6.0 of 10

    TABX is a JAX-based, GPU-accelerated, configurable multi-agent battle simulator that lets researchers vary units, terrain, and physics to benchmark cooperative MARL algorithms.

  9. From Next Token Prediction to (STRIPS) World Models

    cs.AI 2025-09 unverdicted novelty 6.0 of 10

    Transformers trained via next-token prediction on action traces can learn STRIPS action models that support planning over exponentially many unseen initial states and goals.

  10. Deep Potential: Recovering the gravitational potential and local pattern speed in the solar neighborhood with GDR3 using normalizing flows

    astro-ph.GA 2025-07 conditional novelty 6.0 of 10

    Using normalizing flows and a neural network on Gaia DR3 data, the authors recover a local pattern speed of 28.2 km/s/kpc and a total matter density of 0.086 solar masses per cubic parsec within 1 kpc of the Sun.

  11. GCR Spectra Reconstructed with Neutron Monitor Yield Function and Artificial Neural Networks: Comparison of Two Methods

    astro-ph.IM 2026-07 conditional novelty 5.0 of 10

    Neural networks trained on worldwide neutron-monitor counts plus solar indices reconstruct daily proton and helium cosmic-ray spectra for 2006–2022, matching PAMELA and AMS-02 data and beating a yield-function/force-f...

  12. Why Do We Need Warm-up? A Theoretical Perspective

    cs.LG 2025-10 conditional novelty 5.0 of 10

    Under the proposed (H0,H1)-smoothness condition, gradient descent with a warm-up-style adaptive step-size provably converges faster than with any fixed step-size.

  13. Criticality analysis of nuclear binding energy neural networks

    nucl-th 2025-08 conditional novelty 5.0 of 10

    On a two-input nuclear binding energy network, the paper validates ANNFT predictions for variance, kurtosis, and an optimal depth-to-width ratio r*=0.034 under SGD, while adaptive optimizers obscure criticality.

  14. SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam

    cs.LG 2025-07 reject novelty 5.0 of 10

    S3, an optimizer with a p-th order momentum denominator, equal EMA coefficients, and Nesterov acceleration, is claimed to match AdamW's 100k-step perplexity at 50k steps while avoiding loss spikes.

  15. Feature-Enhanced TResNet for Fine-Grained Food Image Classification

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A TResNet variant with style recalibration and criss-cross-style attention reports modest Top-1 accuracy gains on two Chinese food benchmarks.

Pith tools