Pith. sign in

REVIEW 6 cited by

Exponential Moving Average of Weights in Deep Learning: Dynamics and Benefits

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.18704 v1 pith:WD6D2JQT submitted 2024-11-27 cs.LG

classification cs.LG
keywords learningaveragingdeepmodelstrainingweightsaveragedynamics
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Weight averaging of Stochastic Gradient Descent (SGD) iterates is a popular method for training deep learning models. While it is often used as part of complex training pipelines to improve generalization or serve as a `teacher' model, weight averaging lacks proper evaluation on its own. In this work, we present a systematic study of the Exponential Moving Average (EMA) of weights. We first explore the training dynamics of EMA, give guidelines for hyperparameter tuning, and highlight its good early performance, partly explaining its success as a teacher. We also observe that EMA requires less learning rate decay compared to SGD since averaging naturally reduces noise, introducing a form of implicit regularization. Through extensive experiments, we show that EMA solutions differ from last-iterate solutions. EMA models not only generalize better but also exhibit improved i) robustness to noisy labels, ii) prediction consistency, iii) calibration and iv) transfer learning. Therefore, we suggest that an EMA of weights is a simple yet effective plug-in to improve the performance of deep learning models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 18 citations worldwide. Full citation record

  1. Exploring Line Bundle Standard Models with Transformers

    hep-th 2026-06 unverdicted novelty 7.0 of 10

    A Transformer trained by reinforcement learning generates heterotic line-bundle sums that satisfy anomaly-cancellation, stability, and chirality constraints, and its policy transfers usefully across Calabi-Yau geometries.

  2. Beyond Token-Level Cross-Entropy: Fr\'echet Distributional Post-Training for Autoregressive Image Generation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    FD-loss post-training with detached rollout replay and a probability-level straight-through estimator improves FID and FD_r6 across eight ImageNet configurations.

  3. A Generalisable Generative Model for Multi-Detector Calorimeter Simulation

    physics.ins-det 2025-09 conditional novelty 6.0 of 10

    CaloDiT-2 demonstrates that pre-training a transformer-based diffusion model on multiple calorimeter detectors enables 25x less data and 20x less training time when adapting to a new detector.

  4. DMAF-Net: An Effective Modality Rebalancing Framework for Incomplete Multi-Modal Medical Image Segmentation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    DMAF-Net improves incomplete multi-modal MRI segmentation by dynamically masking missing modalities, aligning uni-modal and fused features through covariance, attention, and prototype losses, and reweighting training ...

  5. FootsiesGym: A Fighting Game Benchmark for Two-Player Zero-Sum Imperfect-Information Games

    cs.AI 2026-07 accept novelty 5.0 of 10

    FootsiesGym is an open-source, vectorized fighting-game benchmark for two-player zero-sum imperfect-information RL that isolates non-transitive neutral-game dynamics while remaining tractable on standard hardware.

  6. Demonstration of Efficient Predictive Surrogates for Large-scale Quantum Processors

    quant-ph 2025-07 conditional novelty 5.0 of 10

    Classical surrogates using truncated trigonometric expansions emulate noisy quantum processors and cut measurement overhead in VQE pre-training and Floquet phase identification.

Pith tools