Pith. sign in

REVIEW 1 cited by

Understanding AdamW through Proximal Methods and Scale-Freeness

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.00089 v1 pith:KTBF2AKO submitted 2022-01-31 cs.LG math.OC

classification cs.LGmath.OC
keywords adamwadam-proximaladvantagegradientregularizeradamdeep
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Adam has been widely adopted for training deep neural networks due to less hyperparameter tuning and remarkable performance. To improve generalization, Adam is typically used in tandem with a squared $\ell_2$ regularizer (referred to as Adam-$\ell_2$). However, even better performance can be obtained with AdamW, which decouples the gradient of the regularizer from the update rule of Adam-$\ell_2$. Yet, we are still lacking a complete explanation of the advantages of AdamW. In this paper, we tackle this question from both an optimization and an empirical point of view. First, we show how to re-interpret AdamW as an approximation of a proximal gradient method, which takes advantage of the closed-form proximal mapping of the regularizer instead of only utilizing its gradient information as in Adam-$\ell_2$. Next, we consider the property of "scale-freeness" enjoyed by AdamW and by its proximal counterpart: their updates are invariant to component-wise rescaling of the gradients. We provide empirical evidence across a wide range of deep learning experiments showing a correlation between the problems in which AdamW exhibits an advantage over Adam-$\ell_2$ and the degree to which we expect the gradients of the network to exhibit multiple scales, thus motivating the hypothesis that the advantage of AdamW could be due to the scale-free updates.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Universal crystal material property prediction via multi-view geometric fusion in graph transformers

    cs.LG 2025-07 conditional novelty 6.0 of 10

    MGT fuses SE3-invariant and SO3-equivariant crystal graph representations with a task-adaptive mixture-of-experts router, achieving state-of-the-art accuracy on multiple crystal property prediction benchmarks.

Pith tools