Pith. sign in

REVIEW 1 cited by

Balance is Essence: Accelerating Sparse Training via Adaptive Gradient Correction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.03573 v2 pith:3TBKGNTC submitted 2023-01-09 cs.LG cs.AIcs.CV

classification cs.LGcs.AIcs.CV
keywords trainingsparsegradientmethodaccelerateaccuracyachieveadaptive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite impressive performance, deep neural networks require significant memory and computation costs, prohibiting their application in resource-constrained scenarios. Sparse training is one of the most common techniques to reduce these costs, however, the sparsity constraints add difficulty to the optimization, resulting in an increase in training time and instability. In this work, we aim to overcome this problem and achieve space-time co-efficiency. To accelerate and stabilize the convergence of sparse training, we analyze the gradient changes and develop an adaptive gradient correction method. Specifically, we approximate the correlation between the current and previous gradients, which is used to balance the two gradients to obtain a corrected gradient. Our method can be used with the most popular sparse training pipelines under both standard and adversarial setups. Theoretically, we prove that our method can accelerate the convergence rate of sparse training. Extensive experiments on multiple datasets, model architectures, and sparsities demonstrate that our method outperforms leading sparse training methods by up to \textbf{5.0\%} in accuracy given the same number of training epochs, and reduces the number of training epochs by up to \textbf{52.1\%} to achieve the same accuracy. Our code is available on: \url{https://github.com/StevenBoys/AGENT}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. No More Adam: Learning Rate Scaling at Initialization is All You Need

    cs.LG 2024-12 conditional novelty 5.0 of 10

    A fixed, per-group learning-rate scaling computed at initialization lets SGD with momentum match AdamW on several Transformer tasks while halving optimizer memory.

Pith tools