REVIEW 1 cited by
Local AdaAlter: Communication-Efficient Stochastic Gradient Descent with Adaptive Learning Rates
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
When scaling distributed training, the communication overhead is often the bottleneck. In this paper, we propose a novel SGD variant with reduced communication and adaptive learning rates. We prove the convergence of the proposed algorithm for smooth but non-convex problems. Empirical results show that the proposed algorithm significantly reduces the communication overhead, which, in turn, reduces the training time by up to 30% for the 1B word dataset.
Forward citations
Cited by 1 Pith paper
-
Gradient Correction in Federated Learning with Adaptive Optimization
FAdamGC adds SCAFFOLD-style drift correction into local Adam updates before moment estimation, yielding a communication-efficient federated optimizer for non-IID data.
Discussion (0). Continue with ORCID to comment.