REVIEW 6 cited by
Communication-Efficient Distributed Deep Learning: A Comprehensive Survey
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Distributed deep learning (DL) has become prevalent in recent years to reduce training time by leveraging multiple computing devices (e.g., GPUs/TPUs) due to larger models and datasets. However, system scalability is limited by communication becoming the performance bottleneck. Addressing this communication issue has become a prominent research topic. In this paper, we provide a comprehensive survey of the communication-efficient distributed training algorithms, focusing on both system-level and algorithmic-level optimizations. We first propose a taxonomy of data-parallel distributed training algorithms that incorporates four primary dimensions: communication synchronization, system architectures, compression techniques, and parallelism of communication and computing tasks. We then investigate state-of-the-art studies that address problems in these four dimensions. We also compare the convergence rates of different algorithms to understand their convergence speed. Additionally, we conduct extensive experiments to empirically compare the convergence performance of various mainstream distributed training algorithms. Based on our system-level communication cost analysis, theoretical and experimental convergence speed comparison, we provide readers with an understanding of which algorithms are more efficient under specific distributed environments. Our research also extrapolates potential directions for further optimizations.
Forward citations
Cited by 6 Pith papers
-
Constant Stepsize Local GD for Logistic Regression: Acceleration by Instability
For separable logistic regression, Local GD with any step size and any communication interval converges at rate O~(1/(eta K R)) after O~(eta K M) unstable rounds, beating the general O(1/R) worst-case bound.
-
Label-shift robust federated feature screening for high-dimensional classification
A new label-shift robust utility, LR-FFS, is proposed for federated feature screening, with a unifying framework, distributed estimation, and FDR control.
-
Lion Cub: Minimizing Communication Overhead in Distributed Lion
Lion Cub compresses Lion updates with L1 quantization and sparse momentum synchronization, reducing distributed training time by up to 5.1x at similar convergence.
-
Protocol Learning, Decentralized Frontier Risk and the No-Off Problem
The paper introduces Protocol Learning, a decentralized, incentivized training paradigm, and argues it could reduce frontier risk even as it creates the No-Off Problem.
-
Theoretical Foundations of Communication-Efficient, Robust, and Practical Distributed and Federated Optimization
A thesis proving communication-acceleration guarantees for local-step, compressed, Byzantine-robust, and low-rank federated optimization methods, assembled from the author's own published papers.
-
Semantic-Aware Visual Information Transmission With Key Information Extraction Over Wireless Networks
A foreground-cropping, background-library image transmission system reports PSNR gains over direct deep JSCC, but the gains are confounded by an unequal transmission workload.
Discussion (0). Continue with ORCID to comment.