REVIEW 2 major objections 4 minor 300 references
Theoretical Foundations of Communication-Efficient, Robust, and Practical Distributed and Federated Optimization
T0 review · 2 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read This thesis proves that local gradient steps can accelerate communication in federated optimization, cutting the required number of communication rounds from $O(\kappa \log 1/\varepsilon)$ to $O(\sqrt{\kappa} \log 1/\varepsilon)$ under…
desk verdict A thesis built on seven already-published papers: the ProxSkip/Scaffnew result is real and clean, but the abstract's unqualified claim that local steps accelerate communication is proven only in the exact-gradient, shared-strong-convexity regime. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the randomized prox-skipping iteration of ProxSkip: with probability $p$ the expensive operator (in federated learning, the averaging of client models) is skipped, and the skipped steps are replaced by gradient steps shifted by a control variate that tracks the drift of each client. The proof centers on a two-term Lyapunov function combining distance to the solution and control-variate error, together with firm nonexpansiveness of the proximity operator (the averaging step is a contraction in a precise sense). Choosing $p=1/\sqrt{\kappa}$ balances the contraction factor $\min\{\gamma\mu, p^2\}$ so that iteration complexity stays $\kappa \log 1/\varepsilon$ while expected prox evaluations drop to $\sqrt{\kappa} \log 1/\varepsilon$. Chapter 4 replaces the prox with a finite number of local gradient steps inside a primal-dual method, which is what lets partial participation enter without losing acceleration.
What would settle it
On a strongly convex quadratic with known condition number $\kappa$ (say $\kappa$ from $10^2$ to $10^6$), run Scaffnew with $\gamma=1/L$ and $p=1/\sqrt{\kappa}$ to a fixed accuracy $\varepsilon$ and count communication rounds; if the round count scales as $\kappa \log(1/\varepsilon)$ rather than $\sqrt{\kappa} \log(1/\varepsilon)$, or if a single instance under the paper's assumptions requires $\Omega(\kappa \log 1/\varepsilon)$ rounds, the central claim is refuted.
Extended reading notes
Core claim
The paper's central claim is that local gradient steps alone can accelerate communication in heterogeneous federated optimization, resolving an open problem that had resisted several generations of local-training theory. The method ProxSkip, instantiated as Scaffnew for federated learning, uses a control variate that converges to the gradient at the optimum and a Bernoulli coin with probability $p=1/\sqrt{\kappa}$ that decides whether to skip the averaging round. Under only $L$-smoothness and $\mu$-strong convexity of each local function, the expected number of communications is $O(\sqrt{\kappa} \log 1/\varepsilon)$, breaking the $O(\kappa \log 1/\varepsilon)$ complexity of earlier drift-corrected local methods and matching the first-order distributed-optimization lower bound. The thesis further claims that a primal-dual reformulation, 5GCS, preserves this acceleration under partial client participation, that variance-reduced ProxSkip restores linear convergence with stochastic gradients, and that the same design principles yield the first provable frameworks combining random reshuffling with compression, Byzantine robustness with partial participation, and low-rank adaptation.
Load-bearing premise
The load-bearing premise is that every local objective is $L$-smooth and $\mu$-strongly convex with known shared constants, and that the core acceleration is shown for exact gradients; if $\mu=0$, or if the condition number $\kappa$ is not available to set $p$, the linear-rate argument collapses, and the stochastic version without variance reduction only reaches a neighborhood of the solution.
Editorial extensions
If this is right
- If the thesis is right, local training methods can be optimal in communication without Nesterov-style momentum, so the practical Federated-Averaging-style pipeline needs no extra acceleration machinery.
- No bounded-dissimilarity or data-homogeneity assumption is needed for the speedup; heterogeneous client data do not break the rate.
- Communication acceleration survives partial participation, so methods remain optimal when only a small cohort of devices can be reached each round.
- Variance reduction makes stochastic local updates converge linearly and can lower total cost once local computation is counted, not just communication.
- Gradient-difference compression and clipping extend the same acceleration to compressed communication and Byzantine-robust settings, and RAC-LoRA gives low-rank fine-tuning a convergence guarantee.
Reading between the lines
- The thesis does not spell this out, but the same randomized-skip control variate could serve as a lazy-aggregation rule in asynchronous or time-varying networks, where the server skips aggregation whenever the control-variate error is small rather than by a fixed Bernoulli schedule.
- Because the contraction factor is $\min\{\gamma\mu, p^2\}$, the method suggests a testable practice: when $\kappa$ is unknown, estimating $\mu$ from local Hessian information and adapting $p$ online may preserve most of the $\sqrt{\kappa}$ gain.
- RAC-LoRA's construction implies that standard LoRA's instability may come from the persistent asymmetry between the two low-rank factors; one could test this by monitoring the gradient norm of the frozen-initialized factor and checking whether divergence coincides with its growth.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This dissertation compiles seven previously peer-reviewed works (Chapters 2–8) on communication-efficient distributed and federated optimization. The central claim, developed in Chapter 2, is that local gradient steps can provably accelerate communication: the ProxSkip/Scaffnew method achieves O(√κ log(1/ε)) communication complexity for smooth strongly convex finite-sum problems under heterogeneous data, breaking the O(κ log(1/ε)) barrier of GD and matching the Arjevani–Shamir lower bound. The remaining chapters extend this framework: ProxSkip-VR adds variance reduction and a hierarchical hub architecture; 5GCS supports partial participation with optimal communication complexity; NASTYA analyzes server-side stepsizes and random reshuffling; DIANA-RR and Q-RR combine compression with reshuffling; Byz-VR-MARINA-PP combines Byzantine robustness with partial participation; RAC-LoRA provides a convergence theory for low-rank adaptation. Experimental sections validate the theory on logistic regression and small deep-learning benchmarks.
Significance. If the claims hold, the thesis provides a substantial unification of the theory of local training methods. The core derivation in Chapter 2 is a genuine strength: Lemmas 2 and 3 combine into Theorem 1 with rate E[Ψ(T)] ≤ (1 − min{γμ, p²})^T Ψ(0), and the √κ communication complexity follows from p = 1/√κ with no fitted constants. The extensions to partial participation, Byzantine robustness, and LoRA are concrete algorithmic novelties, and each chapter has survived peer review in ICML, NeurIPS, and AISTATS. The five-generations taxonomy in Chapter 3 is useful for situating the field. However, the significance of the flagship result is restricted to the exact-gradient, strongly convex regime with known κ; the thesis itself notes that the stochastic extension does not achieve linear speedup.
major comments (2)
- [Abstract, §1.5, §2.2.2] The abstract and Section 2.2.2 state, without qualification, that local gradient steps 'can accelerate communication' and that Scaffnew breaks the O(κ log 1/ε) barrier 'without imposing any additional assumptions.' The supporting theorem, Theorem 1 with rate (2.15), requires every local function f_m to be L-smooth and μ-strongly convex with a shared μ > 0 (Assumptions 3 and 12), exact local gradients, and a user-supplied condition number κ = L/μ used to set p = 1/√κ. When μ = 0 the rate collapses to ζ = min{γμ, p²} = p², yet κ is infinite and the claimed O(√κ log 1/ε) complexity is vacuous; in the non-strongly-convex regime no communication acceleration is proved. I ask the authors to rephrase the central claim so that it explicitly refers to the exact-gradient, smooth strongly convex setting with known κ, and to state in the abstract that the stochastic extension does not inherit the same acceleration guarantee.
- [§2.5.1, Theorem 2, Corollary 3] The stochastic extension SProxSkip has an additive noise floor: Theorem 2 gives E[Ψ(T)] ≤ (1 − ζ)^T Ψ(0) + γ²C/ζ, and Corollary 3 yields communication rounds of order √(2C/(εμ²)) log(1/ε) rather than √κ log(1/ε). The limitations paragraph in Section 2.5.1 correctly concedes that the analysis does not achieve linear speedup. This is more than a cosmetic caveat: the introduction motivates local training through the stochastic and deep-learning regime, and the abstract's 'theoretical foundation for this widely used heuristic' promises more than Theorem 2 delivers. I recommend moving this limitation into the summary of contributions and explicitly separating the exact-gradient acceleration result from the stochastic result.
minor comments (4)
- [§2.5.1, Corollary 3] Corollary 3 says 'in order to guarantee E[Ψ(0)] ≤ ε,' which appears to be a typo; the intended statement is presumably about E[Ψ(T)] after T iterations, or about the normalized quantity E[Ψ(T)]/Ψ(0).
- [§2.6, Figure 2.2] The caption of Figure 2.2 states that 'we can obtain the linear speedup, which is more optimistic than we have in theory,' but Section 2.5.1 explicitly says that the SProxSkip analysis does not achieve linear speedup. Please label this as an empirical observation and, ideally, add a sentence reconciling it with the theoretical limitation.
- [§2.2.2 and §2.4.1] The statement that Scaffnew is optimal because it matches the Arjevani–Shamir lower bound would benefit from a precise description of the oracle and architecture class to which that lower bound applies; as written, the reader cannot verify that the lower bound covers the randomized, geometric communication schedule used by Scaffnew.
- [§8.2.2 and Table 8.1] The phrase 'first theoretical framework' for LoRA-type methods is strong; it should be accompanied by a brief comparison with any existing convergence analyses of low-rank adaptation methods so that the novelty claim can be adjudicated by the reader.
Circularity Check
No circularity found; the acceleration theorems are derived from explicit smoothness/strong-convexity assumptions, with optimality claims resting on external lower bounds and experiments not feeding back into the proofs.
full rationale
The central derivation chain is self-contained. Theorem 1 in Chapter 2 proves a contraction bound for ProxSkip from Assumption 1 (L-smoothness and mu-strong convexity) and Assumption 2 (convex regularizer), via Lemmas 2 and 3; the optimal choices gamma = 1/L and p = 1/sqrt(kappa) are obtained by minimizing the explicit expression (2.17) for the expected number of prox evaluations, not by fitting constants to data. The claimed O(sqrt(kappa) log 1/eps) communication complexity follows algebraically from (2.15)-(2.19) and Corollary 1. The optimality statements cite external lower bounds (Arjevani and Shamir 2015, Woodworth and Srebro 2016, Scaman et al. 2019), not a self-citation. Chapter 3's ProxSkip-VR results are proven from Assumptions 7-10 with estimator-dependent constants, and recover ProxSkip/SProxSkip as special cases by setting the estimator parameters, which is a consistency check rather than a circular reduction. Chapter 4's 5GCS results are likewise proved from Assumption 12 (L-smooth and mu-strongly convex clients) and Assumption 13, with communication complexities following from Corollaries 6-8. The thesis is compiled from the author's own prior papers, so self-citations are pervasive, but the load-bearing theorems are reproduced with proofs in the text and appendices; no premise is justified solely by a self-citation. The skepticism expressed in the reader's take concerns scope: the acceleration theorem requires exact gradients and shared mu > 0, and the stochastic extension is acknowledged not to achieve linear speedup. That is a limitation or correctness concern about how far the abstract's unqualified claim extends, not circularity, because the theorems do not assume the conclusion they claim to prove.
Assumptions & free parameters
free parameters (4)
- p, the communication/prox probability in ProxSkip and Scaffnew =
Theory: 1/sqrt(kappa). Experiments: 1/p tuned from 100 to 1000, empirical optimum near 300 in Figure 2.1(c).
- Server and client stepsizes (eta, gamma) of NASTYA =
Tuned in experiments; value trace in Figure 5.2 (middle).
- Clipping multiplier lambda of Byz-VR-MARINA-PP =
Varied in experiments; sensitivity shown in Figure 7.1 (right).
- LoRA rank r and epochs per chain block in RAC-LoRA =
Rank 2 for RoBERTa experiments; epochs per block varied in Table G.2.
assumptions (6)
- domain assumption Each local function f_m is L-smooth and mu-strongly convex with common constants (Assumptions 1, 3, 12).
- standard math Firm nonexpansiveness of the proximity operator (Lemma 1).
- domain assumption Partial participation is modeled as uniformly random cohort sampling, with non-sampled clients freezing their dual variables (Algorithm 7, step 4).
- domain assumption Stochastic gradient estimators satisfy expected smoothness or the parametric variance bounds of Assumptions 5 and 10.
- domain assumption Assumption 19 bounds the honest gradient noise relative to the clipping scale in the Byzantine chapter.
- domain assumption For RAC-LoRA, the fine-tuning objective is smooth (and satisfies the Polyak-Lojasiewicz condition in the PL results), and the randomized chain matrices are drawn from distributions whose supports span all directions.
invented entities (2)
-
Regional hub network tier (ProxSkip-HUB)
-
Randomized asymmetric chain of low-rank blocks (RAC-LoRA)
Cite this review
Pith. "Pith review of Theoretical Foundations of Communication-Efficient, Robust, and Practical Distributed and Federated Optimization." pith.science (2026). https://pith.science/paper/PAKCIVEK
@misc{pith2026260806563,
author = {Pith},
title = {Pith review of: Theoretical Foundations of Communication-Efficient, Robust, and Practical Distributed and Federated Optimization},
year = {2026},
howpublished = {\url{https://pith.science/paper/PAKCIVEK}},
note = {Machine review of arXiv:2608.06563}
}
read the original abstract
Machine learning and optimization have advanced together, with practical demands motivating new theory and theoretical breakthroughs enabling new applications. Modern large-scale training relies on classical optimization principles, but the constraints of distributed systems require these foundations to be reconsidered. This thesis addresses seven challenges at the intersection of theory and practice, focusing on key bottlenecks in federated learning and distributed optimization. First, we introduce ProxSkip and prove that local gradient steps can accelerate communication, providing a theoretical foundation for this widely used heuristic. Second, we develop Variance Reduced ProxSkip, which eliminates the neighborhood error of stochastic local updates while balancing communication and local computation. Third, we show that local steps retain their communication acceleration under partial client participation. Fourth, we prove that server-side stepsizes and sampling without replacement improve convergence in heterogeneous settings. Fifth, for Random Reshuffling, we demonstrate that compressing gradient differences rather than gradients yields better theoretical and practical performance. Sixth, we establish that Byzantine robustness and partial participation can be achieved simultaneously using gradient-difference clipping. Finally, we develop the first theoretical framework for low-rank adaptation based on randomized asymmetric chains, providing new insights into fine-tuning large models. Across these contributions, we introduce novel algorithmic frameworks, establish sharp guarantees under realistic assumptions, and support the theory with numerical experiments.
Figures
Figures from the paper (14 more)
Reference graph
Works this paper leans on
-
[1]
Deep learning with differential privacy
Martin Abadi, Andy Chu, Ian Goodfellow, H Brendan McMahan, Ilya Mironov, Kunal Talwar, and Li Zhang. Deep learning with differential privacy. In2016 ACM SIGSAC Conference on Computer and Communi- cations Security, pages 308–318, 2016.(Cited on pages 24, 28, and 350)
2016
-
[2]
Federated learning in edge computing: a systematic survey.Sensors, 22(2):450, 2022.(Cited on page 24)
Haftay Gebreslasie Abreha, Mohammad Hayajneh, and Mohamed Adel Ser- hani. Federated learning in edge computing: a systematic survey.Sensors, 22(2):450, 2022.(Cited on page 24)
2022
-
[3]
Distributed delayed stochastic optimiza- tion.Advances in neural information processing systems, 24, 2011.(Cited on page 350)
Alekh Agarwal and John C Duchi. Distributed delayed stochastic optimiza- tion.Advances in neural information processing systems, 24, 2011.(Cited on page 350)
2011
-
[4]
Armen Aghajanyan, Luke Zettlemoyer, and Sonal Gupta. Intrinsic dimen- sionality explains the effectiveness of language model fine-tuning.arXiv preprint arXiv:2012.13255, 2020.(Cited on page 131)
arXiv 2012
-
[5]
Kwangjun Ahn, Chulhee Yun, and Suvrit Sra. SGD with shuffling: Optimal rates without component convexity and large epoch requirements.Advances in Neural Information Processing Systems, 33:17526–17535, 2020.(Cited on pages 106 and 423)
2020
-
[6]
Ahmad Ajalloeian and Sebastian U Stich. On the convergence of SGD with biased gradients.arXiv preprint arXiv:2008.00051, 2020.(Cited on page 350)
arXiv 2008
-
[7]
QSGD: Communication-efficient SGD via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vo- jnovic. QSGD: Communication-efficient SGD via gradient quantization and encoding. InAdvances in Neural Information Processing Systems, pages 1709–1720, 2017.(Cited on pages 26, 30, 68, 105, 114, 121, 271, 276, and 349)
2017
-
[8]
Byzantine stochastic gradi- ent descent
Dan Alistarh, Zeyuan Allen-Zhu, and Jerry Li. Byzantine stochastic gradi- ent descent. InAdvances in Neural Information Processing Systems, pages 4618–4628, 2018a.(Cited on pages 28, 117, and 348)
Show all 300 references
-
[9]
The convergence of sparsified gradient methods
Dan Alistarh, Torsten Hoefler, Mikael Johansson, Sarit Khirirat, Nikola Konstantinov, and C´ edric Renggli. The convergence of sparsified gradient methods. InAdvances in Neural Information Processing Systems, 2018b. (Cited on page 91) 148
-
[10]
Byzantine-resilient non-convex stochastic gradient descent
Zeyuan Allen-Zhu, Faeze Ebrahimian, Jerry Li, and Dan Alistarh. Byzantine-resilient non-convex stochastic gradient descent. InInternational Conference on Learning Representations, 2021.(Cited on pages 117 and 348)
2021
-
[11]
Fixing by mixing: A recipe for optimal byzantine ml under heterogeneity
Youssef Allouah, Sadegh Farhadkhani, Rachid Guerraoui, Nirupam Gupta, Rafa¨ el Pinot, and John Stephan. Fixing by mixing: A recipe for optimal byzantine ml under heterogeneity. InInternational Conference on Artificial Intelligence and Statistics, pages 1232–1300, 2023.(Cited o...
2023
-
[12]
Tackling byzantine clients in federated learning.arXiv preprint arXiv:2402.12780, 2024a.(Cited on page 119)
Youssef Allouah, Sadegh Farhadkhani, Rachid Guerraoui, Nirupam Gupta, Rafael Pinot, Geovani Rizk, and Sasha Voitovych. Tackling byzantine clients in federated learning.arXiv preprint arXiv:2402.12780, 2024a.(Cited on page 119)
-
[13]
Youssef Allouah, Rachid Guerraoui, Nirupam Gupta, Rafa¨ el Pinot, and Geovani Rizk. Robust distributed learning: tight error bounds and break- down point under data heterogeneity.Advances in Neural Information Pro- cessing Systems, 36, 2024b.(Cited on pages 28, 127, and 405)
-
[14]
Federated learning for healthcare: Systematic review and architecture proposal.ACM Transactions on Intel- ligent Systems and Technology (TIST), 13(4):1–23, 2022.(Cited on page 104)
Rodolfo Stoffel Antunes, Cristiano Andr´ e da Costa, Arne K¨ uderle, Im- rana Abdullahi Yari, and Bj¨ orn Eskofier. Federated learning for healthcare: Systematic review and architecture proposal.ACM Transactions on Intel- ligent Systems and Technology (TIST), 13(4):1–23, 2022....
2022
-
[15]
Communication complexity of distributed convex learning and optimization
Yossi Arjevani and Ohad Shamir. Communication complexity of distributed convex learning and optimization. InAdvances in Neural Information Pro- cessing Systems, 2015.(Cited on pages 48 and 53)
2015
-
[16]
Lower bounds for non-convex stochastic optimiza- tion.Mathematical Programming, 199(1-2):165–214, 2023.(Cited on pages 123 and 125)
Yossi Arjevani, Yair Carmon, John C Duchi, Dylan J Foster, Nathan Srebro, and Blake Woodworth. Lower bounds for non-convex stochastic optimiza- tion.Mathematical Programming, 199(1-2):165–214, 2023.(Cited on pages 123 and 125)
2023
-
[17]
Distributed deep learning using volunteer computing-like paradigm.arXiv preprint arXiv:2103.08894, 2021.(Cited on page 117)
Medha Atre, Birendra Jha, and Ashwini Rao. Distributed deep learning using volunteer computing-like paradigm.arXiv preprint arXiv:2103.08894, 2021.(Cited on page 117)
2021 arXiv
-
[18]
A little is enough: Cir- cumventing defenses for distributed learning
Gilad Baruch, Moran Baruch, and Yoav Goldberg. A little is enough: Cir- cumventing defenses for distributed learning. InAdvances in Neural Infor- mation Processing Systems, volume 32, 2019.(Cited on pages 117, 118, 130, and 417) 149
2019
-
[19]
Qsparse- local-SGD: Distributed SGD with quantization, sparsification and local computations
Debraj Basu, Deepesh Data, Can Karakus, and Suhas Diggavi. Qsparse- local-SGD: Distributed SGD with quantization, sparsification and local computations. InAdvances in Neural Information Processing Systems, vol- ume 32, 2019.(Cited on pages 268 and 349)
2019
-
[20]
Generalized mono- tone operators and their averaged resolvents.Mathematical Programming, 189(1):55–74, 2021.(Cited on page 49)
Heinz H Bauschke, Walaa M Moursi, and Xianfu Wang. Generalized mono- tone operators and their averaged resolvents.Mathematical Programming, 189(1):55–74, 2021.(Cited on page 49)
2021
-
[21]
MOS-SIAM Series on Optimization, 2017.(Cited on pages 43 and 63)
Amir Beck.First order methods in optimization. MOS-SIAM Series on Optimization, 2017.(Cited on pages 43 and 63)
2017
-
[22]
Demystifying parallel and distributed deep learning: An in-depth concurrency analysis.ACM Computing Surveys (CSUR), 52(4):1–43, 2019.(Cited on page 23)
Tal Ben-Nun and Torsten Hoefler. Demystifying parallel and distributed deep learning: An in-depth concurrency analysis.ACM Computing Surveys (CSUR), 52(4):1–43, 2019.(Cited on page 23)
2019
-
[23]
Practical recommendations for gradient-based training of deep architectures
Yoshua Bengio. Practical recommendations for gradient-based training of deep architectures. InNeural Networks: Tricks of the trade, pages 437–478. Springer, 2012.(Cited on page 91)
2012
-
[24]
signsgd with majority vote is communication efficient and byzantine fault tolerant.arXiv preprint arXiv:1810.05291, 2018.(Cited on page 350)
Jeremy Bernstein, Jiawei Zhao, Kamyar Azizzadenesheli, and Anima Anandkumar. signsgd with majority vote is communication efficient and byzantine fault tolerant.arXiv preprint arXiv:1810.05291, 2018.(Cited on page 350)
2018 arXiv
-
[25]
Incremental proximal methods for large scale convex optimization.Mathematical Programming, 129(2):163–195, 2011.(Cited on page 423)
Dimitri P Bertsekas. Incremental proximal methods for large scale convex optimization.Mathematical Programming, 129(2):163–195, 2011.(Cited on page 423)
2011
-
[26]
On biased compression for distributed learning.Journal of Machine Learning Research, 24(276):1–50, 2023.(Cited on pages 114, 121, 269, and 350)
Aleksandr Beznosikov, Samuel Horv´ ath, Peter Richt´ arik, and Mher Sa- faryan. On biased compression for distributed learning.Journal of Machine Learning Research, 24(276):1–50, 2023.(Cited on pages 114, 121, 269, and 350)
2023
-
[27]
LoRA learns less and forgets less.arXiv preprint arXiv:2405.09673, 2024.(Cited on page 132)
Dan Biderman, Jose Gonzalez Ortiz, Jacob Portes, Mansheej Paul, Philip Greengard, Connor Jennings, Daniel King, Sam Havens, Vitaliy Chiley, Jonathan Frankle, et al. LoRA learns less and forgets less.arXiv preprint arXiv:2405.09673, 2024.(Cited on page 132)
2024 arXiv
-
[28]
Machine learning with adversaries: Byzantine tolerant gradient descent
Peva Blanchard, El Mahdi El Mhamdi, Rachid Guerraoui, and Julien Stainer. Machine learning with adversaries: Byzantine tolerant gradient descent. InAdvances in Neural Information Processing Systems, volume 30, 2017.(Cited on pages 28, 118, and 120) 150
2017
-
[29]
Brendan McMahan, Sarvar Patel, Daniel Ramage, Aaron Segal, and Karn Seth
Keith Bonawitz, Vladimir Ivanov, Ben Kreuter, Antonio Marcedone, H. Brendan McMahan, Sarvar Patel, Daniel Ramage, Aaron Segal, and Karn Seth. Practical secure aggregation for privacy-preserving machine learning. In2017 ACM SIGSAC Conference on Computer and Communi- cations Sec...
2017
-
[30]
Towards federated learning at scale: System design.Proceedings of machine learning and systems, 1:374–388, 2019.(Cited on page 24)
Keith Bonawitz, Hubert Eichner, Wolfgang Grieskamp, Dzmitry Huba, Alex Ingerman, Vladimir Ivanov, Chloe Kiddon, Jakub Koneˇ cn` y, Stefano Mazzocchi, Brendan McMahan, et al. Towards federated learning at scale: System design.Proceedings of machine learning and systems, 1:374–3...
2019
-
[31]
Curiously fast convergence of some stochastic gradient descent algorithms
L´ eon Bottou. Curiously fast convergence of some stochastic gradient descent algorithms. InSymposium on learning and data science, Paris, volume 8, pages 2624–2633, 2009.(Cited on pages 30, 91, and 105)
2009
-
[32]
Democratizing machine learning: Resilient distributed learning with heterogeneous participants
Karim Boubouh, Amine Boussetta, Nirupam Gupta, Alexandre Maurer, and Rafa¨ el Pinot. Democratizing machine learning: Resilient distributed learning with heterogeneous participants. In2022 41st International Sym- posium on Reliable Distributed Systems (SRDS), pages 94–120. IEEE...
2022
-
[33]
Minibatch stochastic three points method for unconstrained smooth minimization
Soumia Boucherouite, Grigory Malinovsky, Peter Richt´ arik, and El Houcine Bergou. Minibatch stochastic three points method for unconstrained smooth minimization. InAAAI Conference on Artificial Intelligence, vol- ume 38, pages 20344–20352, 2024.(Cited on page 39)
2024
-
[34]
Flpytorch: optimization research simulator for federated learning
Konstantin Burlachenko, Samuel Horv´ ath, and Peter Richt´ arik. Flpytorch: optimization research simulator for federated learning. InACM Interna- tional Workshop on Distributed Machine Learning, pages 1–7, 2021.(Cited on pages 115, 274, and 275)
2021
-
[35]
Tighter lower bounds for shuffling SGD: Random permutations and beyond
Jaeyoung Cha, Jaewook Lee, and Chulhee Yun. Tighter lower bounds for shuffling SGD: Random permutations and beyond. InInternational Con- ference on Machine Learning, pages 3855–3912. PMLR, 2023.(Cited on page 423)
2023
-
[36]
A first-order primal-dual algorithm for convex problems with applications to imaging.Journal of Mathematical Imaging and Vision, 40(1):120–145, 2011.(Cited on page 78)
Antonin Chambolle and Thomas Pock. A first-order primal-dual algorithm for convex problems with applications to imaging.Journal of Mathematical Imaging and Vision, 40(1):120–145, 2011.(Cited on page 78)
2011
-
[37]
LIBSVM: A library for support 151 vector machines.ACM transactions on intelligent systems and technology (TIST), 2(3):1–27, 2011.(Cited on pages 57, 71, 87, 102, 129, and 269)
Chih-Chung Chang and Chih-Jen Lin. LIBSVM: A library for support 151 vector machines.ACM transactions on intelligent systems and technology (TIST), 2(3):1–27, 2011.(Cited on pages 57, 71, 87, 102, 129, and 269)
2011
-
[38]
On the outsized importance of learn- ing rates in local update methods.arXiv preprint arXiv:2007.00878, 2020
Zachary Charles and Jakub Koneˇ cn´ y. On the outsized importance of learn- ing rates in local update methods.arXiv preprint arXiv:2007.00878, 2020. (Cited on page 436)
2007 arXiv
-
[39]
On large-cohort training for federated learning.arXiv preprint arXiv:2106.07820, 2021.(Cited on pages 60 and 90)
Zachary Charles, Zachary Garrett, Zhouyuan Huo, Sergei Shmulyian, and Virginia Smith. On large-cohort training for federated learning.arXiv preprint arXiv:2106.07820, 2021.(Cited on pages 60 and 90)
2021 arXiv
-
[40]
Draco: Byzantine-resilient distributed training via redundant gradi- ents
Lingjiao Chen, Hongyi Wang, Zachary Charles, and Dimitris Papailiopou- los. Draco: Byzantine-resilient distributed training via redundant gradi- ents. InInternational Conference on Machine Learning, pages 903–912, 2018.(Cited on page 349)
2018
-
[41]
A primal–dual fixed point algorithm for convex separable minimization with applications to im- age restoration.Inverse Problems, 29(2):025011, 2013.(Cited on page 186)
Peijun Chen, Jianguo Huang, and Xiaoqun Zhang. A primal–dual fixed point algorithm for convex separable minimization with applications to im- age restoration.Inverse Problems, 29(2):025011, 2013.(Cited on page 186)
2013
-
[42]
Optimal client sampling for federated learning.Privacy Preserving Machine Learning (NeurIPS 2020 Workshop), 2020a.(Cited on pages 27, 30, 60, 75, and 90)
Wenlin Chen, Samuel Horv´ ath, and Peter Richt´ arik. Optimal client sampling for federated learning.Privacy Preserving Machine Learning (NeurIPS 2020 Workshop), 2020a.(Cited on pages 27, 30, 60, 75, and 90)
2020
-
[43]
Understanding gradient clipping in private SGD: A geometric perspective
Xiangyi Chen, Steven Z Wu, and Mingyi Hong. Understanding gradient clipping in private SGD: A geometric perspective. InAdvances in Neu- ral Information Processing Systems, volume 33, pages 13773–13782, 2020b. (Cited on page 350)
-
[44]
Yudong Chen, Lili Su, and Jiaming Xu. Distributed statistical machine learning in adversarial settings: Byzantine gradient descent.Proceedings of the ACM on Measurement and Analysis of Computing Systems, 1(2):1–25, 2017.(Cited on pages 130 and 417)
2017
-
[45]
Client selection in feder- ated learning: Convergence analysis and power-of-choice selection strate- gies.arXiv preprint arXiv:2010.01243, 2020.(Cited on pages 30 and 60)
Yae Jee Cho, Jianyu Wang, and Gauri Joshi. Client selection in feder- ated learning: Convergence analysis and power-of-choice selection strate- gies.arXiv preprint arXiv:2010.01243, 2020.(Cited on pages 30 and 60)
2010 arXiv
-
[46]
On the convergence of federated averaging with cyclic client participation
Yae Jee Cho, Pranay Sharma, Gauri Joshi, Zheng Xu, Satyen Kale, and Tong Zhang. On the convergence of federated averaging with cyclic client participation. InInternational Conference on Machine Learning, pages 5677–5721. PMLR, 2023.(Cited on page 423) 152
2023
-
[47]
Emerging trends: A gentle introduction to fine-tuning.Natural Language Engineering, 27(6): 763–778, 2021.(Cited on page 131)
Kenneth Ward Church, Zeyu Chen, and Yanjun Ma. Emerging trends: A gentle introduction to fine-tuning.Natural Language Engineering, 27(6): 763–778, 2021.(Cited on page 131)
2021
-
[48]
Proximal splitting methods in signal processing
Patrick L Combettes and Jean-Christophe Pesquet. Proximal splitting methods in signal processing. InFixed-point algorithms for inverse prob- lems in science and engineering, pages 185–212. Springer, 2011.(Cited on page 43)
2011
-
[49]
Combettes, Laurent Condat, Jean-Christophe Pesquet, and B˘ ang C
Patrick L. Combettes, Laurent Condat, Jean-Christophe Pesquet, and B˘ ang C. V˜ u. A forward-backward view of some primal-dual optimization methods in image recovery. In2014 IEEE International Conference on Image Processing (ICIP), pages 4141–4145. IEEE, 2014.(Cited on page 186)
2014
-
[50]
Murana: A generic framework for stochastic variance-reduced optimization.arXiv preprint arXiv:2106.03056, 2021.(Cited on pages 214, 219, 225, and 230)
Laurent Condat and Peter Richt´ arik. Murana: A generic framework for stochastic variance-reduced optimization.arXiv preprint arXiv:2106.03056, 2021.(Cited on pages 214, 219, 225, and 230)
2021 arXiv
-
[51]
RandProx: Primal-dual opti- mization algorithms with randomized proximal updates.arXiv preprint arXiv:2207.12891, 2022.(Cited on pages 77, 78, 80, and 437)
Laurent Condat and Peter Richt´ arik. RandProx: Primal-dual opti- mization algorithms with randomized proximal updates.arXiv preprint arXiv:2207.12891, 2022.(Cited on pages 77, 78, 80, and 437)
2022 arXiv
-
[52]
Proximal splitting algorithms for convex optimization: A tour of recent advances, with new twists.arXiv preprint arXiv:1912.00137, 2019
Laurent Condat, Daichi Kitahara, Andr´ es Contreras, and Akira Hirabayashi. Proximal splitting algorithms for convex optimization: A tour of recent advances, with new twists.arXiv preprint arXiv:1912.00137, 2019. (Cited on page 186)
1912 arXiv
-
[53]
Distributed proximal splitting algorithms with rates and acceleration.Frontiers in Sig- nal Processing, page 12, 2022.(Cited on page 186)
Laurent Condat, Grigory Malinovsky, and Peter Richt´ arik. Distributed proximal splitting algorithms with rates and acceleration.Frontiers in Sig- nal Processing, page 12, 2022.(Cited on page 186)
2022
-
[54]
TAMUNA: Doubly accelerated federated learning with local training, com- pression, and partial participation.arXiv preprint arXiv:2302.09832, 2023
Laurent Condat, Ivan Agarsk` y, Grigory Malinovsky, and Peter Richt´ arik. TAMUNA: Doubly accelerated federated learning with local training, com- pression, and partial participation.arXiv preprint arXiv:2302.09832, 2023. (Cited on pages 39 and 437)
2023 arXiv
-
[55]
Importance sampling for minibatches
Dominik Csiba and Peter Richt´ arik. Importance sampling for minibatches. Journal of Machine Learning Research, 19(27):1–21, 2018.(Cited on page 59)
2018
-
[56]
Momentum-based variance re- duction in non-convex SGD
Ashok Cutkosky and Francesco Orabona. Momentum-based variance re- duction in non-convex SGD. InAdvances in Neural Information Processing Systems, volume 32, 2019.(Cited on pages 128, 349, and 406) 153
2019
-
[57]
Asynchronous byzantine machine learning (the case of sgd)
Georgios Damaskinos, Rachid Guerraoui, Rhicheek Patra, Mahsa Taziki, et al. Asynchronous byzantine machine learning (the case of sgd). InInter- national Conference on Machine Learning, pages 1145–1154. PMLR, 2018. (Cited on pages 27 and 350)
2018
-
[58]
Aggregathor: Byzantine machine learn- ing via robust gradient aggregation.Proceedings of Machine Learning and Systems, 1:81–106, 2019.(Cited on page 118)
Georgios Damaskinos, El-Mahdi El-Mhamdi, Rachid Guerraoui, Arsany Guirguis, and S´ ebastien Rouault. Aggregathor: Byzantine machine learn- ing via robust gradient aggregation.Proceedings of Machine Learning and Systems, 1:81–106, 2019.(Cited on page 118)
2019
-
[59]
Re- cent theoretical advances in non-convex optimization
Marina Danilova, Pavel Dvurechensky, Alexander Gasnikov, Eduard Gor- bunov, Sergey Guminov, Dmitry Kamzolov, and Innokentiy Shibaev. Re- cent theoretical advances in non-convex optimization. InHigh-Dimensional Optimization and Probability: With a View Towards Data Science, pag...
2022
-
[60]
Byzantine-resilient high-dimensional SGD with local iterations on heterogeneous data
Deepesh Data and Suhas Diggavi. Byzantine-resilient high-dimensional SGD with local iterations on heterogeneous data. InInternational Con- ference on Machine Learning, pages 2478–2488. PMLR, 2021.(Cited on pages 119, 124, 348, and 349)
2021
-
[61]
A three-operator splitting scheme and its optimization applications.Set-Val
Demek Davis and Wotao Yin. A three-operator splitting scheme and its optimization applications.Set-Val. Var. Anal., 25:829–858, 2017.(Cited on page 78)
2017
-
[62]
Large scale distributed deep networks.Advances in neural information processing systems, 25, 2012.(Cited on page 23)
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Marc’aurelio Ranzato, Andrew Senior, Paul Tucker, Ke Yang, et al. Large scale distributed deep networks.Advances in neural information processing systems, 25, 2012.(Cited on page 23)
2012
-
[63]
A simple practical accelerated method for finite sums.29th Conference on Neural Information Processing Systems (NeurIPS), 2016
Aaron Defazio. A simple practical accelerated method for finite sums.29th Conference on Neural Information Processing Systems (NeurIPS), 2016. (Cited on pages 34 and 79)
2016
-
[64]
On the ineffectiveness of variance reduced optimization for deep learning.Advances in Neural Information Processing Systems, 32, 2019.(Cited on page 129)
Aaron Defazio and L´ eon Bottou. On the ineffectiveness of variance reduced optimization for deep learning.Advances in Neural Information Processing Systems, 32, 2019.(Cited on page 129)
2019
-
[65]
SAGA: A fast incremental gradient method with support for non-strongly convex com- posite objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien. SAGA: A fast incremental gradient method with support for non-strongly convex com- posite objectives. InAdvances in Neural Information Processing Systems, volume 27, 2014.(Cited on pages 30, 63, 77, 119, and 349) 154
2014
-
[66]
Optimal dis- tributed online prediction using mini-batches.Journal of Machine Learning Research, 13(1):165–202, January 2012.(Cited on pages 30 and 53)
Ofer Dekel, Ran Gilad-Bachrach, Ohad Shamir, and Lin Xiao. Optimal dis- tributed online prediction using mini-batches.Journal of Machine Learning Research, 13(1):165–202, January 2012.(Cited on pages 30 and 53)
2012
-
[67]
Mast: Model-agnostic sparsified training.arXiv preprint arXiv:2311.16086, 2023a.(Cited on page 39)
Yury Demidovich, Grigory Malinovsky, Egor Shulgin, and Peter Richt´ arik. Mast: Model-agnostic sparsified training.arXiv preprint arXiv:2311.16086, 2023a.(Cited on page 39)
-
[68]
A guide through the Zoo of biased SGD
Yury Demidovich, Grigory Malinovsky, Igor Sokolov, and Peter Richt´ arik. A guide through the Zoo of biased SGD. InAdvances in Neural Information Processing Systems, volume 36, 2023b.(Cited on pages 39, 136, 350, and 430)
-
[69]
Streamlin- ing in the riemannian realm: Efficient riemannian optimization with loop- less variance reduction.arXiv preprint arXiv:2403.06677, 2024a.(Cited on page 39)
Yury Demidovich, Grigory Malinovsky, and Peter Richt´ arik. Streamlin- ing in the riemannian realm: Efficient riemannian optimization with loop- less variance reduction.arXiv preprint arXiv:2403.06677, 2024a.(Cited on page 39)
-
[70]
Methods with local steps and random reshuffling for generally smooth non-convex federated optimization.arXiv preprint arXiv:2412.02781, 2024b.(Cited on page 39)
Yury Demidovich, Petr Ostroukhov, Grigory Malinovsky, Samuel Horv´ ath, Martin Tak´ aˇ c, Peter Richt´ arik, and Eduard Gorbunov. Methods with local steps and random reshuffling for generally smooth non-convex federated optimization.arXiv preprint arXiv:2412.02781, 2024b.(Cite...
-
[71]
Distributed deep learning in open collaborations
Michael Diskin, Alexey Bukhtiyarov, Max Ryabinin, Lucile Saulnier, An- ton Sinitsin, Dmitry Popov, Dmitry V Pyrkin, Maxim Kashirin, Alexander Borzunov, Albert Villanova del Moral, et al. Distributed deep learning in open collaborations. InAdvances in Neural Information Process...
2021
-
[72]
A simple algorithm for a class of nonsmooth convex–concave saddle-point problems.Operations Research Letters, 43(2):209–214, 2015
Yoel Drori, Shoham Sabach, and Marc Teboulle. A simple algorithm for a class of nonsmooth convex–concave saddle-point problems.Operations Research Letters, 43(2):209–214, 2015. ISSN 0167-6377.(Cited on page 186)
2015
-
[73]
The algorithmic foundations of differential privacy.Foundations and trends®in theoretical computer science, 9(3-4): 211–487, 2014.(Cited on pages 23, 26, and 28)
Cynthia Dwork and Aaron Roth. The algorithmic foundations of differential privacy.Foundations and trends®in theoretical computer science, 9(3-4): 211–487, 2014.(Cited on pages 23, 26, and 28)
2014
-
[74]
Cali- brating noise to sensitivity in private data analysis
Cynthia Dwork, Frank McSherry, Kobbi Nissim, and Adam Smith. Cali- brating noise to sensitivity in private data analysis. InTheory of cryptog- raphy conference, pages 265–284. Springer, 2006.(Cited on page 28)
2006
-
[75]
Eichner, T
H. Eichner, T. Koren, H. B. McMahan, N. Srebro, and K. Talwar. Semi- cyclic stochastic gradient descent. InInternational Conference on Machine Learning, 2019.(Cited on page 60) 155
2019
-
[76]
El Mahdi El-Mhamdi, Sadegh Farhadkhani, Rachid Guerraoui, Arsany Guirguis, Lˆ e-Nguyˆ en Hoang, and S´ ebastien Rouault. Collaborative learn- ing in the jungle (decentralized, byzantine, heterogeneous, asynchronous and nonconvex learning).Advances in neural information process...
2021
-
[77]
Spider: Near- optimal non-convex optimization via stochastic path-integrated differential estimator
Cong Fang, Chris Junchi Li, Zhouchen Lin, and Tong Zhang. Spider: Near- optimal non-convex optimization via stochastic path-integrated differential estimator. InAdvances in Neural Information Processing Systems, vol- ume 31, 2018.(Cited on pages 75, 123, 128, and 413)
2018
-
[78]
AFLGuard: Byzantine-robust asynchronous federated learning
Minghong Fang, Jia Liu, Neil Zhenqiang Gong, and Elizabeth S Bentley. AFLGuard: Byzantine-robust asynchronous federated learning. InCom- puter Security Applications Conference, pages 632–646, 2022.(Cited on page 350)
2022
-
[79]
Ef21 with bells & whistles: Practical algorithmic extensions of modern error feedback.arXiv preprint arXiv:2110.03294, 2021.(Cited on pages 273, 274, and 350)
Ilyas Fatkhullin, Igor Sokolov, Eduard Gorbunov, Zhize Li, and Peter Richt´ arik. Ef21 with bells & whistles: Practical algorithmic extensions of modern error feedback.arXiv preprint arXiv:2110.03294, 2021.(Cited on pages 273, 274, and 350)
2021 arXiv
-
[80]
Efficient evaluation of scaled proxi- mal operators.Electronic Transactions on Numerical Analysis, 46:1–23, 03 2016.(Cited on page 44)
Michael Friedlander and Gabriel Goh. Efficient evaluation of scaled proxi- mal operators.Electronic Transactions on Numerical Analysis, 46:1–23, 03 2016.(Cited on page 44)
2016
-
[81]
Stochastic first-and zeroth-order meth- ods for nonconvex stochastic programming.SIAM journal on optimization, 23(4):2341–2368, 2013.(Cited on page 406)
Saeed Ghadimi and Guanghui Lan. Stochastic first-and zeroth-order meth- ods for nonconvex stochastic programming.SIAM journal on optimization, 23(4):2341–2368, 2013.(Cited on page 406)
2013
-
[82]
Distributed new- ton can communicate less and resist Byzantine workers
Avishek Ghosh, Raj Kumar Maity, and Arya Mazumdar. Distributed new- ton can communicate less and resist Byzantine workers. InAdvances in Neu- ral Information Processing Systems, volume 33, pages 18028–18038, 2020. (Cited on page 350)
2020
-
[83]
Communication-efficient and byzantine-robust dis- tributed learning with error feedback.IEEE Journal on Selected Areas in Information Theory, 2(3):942–953, 2021.(Cited on page 350)
Avishek Ghosh, Raj Kumar Maity, Swanand Kadhe, Arya Mazumdar, and Kannan Ramchandran. Communication-efficient and byzantine-robust dis- tributed learning with error feedback.IEEE Journal on Selected Areas in Information Theory, 2(3):942–953, 2021.(Cited on page 350)
2021
-
[84]
Sharp bounds for federated averaging (local sgd) and continuous perspective
Margalit R Glasgow, Honglin Yuan, and Tengyu Ma. Sharp bounds for federated averaging (local sgd) and continuous perspective. InInterna- tional Conference on Artificial Intelligence and Statistics, pages 9050–9090. PMLR, 2022.(Cited on pages 29, 62, 76, 268, and 436) 156
2022
-
[85]
Television by pulse code modulation.Bell System Technical Journal, 30(1):33–49, 1951.(Cited on page 121)
WM Goodall. Television by pulse code modulation.Bell System Technical Journal, 30(1):33–49, 1951.(Cited on page 121)
1951
-
[86]
Stochas- tic optimization with heavy-tailed noise via accelerated gradient clipping
Eduard Gorbunov, Marina Danilova, and Alexander Gasnikov. Stochas- tic optimization with heavy-tailed noise via accelerated gradient clipping. InAdvances in Neural Information Processing Systems, volume 33, pages 15042–15053, 2020a.(Cited on pages 350, 354, and 395)
-
[87]
A unified theory of SGD: Variance reduction, sampling, quantization and coordinate descent
Eduard Gorbunov, Filip Hanzely, and Peter Richt´ arik. A unified theory of SGD: Variance reduction, sampling, quantization and coordinate descent. InInternational Conference on Artificial Intelligence and Statistics, pages 680–690. PMLR, 2020b.(Cited on pages 30, 65, 77, 109, ...
-
[88]
Linearly converging error compensated SGD.Advances in Neural Information Processing Systems, 33:20889–20900, 2020c.(Cited on page 65)
Eduard Gorbunov, Dmitry Kovalev, Dmitry Makarenko, and Peter Richt´ arik. Linearly converging error compensated SGD.Advances in Neural Information Processing Systems, 33:20889–20900, 2020c.(Cited on page 65)
-
[89]
Eduard Gorbunov, Konstantin Burlachenko, Zhize Li, and Peter Richt´ arik. Marina: Faster non-convex distributed learning with compression.Pro- ceedings of the 38th International Conference on Machine Learning, 139: 3788–3798, 18–24 Jul 2021a.(Cited on pages 30, 91, 125, 128, 3...
-
[90]
Local SGD: Unified theory and new efficient methods
Eduard Gorbunov, Filip Hanzely, and Peter Richt´ arik. Local SGD: Unified theory and new efficient methods. InInternational Conference on Artifi- cial Intelligence and Statistics, pages 3556–3564. PMLR, 2021b.(Cited on pages 29, 30, 32, 45, 46, 54, 61, 62, 65, 77, 91, 268, and 437)
-
[91]
Secure distributed training at scale
Eduard Gorbunov, Alexander Borzunov, Michael Diskin, and Max Ryabinin. Secure distributed training at scale. InInternational Confer- ence on Machine Learning, 2022. (arXiv preprint arXiv:2106.11257, 2021). (Cited on pages 117 and 349)
2022 arXiv
-
[92]
Eduard Gorbunov, Samuel Horv´ ath, Peter Richt´ arik, and Gauthier Gidel. Variance reduction is an antidote to Byzantines: Better rates, weaker as- sumptions and communication compression as a cherry on the top.Inter- national Conference on Learning Representations, 2023.(Cite...
2023
-
[93]
Variance-reduced methods for machine learning.Proceedings of the IEEE, 108(11):1968–1983, 2020.(Cited on pages 30 and 349) 157
Robert M Gower, Mark Schmidt, Francis Bach, and Peter Richt´ arik. Variance-reduced methods for machine learning.Proceedings of the IEEE, 108(11):1968–1983, 2020.(Cited on pages 30 and 349) 157
1968
-
[94]
Gower, Peter Richt´ arik, and Francis Bach
Robert M. Gower, Peter Richt´ arik, and Francis Bach. Stochastic quasi- gradient methods: Variance reduction via Jacobian sketching.Mathematical Programming, 188(1):135–192, 2021.(Cited on pages 48, 54, and 77)
2021
-
[95]
SGD: General analysis and improved rates
Robert Mansel Gower, Nicolas Loizou, Xun Qian, Alibek Sailanbayev, Egor Shulgin, and Peter Richt´ arik. SGD: General analysis and improved rates. In International Conference on Machine Learning, pages 5200–5209. PMLR, 2019.(Cited on pages 27, 30, 48, 54, 59, 60, 65, 66, 75, 77...
2019
-
[96]
Accurate, large minibatch sgd: training imagenet in 1 hour.arXiv preprint arXiv:1706.02677, 2017.(Cited on page 104)
Priya Goyal, Piotr Doll´ ar, Ross Girshick, Pieter Noordhuis, Lukasz Wlowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. Accurate, large minibatch sgd: training imagenet in 1 hour.arXiv preprint arXiv:1706.02677, 2017.(Cited on page 104)
2017 arXiv
-
[97]
Can 5th gen- eration local training methods support client sampling? yes! InInterna- tional Conference on Artificial Intelligence and Statistics, pages 1055–1092
Micha l Grudzie´ n, Grigory Malinovsky, and Peter Richt´ arik. Can 5th gen- eration local training methods support client sampling? yes! InInterna- tional Conference on Artificial Intelligence and Statistics, pages 1055–1092. PMLR, 2023a.(Cited on pages 31, 34, and 437)
-
[98]
Improving ac- celerated federated learning with compression and importance sampling
Micha l Grudzie´ n, Grigory Malinovsky, and Peter Richt´ arik. Improving ac- celerated federated learning with compression and importance sampling. arXiv preprint arXiv:2306.03240, 2023b.(Cited on pages 39 and 437)
-
[99]
On the convergence of local descent methods infederated learning.arXiv preprint arXiv:1910.14425, 2019.(Cited on pages 61 and 76)
Farzin Haddadpour and Mehrdad Mahdavi. On the convergence of local descent methods infederated learning.arXiv preprint arXiv:1910.14425, 2019.(Cited on pages 61 and 76)
1910 arXiv
-
[100]
Federated learning with compression: Unified analysis and sharp guarantees
Farzin Haddadpour, Mohammad Mahdi Kamani, Aryan Mokhtari, and Mehrdad Mahdavi. Federated learning with compression: Unified analysis and sharp guarantees. InInternational Conference on Artificial Intelligence and Statistics, pages 2350–2358, 2021.(Cited on pages 111, 114, 268,...
2021
-
[101]
Parameter- efficient fine-tuning for large models: A comprehensive survey.arXiv preprint arXiv:2403.14608, 2024.(Cited on page 131)
Zeyu Han, Chao Gao, Jinyang Liu, Sai Qian Zhang, et al. Parameter- efficient fine-tuning for large models: A comprehensive survey.arXiv preprint arXiv:2403.14608, 2024.(Cited on page 131)
2024 arXiv
-
[102]
Federated learning of a mixture of global and local models.arXiv preprint arXiv:2002.05516, 2020.(Cited on page 53)
Filip Hanzely and Peter Richt´ arik. Federated learning of a mixture of global and local models.arXiv preprint arXiv:2002.05516, 2020.(Cited on page 53)
2002 arXiv
-
[103]
Lower bounds and optimal algorithms for personalized federated learning
Filip Hanzely, Slavom´ ır Hanzely, Samuel Horv´ ath, and Peter Richt´ arik. Lower bounds and optimal algorithms for personalized federated learning. Advances in Neural Information Processing Systems, 33:2304–2315, 2020. (Cited on page 53) 158
2020
-
[104]
Towards a unified view of parameter-efficient transfer learning.arXiv preprint arXiv:2110.04366, 2021.(Cited on page 131)
Junxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick, and Graham Neubig. Towards a unified view of parameter-efficient transfer learning.arXiv preprint arXiv:2110.04366, 2021.(Cited on page 131)
2021 arXiv
-
[105]
Delving deep into rectifiers: Surpassing human-level performance on imagenet classifica- tion
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Delving deep into rectifiers: Surpassing human-level performance on imagenet classifica- tion. InIEEE International Conference on Computer Vision, pages 1026– 1034, 2015.(Cited on page 275)
2015
-
[106]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. InIEEE Conference on Computer Vision and Pattern Recognition, pages 770–778, 2016.(Cited on pages 115, 274, 275, and 417)
2016
-
[107]
Byzantine-robust decentralized learning via ClippedGossip.arXiv preprint arXiv:2202.01545, 2022.(Cited on page 348)
Lie He, Sai Praneeth Karimireddy, and Martin Jaggi. Byzantine-robust decentralized learning via ClippedGossip.arXiv preprint arXiv:2202.01545, 2022.(Cited on page 348)
2022 arXiv
-
[108]
M. R. Hestenes. Multiplier and gradient methods.Journal of Optimization Theory and Applications, 4(5):303–320, 1969.(Cited on page 78)
1969
-
[109]
Variance reduced stochastic gradient descent with neighbors
Thomas Hofmann, Aurelien Lucchi, Simon Lacoste-Julien, and Brian McWilliams. Variance reduced stochastic gradient descent with neighbors. Advances in Neural Information Processing Systems, 28, 2015.(Cited on pages 68, 71, and 77)
2015
-
[110]
Nonconvex variance reduced op- timization with arbitrary sampling
Samuel Horv´ ath and Peter Richt´ arik. Nonconvex variance reduced op- timization with arbitrary sampling. In Kamalika Chaudhuri and Rus- lan Salakhutdinov, editors,International Conference on Machine Learn- ing, volume 97 ofProceedings of Machine Learning Research, pages 2781...
2019
-
[111]
Natural compression for distributed deep learning.arXiv preprint arXiv:1905.10988, 2019a.(Cited on page 68)
Samuel Horv´ ath, Chen-Yu Ho,ˇLudov´ ıt Horv´ ath, Atal Narayan Sahu, Marco Canini, and Peter Richt´ arik. Natural compression for distributed deep learning.arXiv preprint arXiv:1905.10988, 2019a.(Cited on page 68)
1905 arXiv
-
[112]
Stochastic distributed learning with gradient quan- tization and variance reduction.arXiv preprint arXiv:1904.05115, 2019b
Samuel Horv´ ath, Dmitry Kovalev, Konstantin Mishchenko, Sebastian Stich, and Peter Richt´ arik. Stochastic distributed learning with gradient quan- tization and variance reduction.arXiv preprint arXiv:1904.05115, 2019b. (Cited on pages 30, 68, 71, 77, 105, 109, 119, 120, 209,...
1904 arXiv
-
[113]
FedShuffle: Recipes for better use of local work in federated learn- ing.arXiv preprint arXiv:2204.13169, 2022.(Cited on pages 268 and 423)
Samuel Horv´ ath, Maziar Sanjabi, Lin Xiao, Peter Richt´ arik, and Michael Rabbat. FedShuffle: Recipes for better use of local work in federated learn- ing.arXiv preprint arXiv:2204.13169, 2022.(Cited on pages 268 and 423)
2022 arXiv
-
[114]
Measuring the effects of non-identical data distribution for federated visual classification.arXiv preprint arXiv:1909.06335, 2019.(Cited on page 91)
Tzu-Ming Harry Hsu, Hang Qi, and Matthew Brown. Measuring the effects of non-identical data distribution for federated visual classification.arXiv preprint arXiv:1909.06335, 2019.(Cited on page 91)
1909 arXiv
-
[115]
LoRA: Low-rank adaptation of large language models.arXiv preprint arXiv:2106.09685, 2021.(Cited on pages 21, 22, 131, 142, and 419)
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models.arXiv preprint arXiv:2106.09685, 2021.(Cited on pages 21, 22, 131, 142, and 419)
2021 arXiv
-
[116]
Distributed random reshuffling over networks.arXiv preprint arXiv:2112.15287, abs/2112.15287, 2021
Kun Huang, Xiao Li, Andre Milzarek, Shi Pu, and Junwen Qiu. Distributed random reshuffling over networks.arXiv preprint arXiv:2112.15287, abs/2112.15287, 2021. URLhttps://arXiv.org/abs/2112.15287.(Cited on page 268)
2021 arXiv
-
[117]
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. InInternational Conference on Machine Learning, pages 448–456. PMLR, 2015.(Cited on page 275)
2015
-
[118]
Distributed second order methods with fast rates and compressed communication
Rustem Islamov, Xun Qian, and Peter Richt´ arik. Distributed second order methods with fast rates and compressed communication. InInternational Conference on Machine Learning, pages 4617–4628, 2021.(Cited on page 349)
2021
-
[119]
Byzantine-robust and differ- entially private federated optimization under weaker assumptions.arXiv preprint arXiv:2603.23472, 2026.(Cited on page 38)
Rustem Islamov, Grigory Malinovsky, Alexander Gaponov, Aurelien Luc- chi, Peter Richt´ arik, and Eduard Gorbunov. Byzantine-robust and differ- entially private federated optimization under weaker assumptions.arXiv preprint arXiv:2603.23472, 2026.(Cited on page 38)
2026 arXiv
-
[120]
Martin Jaggi, Virginia Smith, Martin Tak´ aˇ c, Jonathan Terhorst, Sanjay Kr- ishnan, Thomas Hofmann, and Michael I. Jordan. Communication-efficient distributed dual coordinate ascent. InAdvances in Neural Information Pro- cessing Systems, 2014.(Cited on pages 76 and 77)
2014
-
[121]
SGD without replacement: Sharper rates for general smooth convex functions.arXiv preprint arXiv:1903.01463, 2019.(Cited on page 423)
Prateek Jain, Dheeraj Nagaraj, and Praneeth Netrapalli. SGD without replacement: Sharper rates for general smooth convex functions.arXiv preprint arXiv:1903.01463, 2019.(Cited on page 423)
1903 arXiv
-
[122]
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang. Accelerating stochastic gradient descent using predictive variance reduction. InAdvances in Neural Information Processing Systems, volume 26, 2013.(Cited on pages 30, 63, 77, and 349) 160
2013
-
[123]
Brendan McMahan, Brendan Avent, Aur´ elien Bellet, Mehdi Bennis, Arjun Nitin Bhagoji, Keith Bonawitz, Zachary Charles, Gra- ham Cormode, Rachel Cummings, et al
Peter Kairouz, H. Brendan McMahan, Brendan Avent, Aur´ elien Bellet, Mehdi Bennis, Arjun Nitin Bhagoji, Keith Bonawitz, Zachary Charles, Gra- ham Cormode, Rachel Cummings, et al. Advances and open problems in federated learning.Foundations and Trends®in Machine Learning, 14(1)...
2021
-
[124]
Stich, and Ananda Theertha Suresh
Sai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank Reddi, Sebastian U. Stich, and Ananda Theertha Suresh. SCAFFOLD: Stochastic controlled averaging for federated learning. InInternational Conference on Machine Learning, pages 5132–5143. PMLR, 2020.(Cited on pages 27...
2020
-
[125]
Learning from history for Byzantine robust optimization
Sai Praneeth Karimireddy, Lie He, and Martin Jaggi. Learning from history for Byzantine robust optimization. InInternational Conference on Machine Learning, pages 5311–5319, 2021a.(Cited on pages 117, 118, 119, 120, 121, 129, 130, 348, 349, 350, and 417)
-
[126]
Reddi, Sebastian U
Sai Praneeth Karimireddy, Martin Jaggi, Satyen Kale, Mehryar Mohri, Sashank J. Reddi, Sebastian U. Stich, and Ananda Theertha Suresh. Break- ing the centralized barrier for cross-device federated learning. InAdvances in Neural Information Processing Systems. Curran Associates,...
-
[127]
Sai Praneeth Karimireddy, Lie He, and Martin Jaggi. Byzantine-robust learning on heterogeneous datasets via bucketing.International Conference on Learning Representations, 2022.(Cited on pages 117, 119, 120, 121, 122, 127, 348, 349, 351, and 354)
2022
-
[128]
Better theory for sgd in the nonconvex world.Transactions on Machine Learning Research, 2023.(Cited on pages 59, 75, 136, 261, 430, and 444)
Ahmed Khaled and Peter Richt´ arik. Better theory for sgd in the nonconvex world.Transactions on Machine Learning Research, 2023.(Cited on pages 59, 75, 136, 261, 430, and 444)
2023
-
[129]
First analysis of local GD on heterogeneous data.arXiv preprint arXiv:1909.04715, 2019
Ahmed Khaled, Konstantin Mishchenko, and Peter Richt´ arik. First analysis of local GD on heterogeneous data.arXiv preprint arXiv:1909.04715, 2019. (Cited on pages 28, 29, 45, 46, 61, 62, 76, and 436)
1909 arXiv
-
[130]
Tighter the- ory for local SGD on identical and heterogeneous data
Ahmed Khaled, Konstantin Mishchenko, and Peter Richt´ arik. Tighter the- ory for local SGD on identical and heterogeneous data. InInternational Conference on Artificial Intelligence and Statistics, 2020.(Cited on pages 28, 45, 46, 52, 61, 62, 76, 91, and 436) 161
2020
-
[131]
Ahmed Khaled, Othmane Sebbouh, Nicolas Loizou, Robert M Gower, and Peter Richt´ arik. Unified analysis of stochastic gradient methods for com- posite convex and smooth optimization.Journal of Optimization Theory and Applications, 199(2):499–540, 2023.(Cited on page 430)
2023
-
[132]
Distributed learning with compressed gradients.arXiv preprint arXiv:1806.06573, 2018.(Cited on pages 68, 105, 107, 120, and 349)
Sarit Khirirat, Hamid Reza Feyzmahdavian, and Mikael Johans- son. Distributed learning with compressed gradients.arXiv preprint arXiv:1806.06573, 2018.(Cited on pages 68, 105, 107, 120, and 349)
2018 arXiv
-
[133]
Mikhail Khodak, Renbo Tu, Tian Li, Liam Li, Maria-Florina F Balcan, Virginia Smith, and Ameet Talwalkar. Federated hyperparameter tun- ing: Challenges, baselines, and connections to weight-sharing.Advances in Neural Information Processing Systems, 34:19184–19197, 2021.(Cited o...
2021
-
[134]
A hybrid gpu cluster and volunteer computing platform for scalable deep learning.The Journal of Supercomputing, 74(7):3236–3263, 2018.(Cited on pages 117 and 122)
Ekasit Kijsipongse, Apivadee Piyatumrong, et al. A hybrid gpu cluster and volunteer computing platform for scalable deep learning.The Journal of Supercomputing, 74(7):3236–3263, 2018.(Cited on pages 117 and 122)
2018
-
[135]
A unified theory of decentralized SGD with changing topol- ogy and local updates
Anastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi, and Se- bastian Stich. A unified theory of decentralized SGD with changing topol- ogy and local updates. InInternational Conference on Machine Learning, pages 5381–5393. PMLR, 2020.(Cited on pages 28, 29, 46, 52...
2020
-
[136]
Anastasiia Koloskova, Tao Lin, and Sebastian U. Stich. An improved anal- ysis of gradient tracking for decentralized machine learning. InAdvances in Neural Information Processing Systems, volume 34. Curran Associates, Inc., 2021.(Cited on page 49)
2021
-
[137]
Federated opti- mization: Distributed optimization beyond the datacenter.arXiv preprint arXiv:1511.03575, 2015.(Cited on pages 24 and 26)
Jakub Koneˇ cny, Brendan McMahan, and Daniel Ramage. Federated opti- mization: Distributed optimization beyond the datacenter.arXiv preprint arXiv:1511.03575, 2015.(Cited on pages 24 and 26)
2015 arXiv
-
[138]
Semi-stochastic gradient descent methods.arXiv preprint arXiv:1312.1666, pages 1–14, 2017
Jakub Koneˇ cn´ y and Peter Richt´ arik. Semi-stochastic gradient descent methods.arXiv preprint arXiv:1312.1666, pages 1–14, 2017. URLhttp: //arxiv.org/abs/1312.1666.(Cited on page 77)
2017 arXiv
-
[139]
Brendan McMahan, Daniel Ramage, and Peter Richt´ arik
Jakub Koneˇ cn´ y, H. Brendan McMahan, Daniel Ramage, and Peter Richt´ arik. Federated optimization: distributed machine learning for on- device intelligence.arXiv preprint arXiv:1610.02527, 2016a.(Cited on pages 23, 24, 58, 74, and 89) 162
-
[140]
Brendan McMahan, Felix Yu, Peter Richt´ arik, Ananda Theertha Suresh, and Dave Bacon
Jakub Koneˇ cn´ y, H. Brendan McMahan, Felix Yu, Peter Richt´ arik, Ananda Theertha Suresh, and Dave Bacon. Federated learning: strate- gies for improving communication efficiency. InNIPS Private Multi-Party Machine Learning Workshop, 2016b.(Cited on pages 23, 24, 45, 58, 74, ...
-
[141]
Don’t jump through hoops and remove those loops: Svrg and katyusha are better without the outer loop
Dmitry Kovalev, Samuel Horv´ ath, and Peter Richt´ arik. Don’t jump through hoops and remove those loops: Svrg and katyusha are better without the outer loop. InAlgorithmic Learning Theory, pages 451–467. PMLR, 2020a. (Cited on pages 63, 68, 71, 77, 210, and 239)
-
[142]
Optimal and practi- cal algorithms for smooth and strongly convex decentralized optimization
Dmitry Kovalev, Adil Salim, and Peter Richt´ arik. Optimal and practi- cal algorithms for smooth and strongly convex decentralized optimization. Neural Information Processing Systems (NeurIPS), 33:18342–18352, 2020b. (Cited on pages 49 and 56)
-
[143]
Lower bounds and optimal algorithms for smooth and strongly convex de- centralized optimization over time-varying networks
Dmitry Kovalev, Elnur Gasanov, Peter Richt´ arik, and Alexander Gasnikov. Lower bounds and optimal algorithms for smooth and strongly convex de- centralized optimization over time-varying networks. InAdvances in Neural Information Processing Systems, 2021a.(Cited on page 49)
-
[144]
ADOM: Accelerated decentralized optimization method for time-varying networks
Dmitry Kovalev, Egor Shulgin, Peter Richt´ arik, Alexander Rogozin, and Alexander Gasnikov. ADOM: Accelerated decentralized optimization method for time-varying networks. InInternational Conference on Ma- chine Learning, pages 5784–5793. PMLR, 2021b.(Cited on page 49)
-
[145]
An optimal algorithm for strongly convex min-min optimization
Dmitry Kovalev, Alexander Gasnikov, and Grigory Malinovsky. An optimal algorithm for strongly convex min-min optimization. InConference on Uncertainty in Artificial Intelligence, 2025.(Cited on page 39)
2025
-
[146]
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009.(Cited on pages 102, 115, 274, 275, and 417)
2009
-
[147]
FederatedScope-LLM: A comprehensive package for fine-tuning large lan- guage models in federated learning
Weirui Kuang, Bingchen Qian, Zitao Li, Daoyuan Chen, Dawei Gao, Xuchen Pan, Yuexiang Xie, Yaliang Li, Bolin Ding, and Jingren Zhou. FederatedScope-LLM: A comprehensive package for fine-tuning large lan- guage models in federated learning. InACM SIGKDD International Con- ferenc...
2024
-
[148]
The Byzantine gener- als problem.ACM Transactions on Programming Languages and Systems, 4(3):382–401, 1982.(Cited on page 117) 163
Leslie Lamport, Robert Shostak, and Marshall Pease. The Byzantine gener- als problem.ACM Transactions on Programming Languages and Systems, 4(3):382–401, 1982.(Cited on page 117) 163
1982
-
[149]
An optimal method for stochastic composite optimization
Guanghui Lan. An optimal method for stochastic composite optimization. Mathematical Programming, 133:365–397, 2012.(Cited on page 46)
2012
-
[150]
The mnist database of handwritten digits
Yann LeCun and Corinna Cortes. The mnist database of handwritten digits. http://yann.lecun.com/exdb/mnist/, 1998. URLhttps://cir.nii.ac.jp/ crid/1571417126193283840.(Cited on pages 130 and 417)
1998
-
[151]
Demystifying gpt-3 language model: A technical overview, 2020
Chuan Li. Demystifying gpt-3 language model: A technical overview, 2020. ”https://lambdalabs.com/blog/demystifying-[]gpt-[]3”.(Cited on page 117)
2020
-
[152]
Mea- suring the intrinsic dimension of objective landscapes.arXiv preprint arXiv:1804.08838, 2018.(Cited on page 131)
Chunyuan Li, Heerad Farkhoor, Rosanne Liu, and Jason Yosinski. Mea- suring the intrinsic dimension of objective landscapes.arXiv preprint arXiv:1804.08838, 2018.(Cited on page 131)
2018 arXiv
-
[153]
Mu Li, Tong Zhang, Yuqiang Chen, and Alexander J. Smola. Efficient mini- batch training for stochastic optimization. InACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD ’14, pages 661–670, New York, NY, USA, 2014. ACM. ISBN 978-1-4503-2956-9. doi:...
2014
-
[154]
Fed- erated learning: challenges, methods, and future directions.IEEE Signal Processing Magazine, 37(3):50–60, 2020a
Tian Li, Anit Kumar Sahu, Ameet Talwalkar, and Virginia Smith. Fed- erated learning: challenges, methods, and future directions.IEEE Signal Processing Magazine, 37(3):50–60, 2020a. doi: 10.1109/MSP.2020.2975749. (Cited on pages 27, 58, and 74)
2020
-
[155]
Communication-efficient local decentralized SGD methods.arXiv preprint arXiv:1910.09126, 2019.(Cited on page 61)
Xiang Li, Wenhao Yang, Shusen Wang, and Zhihua Zhang. Communication-efficient local decentralized SGD methods.arXiv preprint arXiv:1910.09126, 2019.(Cited on page 61)
1910 arXiv
-
[156]
On the convergence of FedAvg on non-IID data
Xiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang, and Zhihua Zhang. On the convergence of FedAvg on non-IID data. InInternational Confer- ence on Learning Representations, 2020b.(Cited on pages 61 and 76)
-
[157]
Acceleration for compressed gradient descent in distributed and federated optimization
Zhize Li, Dmitry Kovalev, Xun Qian, and Peter Richt´ arik. Acceleration for compressed gradient descent in distributed and federated optimization. InInternational Conference on Machine Learning, pages 5895–5904, 2020c. (Cited on page 349)
-
[158]
PAGE: A simple and optimal probabilistic gradient estimator for nonconvex optimiza- tion
Zhize Li, Hongyan Bao, Xiangliang Zhang, and Peter Richt´ arik. PAGE: A simple and optimal probabilistic gradient estimator for nonconvex optimiza- tion. InInternational Conference on Machine Learning, pages 6286–6295, 2021.(Cited on pages 75, 119, 128, 349, 355, 406, and 413) 164
2021
-
[159]
Stich, and Martin Jaggi
Tao Lin, Sebastian U. Stich, and Martin Jaggi. Don’t use large mini- batches, use local SGD. InInternational Conference on Learning Rep- resentations, 2018.(Cited on page 46)
2018
-
[160]
Quasi-global momentum: Accelerating decentralized deep learning on het- erogeneous data
Tao Lin, Sai Praneeth Karimireddy, Sebastian Stich, and Martin Jaggi. Quasi-global momentum: Accelerating decentralized deep learning on het- erogeneous data. InInternational Conference on Machine Learning, volume 139, pages 6654–6665. PMLR, 2021.(Cited on page 53)
2021
-
[161]
Federated learning meets natural language processing: a survey.arXiv preprint arXiv:2107.12603, abs/2107.12603, 2021
Ming Liu, Stella Ho, Mengqi Wang, Longxiang Gao, Yuan Jin, and He Zhang. Federated learning meets natural language processing: a survey.arXiv preprint arXiv:2107.12603, abs/2107.12603, 2021. URL https://arXiv.org/abs/2107.12603.(Cited on page 104)
2021 arXiv
-
[162]
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu. RoBERTa: A robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692, 2019.(Cited on page 142)
1907 arXiv
-
[163]
A topological property of real analytic subsets.Coll
Stanislaw Lojasiewicz. A topological property of real analytic subsets.Coll. du CNRS, Les ´ equations aux d´ eriv´ ees partielles, 117:87–89, 1963.(Cited on page 122)
1963
-
[164]
NEXT: In-network nonconvex op- timization.IEEE Transactions on Signal and Information Processing over Networks, 2(2):120–136, 2016
Paolo Di Lorenzo and Gesualdo Scutari. NEXT: In-network nonconvex op- timization.IEEE Transactions on Signal and Information Processing over Networks, 2(2):120–136, 2016. doi: 10.1109/TSIPN.2016.2524588.(Cited on page 48)
2016
-
[165]
On a generalization of the iterative soft-thresholding algorithm for the case of non-separable penalty.Inverse Problems, 27(12), 2011.(Cited on page 186)
Ignace Loris and Caroline Verhoeven. On a generalization of the iterative soft-thresholding algorithm for the case of non-separable penalty.Inverse Problems, 27(12), 2011.(Cited on page 186)
2011
-
[166]
Decoupled weight decay regularization.arXiv preprint arXiv:1711.05101, 2017.(Cited on page 142)
I Loshchilov. Decoupled weight decay regularization.arXiv preprint arXiv:1711.05101, 2017.(Cited on page 142)
2017 arXiv
-
[167]
A general analysis of example-selection for stochastic gradient descent
Yucheng Lu, Si Yi Meng, and Christopher De Sa. A general analysis of example-selection for stochastic gradient descent. InInternational Confer- ence on Learning Representations, volume 10, 2022.(Cited on page 423)
2022
-
[168]
Russell Luke.Proximal Methods for Image Processing, pages 165–202
D. Russell Luke.Proximal Methods for Image Processing, pages 165–202. Springer International Publishing, Cham, 2020. ISBN 978-3-030-34413-9. doi: 10.1007/978-[]3-[]030-[]34413-[]9 6.(Cited on page 43)
2020 doi
-
[169]
Privacy and robustness in federated learning: Attacks 165 and defenses.arXiv preprint arXiv:2012.06337, 2020.(Cited on pages 28, 348, and 349)
Lingjuan Lyu, Han Yu, Xingjun Ma, Lichao Sun, Jun Zhao, Qiang Yang, and Philip S Yu. Privacy and robustness in federated learning: Attacks 165 and defenses.arXiv preprint arXiv:2012.06337, 2020.(Cited on pages 28, 348, and 349)
2012 arXiv
-
[170]
Jordan, Peter Richt´ arik, and Martin Tak´ aˇ c
Chenxin Ma, Virginia Smith, Martin Jaggi, Michael I. Jordan, Peter Richt´ arik, and Martin Tak´ aˇ c. Adding vs. averaging in distributed primal- dual optimization. InInternational Conference on Machine Learning, pages 1973–1982, 2015.(Cited on pages 76 and 77)
1973
-
[171]
Jordan, Peter Richt´ arik, and Martin Tak´ aˇ c
Chenxin Ma, Jakub Koneˇ cn´ y, Martin Jaggi, Virginia Smith, Michael I. Jordan, Peter Richt´ arik, and Martin Tak´ aˇ c. Distributed optimization with arbitrary local solvers.Optimization Methods and Software, 32(4):813–848, 2017.(Cited on pages 76 and 77)
2017
-
[172]
Stability and convergence of stochastic gradient clipping: Beyond Lipschitz continuity and smoothness
Vien V Mai and Mikael Johansson. Stability and convergence of stochastic gradient clipping: Beyond Lipschitz continuity and smoothness. InInter- national Conference on Machine Learning, pages 7325–7335, 2021.(Cited on page 350)
2021
-
[173]
From local SGD to local fixed-point methods for federated learning
Grigory Malinovskiy, Dmitry Kovalev, Elnur Gasanov, Laurent Condat, and Peter Richtarik. From local SGD to local fixed-point methods for federated learning. InInternational Conference on Machine Learning, pages 6692–6701. PMLR, 2020.(Cited on pages 52, 62, 76, and 436)
2020
-
[174]
Federated random reshuffling with compression and variance reduction.arXiv preprint arXiv:2205.03914, 2022.(Cited on pages 39, 112, 113, 268, 312, 313, and 423)
Grigory Malinovsky and Peter Richt´ arik. Federated random reshuffling with compression and variance reduction.arXiv preprint arXiv:2205.03914, 2022.(Cited on pages 39, 112, 113, 268, 312, 313, and 423)
2022 arXiv
-
[175]
Variance reduced Prox- Skip: Algorithm, theory and application to federated learning
Grigory Malinovsky, Kai Yi, and Peter Richt´ arik. Variance reduced Prox- Skip: Algorithm, theory and application to federated learning. InAdvances in Neural Information Processing Systems, 2022.(Cited on pages 31, 33, 76, 77, 78, 87, and 437)
2022
-
[176]
Federated learning with regularized client participation.arXiv preprint arXiv:2302.03662, 2023a.(Cited on pages 39 and 423)
Grigory Malinovsky, Samuel Horv´ ath, Konstantin Burlachenko, and Peter Richt´ arik. Federated learning with regularized client participation.arXiv preprint arXiv:2302.03662, 2023a.(Cited on pages 39 and 423)
-
[177]
Server- side stepsizes and sampling without replacement provably help in federated optimization
Grigory Malinovsky, Konstantin Mishchenko, and Peter Richt´ arik. Server- side stepsizes and sampling without replacement provably help in federated optimization. InInternational Workshop on Distributed Machine Learning, pages 85–104, 2023b.(Cited on pages 31, 35, 106, 268, 28...
-
[178]
Random reshuffling with variance reduction: New analysis and better rates
Grigory Malinovsky, Alibek Sailanbayev, and Peter Richt´ arik. Random reshuffling with variance reduction: New analysis and better rates. InCon- ference on Uncertainty in Artificial Intelligence, pages 1347–1357. PMLR, 2023c.(Cited on pages 39, 113, 313, and 423)
-
[179]
Randomized asymmetric chain of lora: The first meaningful theoretical framework for low-rank adaptation.arXiv preprint arXiv:2410.08305, 2024a.(Cited on pages 31 and 38)
Grigory Malinovsky, Umberto Michieli, Hasan Abed Al Kader Hammoud, Taha Ceritli, Hayder Elesedy, Mete Ozay, and Peter Richt´ arik. Randomized asymmetric chain of lora: The first meaningful theoretical framework for low-rank adaptation.arXiv preprint arXiv:2410.08305, 2024a.(Ci...
-
[180]
Grigory Malinovsky, Peter Richt´ arik, Samuel Horv´ ath, and Eduard Gor- bunov. Byzantine robustness and partial participation can be achieved at once: Just clip gradient differences.Advances in Neural Information Pro- cessing Systems, 37:34900–34979, 2024b.(Cited on pages 31 and 37)
-
[181]
Adaptive gradient descent without descent
Yura Malitsky and Konstantin Mishchenko. Adaptive gradient descent without descent. InInternational Conference on Machine Learning, vol- ume 119, pages 6702–6712. PMLR, 2020.(Cited on pages 93 and 102)
2020
-
[182]
Mangasarian
L. Mangasarian. Parallel gradient distribution in unconstrained optimiza- tion.SIAM Journal on Control and Optimization, 33(6):1916–1925, 1995a. (Cited on page 76)
1916
-
[183]
Parallel gradient distribution in unconstrained optimiza- tion.SIAM Journal on Control and Optimization, 33(6):1916–1925, 1995b
LO Mangasarian. Parallel gradient distribution in unconstrained optimiza- tion.SIAM Journal on Control and Optimization, 33(6):1916–1925, 1995b. (Cited on page 26)
1916
-
[184]
Mangasarian and Mikhail V
Olvi L. Mangasarian and Mikhail V. Solodov. Backpropagation conver- gence via deterministic nonmonotone perturbed minimization. In J. Cowan, G. Tesauro, and J. Alspector, editors,Advances in Neural Information Pro- cessing Systems, volume 6. Morgan-Kaufmann, 1994.(Cited on page 45)
1994
-
[185]
Gradskip: Communication-accelerated local gradient methods with better computa- tional complexity.arXiv preprint arXiv:2210.16402, 2022.(Cited on page 437)
Artavazd Maranjyan, Mher Safaryan, and Peter Richt´ arik. Gradskip: Communication-accelerated local gradient methods with better computa- tional complexity.arXiv preprint arXiv:2210.16402, 2022.(Cited on page 437)
2022 arXiv
-
[186]
Distributed training strate- gies for the structured perceptron
Ryan McDonald, Keith Hall, and Gideon Mann. Distributed training strate- gies for the structured perceptron. InHuman Language Technologies: The 2010 Annual Conference of the North American Chapter of the Association for Computational Linguistics, pages 456–464. Association for...
2010
-
[187]
Brendan McMahan and Daniel Ramage
H. Brendan McMahan and Daniel Ramage. Federated learning: Collabo- rative machine learning without centralized training data. GoogleAIBlog, April 2017.(Cited on pages 23, 24, 58, and 74)
2017
-
[188]
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Ag¨ uera y Arcas
H. Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Ag¨ uera y Arcas. Communication-efficient learning of deep networks from decentralized data. InInternational Conference on Artificial Intelli- gence and Statistics, pages 1273–1282. PMLR, 2017.(Cited on p...
2017
-
[189]
Distributed learning with compressed gradient differences.arXiv preprint arXiv:1901.09269, 2019.(Cited on pages 30, 63, 68, 71, 77, 105, 109, 114, 116, 209, 271, 276, and 349)
Konstantin Mishchenko, Eduard Gorbunov, Martin Tak´ aˇ c, and Peter Richt´ arik. Distributed learning with compressed gradient differences.arXiv preprint arXiv:1901.09269, 2019.(Cited on pages 30, 63, 68, 71, 77, 105, 109, 114, 116, 209, 271, 276, and 349)
1901 arXiv
-
[190]
Konstantin Mishchenko, Ahmed Khaled, and Peter Richt´ arik. Random Reshuffling: Simple analysis with vast improvements.Advances in Neural Information Processing Systems, 33:17309–17320, 2020.(Cited on pages 30, 91, 95, 97, 98, 105, 108, 109, 116, 259, 265, 423, and 426)
2020
-
[191]
Proximal and federated random reshuffling
Konstantin Mishchenko, Ahmed Khaled, and Peter Richt´ arik. Proximal and federated random reshuffling. InInternational Conference on Machine Learning, pages 15718–15749. PMLR, 2022a.(Cited on pages 91, 92, 93, 101, 102, 107, 109, 112, 266, 268, 281, 312, 329, 334, and 423)
-
[192]
Konstantin Mishchenko, Grigory Malinovsky, Sebastian Stich, and Peter Richt´ arik. ProxSkip: Yes! Local gradient steps provably lead to commu- nication acceleration! Finally! InInternational Conference on Machine Learning, 2022b.(Cited on pages 31, 33, 61, 62, 63, 66, 67, 77, ...
-
[193]
Linear convergence in federated learning: Tackling client heterogeneity and sparse gradients.Advances in Neural Information Processing Systems, 34, 2021
Aritra Mitra, Rayana Jaafar, George Pappas, and Hamed Hassani. Linear convergence in federated learning: Tackling client heterogeneity and sparse gradients.Advances in Neural Information Processing Systems, 34, 2021. (Cited on pages 29, 32, 45, 46, 53, 61, 62, 77, and 268)
2021
-
[194]
Microadam: Accurate adaptive optimization with low space overhead and provable convergence
Ionut-Vlad Modoranu, Mher Safaryan, Grigory Malinovsky, Eldar Kurtic, Thomas Robert, Peter Richt´ arik, and Dan Alistarh. Microadam: Accurate adaptive optimization with low space overhead and provable convergence. Advances in Neural Information Processing Systems, 37:1–43, 202...
2024
-
[195]
Philipp Moritz, Robert Nishihara, Ion Stoica, and Michael I. Jordan. SparkNet: Training deep networks in Spark. InInternational Conference on Learning Representations, 2016.(Cited on pages 59, 61, and 76)
2016
-
[196]
Jordan, and Ion Stoica
Philipp Moritz, Robert Nishihara, Stephanie Wang, Alexey Tumanov, Richard Liaw, Eric Liang, Melih Elibol, Zongheng Yang, William Paul, Michael I. Jordan, and Ion Stoica. Ray: A distributed framework for emerg- ing AI applications. InUSENIX Symposium on Operating Systems Design...
2018
-
[197]
Bias-variance reduced local SGD for less heterogeneous federated learning
Tomoya Murata and Taiji Suzuki. Bias-variance reduced local SGD for less heterogeneous federated learning. In Marina Meila and Tong Zhang, editors, International Conference on Machine Learning, volume 139 ofProceedings of Machine Learning Research, pages 7872–7881. PMLR, 2021....
2021
-
[198]
Algorithms of robust stochastic optimization based on mirror descent method.Automation and Remote Control, 80:1607–1627, 2019.(Cited on page 350)
Alexander V Nazin, Arkadi S Nemirovsky, Alexandre B Tsybakov, and Anatoli B Juditsky. Algorithms of robust stochastic optimization based on mirror descent method.Automation and Remote Control, 80:1607–1627, 2019.(Cited on page 350)
2019
-
[199]
Linear convergence of first order methods for non-strongly convex optimization.Mathematical Programming, 175:69–107, 2019.(Cited on page 122)
Ion Necoara, Yurii Nesterov, and Fran¸ cois Glineur. Linear convergence of first order methods for non-strongly convex optimization.Mathematical Programming, 175:69–107, 2019.(Cited on page 122)
2019
-
[200]
Convergence rate of incremental sub- gradient algorithms.Stochastic optimization: algorithms and applications, pages 223–264, 2001.(Cited on page 423)
Angelia Nedi´ c and Dimitri Bertsekas. Convergence rate of incremental sub- gradient algorithms.Stochastic optimization: algorithms and applications, pages 223–264, 2001.(Cited on page 423)
2001
-
[201]
Distributed subgradient methods for multi-agent optimization.IEEE Transactions on automatic control, 54(1): 48–61, 2009.(Cited on page 26)
Angelia Nedic and Asuman Ozdaglar. Distributed subgradient methods for multi-agent optimization.IEEE Transactions on automatic control, 54(1): 48–61, 2009.(Cited on page 26)
2009
-
[202]
Distributed asynchronous incremental subgradient methods.Studies in Computational Mathematics, 8(C):381–407, 2001.(Cited on page 350)
Angelia Nedi´ c, Dimitri P Bertsekas, and Vivek S Borkar. Distributed asynchronous incremental subgradient methods.Studies in Computational Mathematics, 8(C):381–407, 2001.(Cited on page 350)
2001
-
[203]
Achieving geometric conver- gence for distributed optimization over time-varying graphs.SIAM Journal on Optimization, 27, 07 2016
Angelia Nedi´ c, Alex Olshevsky, and Wei Shi. Achieving geometric conver- gence for distributed optimization over time-varying graphs.SIAM Journal on Optimization, 27, 07 2016. doi: 10.1137/16M1084316.(Cited on page 48)
2016 doi
-
[204]
Robust stochastic approximation approach to stochastic program- 169 ming.SIAM Journal on Optimization, 19(4):1574–1609, 2009.(Cited on page 406)
Arkadii S Nemirovski, Anatoli B Juditsky, Guanghui Lan, and Alexander Shapiro. Robust stochastic approximation approach to stochastic program- 169 ming.SIAM Journal on Optimization, 19(4):1574–1609, 2009.(Cited on page 406)
2009
-
[205]
Kluwer Academic Publishers, 2004.(Cited on pages 24, 45, 60, 136, and 419)
Yurii Nesterov.Introductory lectures on convex optimization: a basic course (Applied Optimization). Kluwer Academic Publishers, 2004.(Cited on pages 24, 45, 60, 136, and 419)
2004
-
[206]
Gradient methods for minimizing composite functions
Yurii Nesterov. Gradient methods for minimizing composite functions. Mathematical Programming, 140(1):125–161, 2013.(Cited on pages 43 and 63)
2013
-
[207]
Sarah: A novel method for machine learning problems using stochastic recursive gradient
Lam M Nguyen, Jie Liu, Katya Scheinberg, and Martin Tak´ aˇ c. Sarah: A novel method for machine learning problems using stochastic recursive gradient. InInternational Conference on Machine Learning, pages 2613– 2621, 2017.(Cited on pages 75 and 349)
2017
-
[208]
A unified convergence analysis for shuffling-type gradient methods.Journal of Machine Learning Research, 22(207):1–44, 2021.(Cited on page 423)
Lam M Nguyen, Quoc Tran-Dinh, Dzung T Phan, Phuong Ha Nguyen, and Marten Van Dijk. A unified convergence analysis for shuffling-type gradient methods.Journal of Machine Learning Research, 22(207):1–44, 2021.(Cited on page 423)
2021
-
[209]
Improved convergence in high probability of clipped gradient methods with heavy tails.arXiv preprint arXiv:2304.01119, 2023.(Cited on page 350)
Ta Duy Nguyen, Alina Ene, and Huy L Nguyen. Improved convergence in high probability of clipped gradient methods with heavy tails.arXiv preprint arXiv:2304.01119, 2023.(Cited on page 350)
2023 arXiv
-
[210]
Billion-scale federated learning on mobile clients: A submodel design with tunable privacy
Chaoyue Niu, Fan Wu, Shaojie Tang, Lifeng Hua, Rongfei Jia, Chengfei Lv, Zhihua Wu, and Guihai Chen. Billion-scale federated learning on mobile clients: A submodel design with tunable privacy. InInternational Con- ference on Mobile Computing and Networking, pages 1–14, 2020.(C...
2020
-
[211]
Trade-offs of local SGD at scale: an empirical study
Jose Javier Gonzalez Ortiz, Jonathan Frankle, Mike Rabbat, Ari Morcos, and Nicolas Ballas. Trade-offs of local SGD at scale: an empirical study. arXiv preprint arXiv:2110.08133, abs/2110.08133, 2021. URLhttps:// arXiv.org/abs/2110.08133.(Cited on page 268)
2021 arXiv
-
[212]
Proximal algorithms.Foundations and Trends in Optimization, 1(3):127–239, jan 2014
Neal Parikh and Stephen Boyd. Proximal algorithms.Foundations and Trends in Optimization, 1(3):127–239, jan 2014. ISSN 2167-3888.(Cited on pages 43, 44, and 45)
2014
-
[213]
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio. On the difficulty of training recurrent neural networks. InInternational Conference on Machine Learning, pages 1310–1318, 2013.(Cited on page 350) 170
2013
-
[214]
Automatic differentiation in PyTorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. Automatic differentiation in PyTorch. 2017.(Cited on page 102)
2017
-
[215]
Constantin Philippenko and Aymeric Dieuleveut. Bidirectional com- pression in heterogeneous settings for distributed or federated learning with partial participation: tight convergence guarantees.arXiv preprint arXiv:2006.14591, 2020.(Cited on pages 30 and 68)
2006 arXiv
-
[216]
Robust aggregation for federated learning.IEEE Transactions on Signal Processing, 70:1142– 1154, 2022.(Cited on pages 118, 130, and 417)
Krishna Pillutla, Sham M Kakade, and Zaid Harchaoui. Robust aggregation for federated learning.IEEE Transactions on Signal Processing, 70:1142– 1154, 2022.(Cited on pages 118, 130, and 417)
2022
-
[217]
Gradient methods for the minimisation of functionals
Boris T Polyak. Gradient methods for the minimisation of functionals. USSR Computational Mathematics and Mathematical Physics, 3(4):864– 878, 1963.(Cited on page 122)
1963
-
[218]
Parallel training of DNNs with natural gradient and parameter averaging
Daniel Povey, Xiaohui Zhang, and Sanjeev Khudanpur. Parallel training of DNNs with natural gradient and parameter averaging. InICLR Workshop, 2015.(Cited on pages 59, 61, and 76)
2015
-
[219]
M. J. D. Powell. A method for nonlinear constraints in minimization prob- lems.Optimization, pages 283–298, 1969.(Cited on page 78)
1969
-
[220]
Error compensated distributed SGD can be accelerated
Xun Qian, Peter Richt´ arik, and Tong Zhang. Error compensated distributed SGD can be accelerated. InAdvances in Neural Information Processing Systems, volume 34, 2021.(Cited on page 349)
2021
-
[221]
Communication-efficient gluon in federated learning.arXiv preprint arXiv:2604.10689, 2026.(Cited on page 38)
Xun Qian, Alexander Gaponov, Grigory Malinovsky, and Peter Richt´ arik. Communication-efficient gluon in federated learning.arXiv preprint arXiv:2604.10689, 2026.(Cited on page 38)
2026 arXiv
-
[222]
Detox: A redundancy-based framework for faster and more ro- bust gradient aggregation
Shashank Rajput, Hongyi Wang, Zachary Charles, and Dimitris Papail- iopoulos. Detox: A redundancy-based framework for faster and more ro- bust gradient aggregation. InAdvances in Neural Information Processing Systems, volume 32, 2019.(Cited on page 349)
2019
-
[223]
Papailiopoulos
Shashank Rajput, Anant Gupta, and Dimitris S. Papailiopoulos. Closing the convergence gap of SGD without replacement. InInternational Confer- ence on Machine Learning, volume 119 ofProceedings of Machine Learning Research, pages 7964–7973. PMLR, 2020. URLhttp://proceedings.mlr...
2020
-
[224]
Toward a noncommutative arithmetic-geometric mean inequality: Conjectures, case-studies, and con- sequences
Benjamin Recht and Christopher R´ e. Toward a noncommutative arithmetic-geometric mean inequality: Conjectures, case-studies, and con- sequences. InConference on Learning Theory, pages 11–1. JMLR Workshop and Conference Proceedings, 2012.(Cited on page 423)
2012
-
[225]
Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett, Keith Rush, Jakub Koneˇ cn´ y, Sanjiv Kumar, and H
Sashank J. Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett, Keith Rush, Jakub Koneˇ cn´ y, Sanjiv Kumar, and H. Brendan McMahan. Adaptive federated optimization. InInternational Conference on Learning Represen- tations, 2020.(Cited on pages 91 and 93)
2020
-
[226]
ByGARS: Byzantine SGD with arbitrary number of attackers.arXiv preprint arXiv:2006.13421, 2020.(Cited on page 349)
Jayanth Regatti, Hao Chen, and Abhishek Gupta. ByGARS: Byzantine SGD with arbitrary number of attackers.arXiv preprint arXiv:2006.13421, 2020.(Cited on page 349)
2006 arXiv
-
[227]
Fedpaq: A communication-efficient federated learn- ing method with periodic averaging and quantization
Amirhossein Reisizadeh, Aryan Mokhtari, Hamed Hassani, Ali Jadbabaie, and Ramtin Pedarsani. Fedpaq: A communication-efficient federated learn- ing method with periodic averaging and quantization. In Silvia Chiappa and Roberto Calandra, editors,International Conference on Artif...
2021
-
[228]
Parallel coordinate descent methods for big data optimization.Mathematical Programming, 156(1-2):433–484, 2016.(Cited on page 131)
Peter Richt´ arik and Martin Tak´ aˇ c. Parallel coordinate descent methods for big data optimization.Mathematical Programming, 156(1-2):433–484, 2016.(Cited on page 131)
2016
-
[229]
EF21: A new, simpler, theoretically better, and practically faster error feedback
Peter Richt´ arik, Igor Sokolov, and Ilyas Fatkhullin. EF21: A new, simpler, theoretically better, and practically faster error feedback. InAdvances in Neural Information Processing Systems, volume 34, 2021.(Cited on pages 91, 350, and 351)
2021
-
[230]
Error feed- back reloaded: From quadratic to arithmetic mean of smoothness constants
Peter Richt´ arik, Elnur Gasanov, and Konstantin Burlachenko. Error feed- back reloaded: From quadratic to arithmetic mean of smoothness constants. arXiv preprint arXiv:2402.10774, 2024.(Cited on page 274)
2024 arXiv
-
[231]
A stochastic approximation method
Herbert Robbins and Sutton Monro. A stochastic approximation method. The annals of mathematical statistics, pages 400–407, 1951.(Cited on pages 30, 75, and 105)
1951
-
[232]
Picture coding using pseudo-random noise.IRE Trans- actions on Information Theory, 8(2):145–154, 1962.(Cited on page 121) 172
Lawrence Roberts. Picture coding using pseudo-random noise.IRE Trans- actions on Information Theory, 8(2):145–154, 1962.(Cited on page 121) 172
1962
-
[233]
Dy- namic federated learning model for identifying adversarial clients.arXiv preprint arXiv:2007.15030, 2020.(Cited on page 349)
Nuria Rodr´ ıguez-Barroso, Eugenio Mart´ ınez-C´ amara, M Luz´ on, Ger- ardo Gonz´ alez Seco, Miguel´Angel Veganzones, and Francisco Herrera. Dy- namic federated learning model for identifying adversarial clients.arXiv preprint arXiv:2007.15030, 2020.(Cited on page 349)
2007 arXiv
-
[234]
An overview of gradient descent optimization algorithms
Sebastian Ruder. An overview of gradient descent optimization algorithms. arXiv preprint arXiv:1609.04747, 2016.(Cited on page 135)
2016 arXiv
-
[235]
Towards crowdsourced training of large neural networks using decentralized mixture-of-experts
Max Ryabinin and Anton Gusev. Towards crowdsourced training of large neural networks using decentralized mixture-of-experts. In H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, editors,Advances in Neural Information Processing Systems, volume 33, pages 3659–367...
2020
-
[236]
Communica- tion acceleration of local gradient methods via an accelerated primal-dual algorithm with inexact prox
Abdurakhmon Sadiev, Dmitry Kovalev, and Peter Richt´ arik. Communica- tion acceleration of local gradient methods via an accelerated primal-dual algorithm with inexact prox. InAdvances in Neural Information Processing Systems, 2022.(Cited on pages 78 and 437)
2022
-
[237]
High-probability bounds for stochastic optimization and variational inequalities: the case of unbounded variance.arXiv preprint arXiv:2302.00999, 2023.(Cited on page 350)
Abdurakhmon Sadiev, Marina Danilova, Eduard Gorbunov, Samuel Horv´ ath, Gauthier Gidel, Pavel Dvurechensky, Alexander Gasnikov, and Peter Richt´ arik. High-probability bounds for stochastic optimization and variational inequalities: the case of unbounded variance.arXiv preprin...
2023 arXiv
-
[238]
Abdurakhmon Sadiev, Grigory Malinovsky, Eduard Gorbunov, Igor Sokolov, Ahmed Khaled, Konstantin Burlachenko, and Peter Richt´ arik. Don’t compress gradients in random reshuffling: compress gradient dif- ferences.Advances in Neural Information Processing Systems, 37:84523– 8460...
2024
-
[239]
Improved convergence in parameter-agnostic error feedback through momentum.arXiv preprint arXiv:2511.14501, 2025.(Cited on page 38)
Abdurakhmon Sadiev, Yury Demidovich, Igor Sokolov, Grigory Mali- novsky, Sarit Khirirat, and Peter Richt´ arik. Improved convergence in parameter-agnostic error feedback through momentum.arXiv preprint arXiv:2511.14501, 2025.(Cited on page 38)
2025
-
[240]
Smoothness matrices beat smoothness constants: Better communication compression techniques for distributed optimization
Mher Safaryan, Filip Hanzely, and Peter Richt´ arik. Smoothness matrices beat smoothness constants: Better communication compression techniques for distributed optimization. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan, editors,Advances in Neur...
2021
-
[241]
FedNL: Making Newton-type methods applicable to federated learning
Mher Safaryan, Rustem Islamov, Xun Qian, and Peter Richt´ arik. FedNL: Making Newton-type methods applicable to federated learning. InIn- ternational Conference on Machine Learning, 2022. (arXiv preprint arXiv:2106.02969, 2021).(Cited on page 349)
2022 arXiv
-
[242]
Random shuffling beats SGD only after many epochs on ill-conditioned problems.Advances in Neural Information Processing Systems, 34:15151–15161, 2021.(Cited on pages 106 and 423)
Itay Safran and Ohad Shamir. Random shuffling beats SGD only after many epochs on ill-conditioned problems.Advances in Neural Information Processing Systems, 34:15151–15161, 2021.(Cited on pages 106 and 423)
2021
-
[243]
Dualize, split, randomize: fast nonsmooth optimization algorithms
Adil Salim, Laurent Condat, Konstantin Mishchenko, and Peter Richt´ arik. Dualize, split, randomize: fast nonsmooth optimization algorithms. In OPT2020: 12th Annual Workshop on Optimization for Machine Learning (NeurIPS 2020 Workshop), 2020.(Cited on pages 78 and 186)
2020
-
[244]
Optimal convergence rates for convex distributed optimization in networks.Journal of Machine Learning Research, 20:1–31, 2019.(Cited on page 77)
Kevin Scaman, Francis Bach, S´ ebastien Bubeck, Yin Lee, and Laurent Massouli´ e. Optimal convergence rates for convex distributed optimization in networks.Journal of Machine Learning Research, 20:1–31, 2019.(Cited on page 77)
2019
-
[245]
Minimizing finite sums with the stochastic average gradient.Mathematical Programming, 162(1): 83–112, 2017.(Cited on pages 30 and 349)
Mark Schmidt, Nicolas Le Roux, and Francis Bach. Minimizing finite sums with the stochastic average gradient.Mathematical Programming, 162(1): 83–112, 2017.(Cited on pages 30 and 349)
2017
-
[246]
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu. 1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns. InFifteenth Annual Conference of the International Speech Communication Association. Citeseer, 2014.(Cited on page 350)
2014
-
[247]
Cambridge University Press, 2014.(Cited on pages 43 and 44)
Shai Shalev-Shwartz and Shai Ben-David.Understanding machine learning: from theory to algorithms. Cambridge University Press, 2014.(Cited on pages 43 and 44)
2014
-
[248]
Convergence analysis of gradient descent stochastic algorithms.Journal of Optimization Theory and Applications, 91:439–454, 1996.(Cited on page 135)
A Shapiro and Y Wardi. Convergence analysis of gradient descent stochastic algorithms.Journal of Optimization Theory and Applications, 91:439–454, 1996.(Cited on page 135)
1996
-
[249]
Shifted compression framework: Gen- eralizations and improvements
Egor Shulgin and Peter Richt´ arik. Shifted compression framework: Gen- eralizations and improvements. InOPT2021: 13th Annual Workshop on Optimization for Machine Learning, 2021.(Cited on pages 71 and 209) 174
2021
-
[250]
First provable guarantees for practical private fl: Beyond restrictive assumptions
Egor Shulgin, Grigory Malinovsky, Sarit Khirirat, and Peter Richt´ arik. First provable guarantees for practical private fl: Beyond restrictive assumptions. arXiv preprint arXiv:2512.21521, 2025.(Cited on page 39)
2025
-
[251]
Smith, Benoit Dherin, David Barrett, and Soham De
Samuel L. Smith, Benoit Dherin, David Barrett, and Soham De. On the origin of implicit regularization in stochastic gradient descent. InInterna- tional Conference on Learning Representations, 2020.(Cited on page 101)
2020
-
[252]
Sebastian U. Stich. Unified optimal analysis of the (stochastic) gradient method.arXiv preprint arXiv:1907.04232, abs/1907.04232, 2019a. URL https://arXiv.org/abs/1907.04232.(Cited on page 106)
1907 arXiv
-
[253]
Sebastian U. Stich. Local SGD converges fast and communicates little. InInternational Conference on Learning Representations, 2019b.(Cited on pages 26, 29, and 45)
-
[254]
Sebastian U. Stich. On communication compression for distributed op- timization on heterogeneous data.arXiv preprint arXiv:2009.02388, abs/2009.02388, 2020. URLhttps://arXiv.org/abs/2009.02388.(Cited on page 105)
2009 arXiv
-
[255]
Stich and Sai Praneeth Karimireddy
Sebastian U. Stich and Sai Praneeth Karimireddy. The error-feedback framework: Better rates for SGD with delayed gradients and compressed communication.arXiv preprint arXiv:1909.05350, 2019.(Cited on page 91)
1909 arXiv
-
[256]
Sparsified SGD with memory
Sebastian U Stich, Jean-Baptiste Cordonnier, and Martin Jaggi. Sparsified SGD with memory. InAdvances in Neural Information Processing Systems, volume 31, 2018.(Cited on pages 26, 30, 105, 121, and 350)
2018
-
[257]
Fault-tolerant multi-agent optimization: opti- mal iterative distributed algorithms
Lili Su and Nitin H Vaidya. Fault-tolerant multi-agent optimization: opti- mal iterative distributed algorithms. In2016 ACM Symposium on Princi- ples of Distributed Computing, pages 425–434, 2016.(Cited on page 117)
2016
-
[258]
Optimization for deep learning: An overview.Journal of the Operations Research Society of China, 8(2):249–294, 2020.(Cited on pages 91 and 136)
Ruo-Yu Sun. Optimization for deep learning: An overview.Journal of the Operations Research Society of China, 8(2):249–294, 2020.(Cited on pages 91 and 136)
2020
-
[259]
Improving LoRA in privacy-preserving federated learning.arXiv preprint arXiv:2403.12313, 2024.(Cited on pages 133, 135, 136, and 143)
Youbang Sun, Zitao Li, Yaliang Li, and Bolin Ding. Improving LoRA in privacy-preserving federated learning.arXiv preprint arXiv:2403.12313, 2024.(Cited on pages 133, 135, 136, and 143)
2024 arXiv
-
[260]
Recent advances in LoRa: A comprehensive survey.ACM Transactions on Sensor Networks, 18(4):1–44, 2022.(Cited on page 132) 175
Zehua Sun, Huanqi Yang, Kai Liu, Zhimeng Yin, Zhenjiang Li, and Weitao Xu. Recent advances in LoRa: A comprehensive survey.ACM Transactions on Sensor Networks, 18(4):1–44, 2022.(Cited on page 132) 175
2022
-
[261]
Permutation com- pressors for provably faster distributed nonconvex optimization
Rafa l Szlendak, Alexander Tyurin, and Peter Richt´ arik. Permutation com- pressors for provably faster distributed nonconvex optimization. InInter- national Conference on Learning Representations, 2022.(Cited on page 353)
2022
-
[262]
Mini- batch primal and dual methods for SVMs
Martin Tak´ aˇ c, Avleen Bijral, Peter Richt´ arik, and Nathan Srebro. Mini- batch primal and dual methods for SVMs. InInternational Conference on Machine Learning, pages 537–552, 2013.(Cited on page 59)
2013
-
[263]
D 2: De- centralized training over decentralized data
Hanlin Tang, Xiangru Lian, Ming Yan, Ce Zhang, and Ji Liu. D 2: De- centralized training over decentralized data. InInternational Conference on Machine Learning, volume 80, pages 4848–4856. PMLR, 2018.(Cited on page 49)
2018
-
[264]
Communication-efficient distributed deep learning: a comprehensive sur- vey.arXiv preprint arXiv:2003.06307, abs/2003.06307, 2020
Zhenheng Tang, Shaohuai Shi, Xiaowen Chu, Wei Wang, and Bo Li. Communication-efficient distributed deep learning: a comprehensive sur- vey.arXiv preprint arXiv:2003.06307, abs/2003.06307, 2020. URLhttps: //arXiv.org/abs/2003.06307.(Cited on page 105)
2003 arXiv
-
[265]
Llama: Open and efficient foundation language models.arXiv preprint arXiv:2302.13971, 2023.(Cited on page 23)
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie- Anne Lachaux, Timoth´ ee Lacroix, Baptiste Rozi` ere, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models.arXiv preprint arXiv:2302.13971, 2023.(Cited on page 23)
2023 arXiv
-
[266]
Revisiting stochastic proximal point methods: Generalized smoothness and similarity.arXiv preprint arXiv:2502.03401, 2025.(Cited on page 39)
Zhirayr Tovmasyan, Grigory Malinovsky, Laurent Condat, and Peter Richt´ arik. Revisiting stochastic proximal point methods: Generalized smoothness and similarity.arXiv preprint arXiv:2502.03401, 2025.(Cited on page 39)
2025
-
[267]
Sharper rates and flexible framework for nonconvex SGD with client and data sampling.arXiv preprint arXiv:2206.02275, 2022.(Cited on page 75)
Alexander Tyurin, Lukang Sun, Konstantin Burlachenko, and Peter Richt´ arik. Sharper rates and flexible framework for nonconvex SGD with client and data sampling.arXiv preprint arXiv:2206.02275, 2022.(Cited on page 75)
2022 arXiv
-
[268]
Fast and faster conver- gence of SGD for over-parameterized models and an accelerated perceptron
Sharan Vaswani, Francis Bach, and Mark Schmidt. Fast and faster conver- gence of SGD for over-parameterized models and an accelerated perceptron. InInternational Conference on Artificial Intelligence and Statistics, pages 1195–1204, 2019.(Cited on page 354)
2019
-
[269]
Powersgd: Prac- tical low-rank gradient compression for distributed optimization
Thijs Vogels, Sai Praneeth Karimireddy, and Martin Jaggi. Powersgd: Prac- tical low-rank gradient compression for distributed optimization. InAd- vances in Neural Information Processing Systems, volume 32, 2019.(Cited on page 350) 176
2019
-
[270]
Stich, and Martin Jaggi
Thijs Vogels, Lie He, Anastasia Koloskova, Tao Lin, Sai Praneeth Karim- ireddy, Sebastian U. Stich, and Martin Jaggi. Relaysum for decentralized deep learning on heterogeneous data. InAdvances in Neural Information Processing Systems. Curran Associates, Inc., 2021.(Cited on page 49)
2021
-
[271]
Transfer learning with adaptive fine- tuning.IEEE Access, 8:196197–196211, 2020.(Cited on page 131)
Grega Vrbanˇ ciˇ c and Vili Podgorelec. Transfer learning with adaptive fine- tuning.IEEE Access, 8:196197–196211, 2020.(Cited on page 131)
2020
-
[272]
Glue: A multi-task benchmark and analysis platform for natu- ral language understanding.arXiv preprint arXiv:1804.07461, 2018.(Cited on pages 135 and 141)
Alex Wang. Glue: A multi-task benchmark and analysis platform for natu- ral language understanding.arXiv preprint arXiv:1804.07461, 2018.(Cited on pages 135 and 141)
2018 arXiv
-
[273]
Jianyu Wang, Zachary Charles, Zheng Xu, Gauri Joshi, H. Brendan McMa- han, Blaise Aguera y Arcas, Maruan Al-Shedivat, Galen Andrew, Salman Avestimehr, Katharine Daly, Deepesh Data, Suhas Diggavi, Hubert Eich- ner, Advait Gadhikar, Zachary Garrett, Antonious M. Girgis, Filip Ha...
2021 arXiv
-
[274]
Gra- dient sparsification for communication-efficient distributed optimiza- tion
Jianqiao Wangni, Jialei Wang, Ji Liu, and Tong Zhang. Gra- dient sparsification for communication-efficient distributed optimiza- tion. In S. Bengio, H. Wallach, H. Larochelle, K. Grau- man, N. Cesa-Bianchi, and R. Garnett, editors,Advances in Neu- ral Information Processing S...
2018
-
[275]
Terngrad: Ternary gradients to reduce communication in dis- tributed deep learning
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li. Terngrad: Ternary gradients to reduce communication in dis- tributed deep learning. InAdvances in Neural Information Processing Sys- tems, pages 1509–1519, 2017.(Cited on pages 30 and 349) 177
2017
-
[276]
Is local SGD better than minibatch SGD? InInternational Conference on Machine Learning, pages 10334–10343
Blake Woodworth, Kumar Kshitij Patel, Sebastian Stich, Zhen Dai, Brian Bullins, Brendan Mcmahan, Ohad Shamir, and Nathan Srebro. Is local SGD better than minibatch SGD? InInternational Conference on Machine Learning, pages 10334–10343. PMLR, 2020a.(Cited on pages 53, 268, and 436)
-
[277]
The min-max complexity of distributed stochastic convex optimization with intermittent communication.arXiv preprint arXiv:2102.01583, abs/2102.01583, 2021
Blake Woodworth, Brian Bullins, Ohad Shamir, and Nathan Srebro. The min-max complexity of distributed stochastic convex optimization with intermittent communication.arXiv preprint arXiv:2102.01583, abs/2102.01583, 2021. URLhttps://arXiv.org/abs/2102.01583.(Cited on pages 46 and 268)
2021 arXiv
-
[278]
Tight complexity bounds for optimizing composite objectives
Blake E Woodworth and Nati Srebro. Tight complexity bounds for optimizing composite objectives. In D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett, editors,Advances in Neu- ral Information Processing Systems, volume 29. Curran Associates, Inc., 2016. URLhttps://proce...
2016
-
[279]
Blake E Woodworth, Kumar Kshitij Patel, and Nati Srebro. Minibatch vs local sgd for heterogeneous distributed learning.Advances in Neural Information Processing Systems, 33:6281–6292, 2020b.(Cited on pages 26, 28, 29, 46, 62, 76, 91, 92, 93, 268, and 436)
-
[280]
Feder- ated variance-reduced stochastic gradient descent with robustness to byzan- tine attacks.IEEE Transactions on Signal Processing, 68:4583–4596, 2020
Zhaoxian Wu, Qing Ling, Tianyi Chen, and Georgios B Giannakis. Feder- ated variance-reduced stochastic gradient descent with robustness to byzan- tine attacks.IEEE Transactions on Signal Processing, 68:4583–4596, 2020. (Cited on pages 28, 117, 119, and 122)
2020
-
[281]
Chain of LoRA: Effi- cient fine-tuning of language models via residual learning.arXiv preprint arXiv:2401.04151, 2024.(Cited on pages 132, 133, 134, and 141)
Wenhan Xia, Chengwei Qin, and Elad Hazan. Chain of LoRA: Effi- cient fine-tuning of language models via residual learning.arXiv preprint arXiv:2401.04151, 2024.(Cited on pages 132, 133, 134, and 141)
2024 arXiv
-
[282]
Fall of empires: Break- ing byzantine-tolerant sgd by inner product manipulation
Cong Xie, Oluwasanmi Koyejo, and Indranil Gupta. Fall of empires: Break- ing byzantine-tolerant sgd by inner product manipulation. InConference on Uncertainty in Artificial Intelligence, pages 261–270, 2020a.(Cited on pages 117 and 118)
-
[283]
Zeno++: Robust fully asyn- chronous sgd
Cong Xie, Sanmi Koyejo, and Indranil Gupta. Zeno++: Robust fully asyn- chronous sgd. InInternational Conference on Machine Learning, pages 10495–10503. PMLR, 2020b.(Cited on page 350)
-
[284]
Parameter-efficient fine-tuning methods for pretrained language models: A 178 critical review and assessment.arXiv preprint arXiv:2312.12148, 2023
Lingling Xu, Haoran Xie, Si-Zhao Joe Qin, Xiaohui Tao, and Fu Lee Wang. Parameter-efficient fine-tuning methods for pretrained language models: A 178 critical review and assessment.arXiv preprint arXiv:2312.12148, 2023. (Cited on page 131)
2023 arXiv
-
[285]
Towards building a robust and fair federated learning system.arXiv preprint arXiv:2011.10464, 2020.(Cited on pages 27 and 349)
Xinyi Xu and Lingjuan Lyu. Towards building a robust and fair federated learning system.arXiv preprint arXiv:2011.10464, 2020.(Cited on pages 27 and 349)
2011 arXiv
-
[286]
Buffered asynchronous sgd for byzantine learning.Journal of Machine Learning Research, 24(204):1–62, 2023.(Cited on page 351)
Yi-Rui Yang and Wu-Jun Li. Buffered asynchronous sgd for byzantine learning.Journal of Machine Learning Research, 24(204):1–62, 2023.(Cited on page 351)
2023
-
[287]
Federated learning for 6g: Applications, challenges, and opportunities.Engineering, 8:33–41, 2022.(Cited on page 104)
Zhaohui Yang, Mingzhe Chen, Kai-Kit Wong, H Vincent Poor, and Shuguang Cui. Federated learning for 6g: Applications, challenges, and opportunities.Engineering, 8:33–41, 2022.(Cited on page 104)
2022
-
[288]
Byzantine-robust distributed learning: Towards optimal statistical rates
Dong Yin, Yudong Chen, Ramchandran Kannan, and Peter Bartlett. Byzantine-robust distributed learning: Towards optimal statistical rates. InInternational Conference on Machine Learning, pages 5650–5659, 2018. (Cited on pages 28, 118, and 120)
2018
-
[289]
Large batch optimization for deep learning: Training bert in 76 minutes.arXiv preprint arXiv:1904.00962, 2019.(Cited on page 104)
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh. Large batch optimization for deep learning: Training bert in 76 minutes.arXiv preprint arXiv:1904.00962, 2019.(Cited on page 104)
1904 arXiv
-
[290]
On the linear speedup analysis of commu- nication efficient momentum SGD for distributed non-convex optimization
Hao Yu, Rong Jin, and Sen Yang. On the linear speedup analysis of commu- nication efficient momentum SGD for distributed non-convex optimization. InInternational Conference on Machine Learning, 2019.(Cited on page 61)
2019
-
[291]
Federated accelerated stochastic gradient descent
Honglin Yuan and Tengyu Ma. Federated accelerated stochastic gradient descent. In H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, editors,Advances in Neural Information Processing Systems, volume 33, pages 5332–5344. Curran Associates, Inc., 2020.(Cited on page 53)
2020
-
[292]
Alghunaim
Kun Yuan and Sulaiman A. Alghunaim. Removing data heterogeneity in- fluence enhances network topology dependence of decentralized SGD.arXiv preprint arXiv:2105.08023, 2021.(Cited on page 49)
2021 arXiv
-
[293]
Minibatch vs local SGD with shuffling: Tight convergence bounds and beyond.arXiv preprint arXiv:2110.10342, 2021.(Cited on pages 91, 268, and 423) 179
Chulhee Yun, Shashank Rajput, and Suvrit Sra. Minibatch vs local SGD with shuffling: Tight convergence bounds and beyond.arXiv preprint arXiv:2110.10342, 2021.(Cited on pages 91, 268, and 423) 179
2021 arXiv
-
[294]
Parallel SGD: When does averaging help?arXiv preprint arXiv:1606.07365, 2016.(Cited on page 45)
Jian Zhang, Christopher De Sa, Ioannis Mitliagkas, and Christopher R´ e. Parallel SGD: When does averaging help?arXiv preprint arXiv:1606.07365, 2016.(Cited on page 45)
2016 arXiv
-
[295]
Why gradi- ent clipping accelerates training: A theoretical justification for adaptivity
Jingzhao Zhang, Tianxing He, Suvrit Sra, and Ali Jadbabaie. Why gradi- ent clipping accelerates training: A theoretical justification for adaptivity. InInternational Conference on Learning Representations, 2020a. (arXiv preprint arXiv:1905.11881, 2019).(Cited on page 350)
1905 arXiv
-
[296]
Why are adaptive methods good for attention models? InAdvances in Neural Information Processing Systems, volume 33, pages 15383–15393, 2020b.(Cited on pages 350 and 395)
Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank Reddi, Sanjiv Kumar, and Suvrit Sra. Why are adaptive methods good for attention models? InAdvances in Neural Information Processing Systems, volume 33, pages 15383–15393, 2020b.(Cited on pages 350 and 395)
-
[297]
Regularization and variable selection via the elastic net.Journal of the Royal Statistical Society B, 67:301–320, 2005
Hui Zhou and Trevor Hastie. Regularization and variable selection via the elastic net.Journal of the Royal Statistical Society B, 67:301–320, 2005. (Cited on page 43)
2005
-
[298]
On the fenchel duality between strong convexity and lips- chitz continuous gradient.arXiv preprint arXiv:1803.06573, 2018.(Cited on page 136)
Xingyu Zhou. On the fenchel duality between strong convexity and lips- chitz continuous gradient.arXiv preprint arXiv:1803.06573, 2018.(Cited on page 136)
2018 arXiv
-
[299]
Broadcast: Reducing both stochastic and compression noise to robustify communication-efficient federated learning
Heng Zhu and Qing Ling. Broadcast: Reducing both stochastic and compression noise to robustify communication-efficient federated learning. arXiv preprint arXiv:2104.06685, 2021.(Cited on pages 117 and 350)
2021 arXiv
-
[300]
Asymmetry in low-rank adapters of foundation models.arXiv preprint arXiv:2402.16842, 2024.(Cited on pages 132, 134, and 141)
Jiacheng Zhu, Kristjan Greenewald, Kimia Nadjahi, Haitz S´ aez de Oc´ ariz Borde, Rickard Br¨ uel Gabrielsson, Leshem Choshen, Marzyeh Ghassemi, Mikhail Yurochkin, and Justin Solomon. Asymmetry in low-rank adapters of foundation models.arXiv preprint arXiv:2402.16842, 2024.(Ci...
2024 arXiv
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.