REVIEW 13 cited by
A Field Guide to Federated Optimization
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Federated learning and analytics are a distributed approach for collaboratively learning models (or statistics) from decentralized data, motivated by and designed for privacy protection. The distributed learning process can be formulated as solving federated optimization problems, which emphasize communication efficiency, data heterogeneity, compatibility with privacy and system requirements, and other constraints that are not primary considerations in other problem settings. This paper provides recommendations and guidelines on formulating, designing, evaluating and analyzing federated optimization algorithms through concrete examples and practical implementation, with a focus on conducting effective simulations to infer real-world performance. The goal of this work is not to survey the current literature, but to inspire researchers and practitioners to design federated learning algorithms that can be used in various practical applications.
Forward citations
Cited by 13 Pith papers
-
Tessera: Unlocking Heterogeneous GPUs through Kernel-Granularity Disaggregation
Tessera performs kernel-granularity disaggregation on heterogeneous GPUs, achieving up to 2.3x throughput and 1.6x cost efficiency gains for large model inference while generalizing beyond prior methods.
-
What's in a Smoothness Constant? Tighter Rates for Local SGD with Bounded Second-order Heterogeneity
Local SGD provably improves over Mini-batch SGD under bounded second-order heterogeneity in the general convex setting, with nearly tight upper and lower bounds.
-
Efficient Distributed Optimization under Heavy-Tailed Noise
Coordinate-wise two-sided clipping at inner and outer optimizers (Bi2Clip) achieves provable convergence under heavy-tailed noise with unbounded variance, while needing no preconditioner memory.
-
How Context Attribution Handles What the Model Already Knows
Context attribution methods cannot disentangle in-context from in-weight knowledge and assign unfaithful scores under overlap; new metrics and WMDP-Cyber++ quantify the failure.
-
One-Shot Clustering for Federated Learning Under Clustering-Agnostic Assumption
OCFL automatically picks the clustering round by detecting a rise in the p-norm of the pairwise cosine-distance matrix of client gradients, and with density-based clustering it recovers client cohorts earlier and more...
-
Federated Majorize-Minimization: Beyond Parameter Aggregation
By averaging surrogate-function parameters across clients and then minimizing the aggregated surrogate on the server, federated learning can converge under heterogeneity where parameter averaging diverges.
-
Beyond Communication Overhead: A Multilevel Monte Carlo Approach for Mitigating Compression Bias in Distributed Learning
A multilevel Monte Carlo framework debiases biased gradient compressors, preserving SGD convergence guarantees while reducing communication cost, with adaptive variance-minimizing level selection.
-
DES-LOC: Desynced Low Communication Adaptive Optimizers for Training Foundation Models
DES-LOC synchronizes model parameters and Adam/ADOPT momentum states on separate schedules, matching Local Adam quality with about 2x less communication and 170x less than DDP in tests up to 1.7B parameters.
-
FedMHO: Heterogeneous One-Shot Federated Learning Towards Resource-Constrained Edge Devices
FedMHO is a hybrid one-shot federated learning framework where resource-sufficient clients contribute deep classifiers and resource-constrained clients contribute lightweight generative models, fused on the server int...
-
Aequa: Fair Model Rewards in Collaborative Learning via Slimmable Networks
Aequa allocates model widths (and thus accuracies) to federated learning participants in proportion to their estimated contributions, using slimmable networks and a simulated annealing optimizer.
-
What Makes Local Updates Effective: The Role of Data Heterogeneity and Smoothness
Under bounded second-order heterogeneity, local updates are shown to achieve faster convergence than mini-batch SGD in several convex and non-convex regimes, with matching lower bounds.
-
Strategies for Improving Communication Efficiency in Distributed and Federated Learning: Compression, Local Training, and Personalization
A PhD dissertation showing unified compression theory, personalized accelerated local training, and pruning methods that reduce communication costs in federated learning and maintain accuracy in LLM pruning.
-
How to DP-fy Your Data: A Practical Guide to Generating Synthetic Data With Differential Privacy
A practical, extremely thorough survey of differentially private synthetic data generation: methods, privacy units, evaluation metrics, and end-to-end system components across four data modalities.
Discussion (0). Continue with ORCID to comment.