Pith. sign in

REVIEW 12 cited by

Mime: Mimicking Centralized Stochastic Algorithms in Federated Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2008.03606 v2 pith:M5Y2QUPP submitted 2020-08-08 cs.LG cs.DCmath.OCstat.ML

classification cs.LGcs.DCmath.OCstat.ML
keywords centralizedmimefederatedsettinglearningmomentumalgorithmalgorithms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Federated learning (FL) is a challenging setting for optimization due to the heterogeneity of the data across different clients which gives rise to the client drift phenomenon. In fact, obtaining an algorithm for FL which is uniformly better than simple centralized training has been a major open problem thus far. In this work, we propose a general algorithmic framework, Mime, which i) mitigates client drift and ii) adapts arbitrary centralized optimization algorithms such as momentum and Adam to the cross-device federated learning setting. Mime uses a combination of control-variates and server-level statistics (e.g. momentum) at every client-update step to ensure that each local update mimics that of the centralized method run on iid data. We prove a reduction result showing that Mime can translate the convergence of a generic algorithm in the centralized setting into convergence in the federated setting. Further, we show that when combined with momentum based variance reduction, Mime is provably faster than any centralized method--the first such result. We also perform a thorough experimental exploration of Mime's performance on real world datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. What's in a Smoothness Constant? Tighter Rates for Local SGD with Bounded Second-order Heterogeneity

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Local SGD provably improves over Mini-batch SGD under bounded second-order heterogeneity in the general convex setting, with nearly tight upper and lower bounds.

  2. FedACT: Federated Adaptive Coordinate Trust Modulation for Robust Transformer Training under Data Heterogeneity

    cs.LG 2026-07 conditional novelty 6.5 of 10

    Global-aware coordinate trust modulation after corrected AdamW updates improves federated Transformer and LLM training under data heterogeneity over strong adaptive baselines.

  3. LoRDO: Distributed Low-Rank Optimization with Infrequent Communication

    cs.LG 2026-02 conditional novelty 6.0 of 10

    LoRDO combines global low-rank projections with full-rank quasi-hyperbolic momentum to let infrequent-synchronization distributed training match low-rank DDP at roughly 10x less communication.

  4. Divergence-Based Adaptive Aggregation for Byzantine Robust Federated Learning

    cs.DC 2026-01 reject novelty 6.0 of 10

    The paper proposes divergence-based update calibration (DRAG/BR-DRAG) with convergence theorems, but DRAG's theorem excludes the hyperparameter settings used in its own experiments.

  5. Delayed Momentum Aggregation: Communication-efficient Byzantine-robust Federated Learning with Partial Participation

    cs.LG 2025-09 conditional novelty 6.0 of 10

    D-Byz-SGDM aggregates cached momentum from non-sampled clients together with fresh momentum from sampled clients, preserving Byzantine robustness under partial participation and achieving an optimal O(cδζ²/p) stationa...

  6. Enhancing Privacy in Decentralized Min-Max Optimization: A Differentially Private Approach

    cs.LG 2025-08 reject novelty 6.0 of 10

    DPMixSGD injects calibrated Gaussian noise into local gradient estimates to make decentralized nonconvex-strongly-concave min-max optimization differentially private, while claiming to preserve the STORM convergence rate.

  7. Federated Learning on Riemannian Manifolds: A Gradient-Free Projection-Based Approach

    math.OC 2025-07 conditional novelty 6.0 of 10

    A projection-based zeroth-order federated learning algorithm on Riemannian manifolds achieves sublinear convergence with linear speedup, using only Euclidean random perturbations.

  8. FedWCM: Unleashing the Potential of Momentum-based Federated Learning in Long-Tailed Scenarios

    cs.LG 2025-07 reject novelty 6.0 of 10

    FedWCM uses per-client data-distribution scores to adapt momentum and aggregation weights in federated learning, showing empirical gains over FedAvg and FedCM on long-tailed non-IID datasets, but its convergence proof...

  9. Tackling Heterogeneity in Federated Learning via Variance-Reduced Boltzmann Sampling within Homogeneous Social Coalitions

    cs.LG 2025-06 reject novelty 5.0 of 10

    A variance-reduction-based client selection with coalition clustering yields modest accuracy gains over baselines in heterogeneous federated learning, but its convergence guarantee rests on an assumption that the poli...

  10. Generalizable Federated Learning using Client Adaptive Focal Modulation

    cs.CV 2025-08 reject novelty 4.0 of 10

    The abstract describes AdaptFED, a claimed federated learning method, but the full text is an unrelated graph theory paper, so the claimed results are absent from the submission.

  11. What Makes Local Updates Effective: The Role of Data Heterogeneity and Smoothness

    cs.LG 2025-06 conditional novelty 4.0 of 10

    Under bounded second-order heterogeneity, local updates are shown to achieve faster convergence than mini-batch SGD in several convex and non-convex regimes, with matching lower bounds.

  12. Strategies for Improving Communication Efficiency in Distributed and Federated Learning: Compression, Local Training, and Personalization

    cs.LG 2025-09 conditional novelty 3.0 of 10

    A PhD dissertation showing unified compression theory, personalized accelerated local training, and pruning methods that reduce communication costs in federated learning and maintain accuracy in LLM pruning.

Pith tools