Pith. sign in

REVIEW 6 cited by

Mitigating Data Heterogeneity in Federated Learning with Data Augmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.09979 v1 pith:OCAGA7TG submitted 2022-06-20 cs.LG

classification cs.LG
keywords datafederatedaugmentationheterogeneityclientdomainlearningperformance
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Federated Learning (FL) is a prominent framework that enables training a centralized model while securing user privacy by fusing local, decentralized models. In this setting, one major obstacle is data heterogeneity, i.e., each client having non-identically and independently distributed (non-IID) data. This is analogous to the context of Domain Generalization (DG), where each client can be treated as a different domain. However, while many approaches in DG tackle data heterogeneity from the algorithmic perspective, recent evidence suggests that data augmentation can induce equal or greater performance. Motivated by this connection, we present federated versions of popular DG algorithms, and show that by applying appropriate data augmentation, we can mitigate data heterogeneity in the federated setting, and obtain higher accuracy on unseen clients. Equipped with data augmentation, we can achieve state-of-the-art performance using even the most basic Federated Averaging algorithm, with much sparser communication.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Non-IID data in Federated Learning: A Survey with Taxonomy, Metrics, Methods, Frameworks and Future Directions

    cs.LG 2024-11 conditional novelty 6.0 of 10

    A comprehensive survey that organizes non-IID data in federated learning into taxonomies of skew types, partition protocols, and metrics, with a meta-analysis of 235 selected papers.

  2. Topology-Aware Knowledge Propagation in Decentralized Learning

    cs.LG 2025-05 conditional novelty 5.0 of 10

    Topology-aware aggregation using degree or betweenness centrality improves out-of-distribution knowledge propagation in decentralized learning across 36 simulated topologies and five datasets.

  3. Federated Deconfounding and Debiasing Learning for Out-of-Distribution Generalization

    cs.CV 2025-05 conditional novelty 5.0 of 10

    FedDDL improves federated out-of-distribution generalization by generating background-mixed counterfactual samples and aligning clients with causal prototypes, yielding an average Top-1 gain of about 4.5 percent over ...

  4. FedPhD: Federated Pruning with Hierarchical Learning of Diffusion Models

    cs.LG 2025-07 conditional novelty 4.0 of 10

    A hierarchical federated learning method with distribution-aware aggregation and structured pruning trains diffusion models under non-IID data with lower communication cost.

  5. Addressing Label Shift in Distributed Learning via Entropy Regularization

    cs.LG 2025-02 conditional novelty 4.0 of 10

    VRLS uses entropy regularization of the predictor to improve test-to-train label density ratio estimation, and extends it to multi-node IW-ERM for distributed label shift.

  6. Fed-AugMix: Balancing Privacy and Utility via Data Augmentation

    cs.CR 2024-12 conditional novelty 4.0 of 10

    Fed-AugMix applies AugMix data augmentation with a Jensen-Shannon consistency loss at federated clients, empirically degrading gradient-inversion reconstruction quality while preserving or improving model accuracy.

Pith tools