REVIEW 6 cited by
Multimodal Federated Learning via Contrastive Representation Ensemble
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
With the increasing amount of multimedia data on modern mobile systems and IoT infrastructures, harnessing these rich multimodal data without breaching user privacy becomes a critical issue. Federated learning (FL) serves as a privacy-conscious alternative to centralized machine learning. However, existing FL methods extended to multimodal data all rely on model aggregation on single modality level, which restrains the server and clients to have identical model architecture for each modality. This limits the global model in terms of both model complexity and data capacity, not to mention task diversity. In this work, we propose Contrastive Representation Ensemble and Aggregation for Multimodal FL (CreamFL), a multimodal federated learning framework that enables training larger server models from clients with heterogeneous model architectures and data modalities, while only communicating knowledge on public dataset. To achieve better multimodal representation fusion, we design a global-local cross-modal ensemble strategy to aggregate client representations. To mitigate local model drift caused by two unprecedented heterogeneous factors stemming from multimodal discrepancy (modality gap and task gap), we further propose two inter-modal and intra-modal contrasts to regularize local training, which complements information of the absent modality for uni-modal clients and regularizes local clients to head towards global consensus. Thorough evaluations and ablation studies on image-text retrieval and visual question answering tasks showcase the superiority of CreamFL over state-of-the-art FL methods and its practical value.
Forward citations
Cited by 6 Pith papers
-
FedTaste: Topology-Aware Structural Transfer for Multimodal Federated Learning with Missing Modalities
FedTaste aligns missing-modality clients to a server-built semantic graph distilled from full-modality CLIP clients by updating only lightweight prompts, reporting state-of-the-art federated retrieval numbers.
-
Multimodal Federated Learning With Missing Modalities through Feature Imputation Network
A federated feature imputation network that synthesizes missing modality bottleneck features improves multimodal federated learning accuracy over naive and generative baselines.
-
Sheaf-Based Decentralized Multimodal Learning for Next-Generation Wireless Communication Systems
A sheaf-Laplacian regularized decentralized multimodal federated learning algorithm with local attention is proposed and claimed to converge to a stationary point, with reported gains on two wireless tasks.
-
FedNano: Toward Lightweight Federated Tuning for Pretrained Multimodal Large Language Models
FedNano centralizes the frozen LLM on the server, trains lightweight NanoAdapters on clients, and reports higher federated VQA accuracy than FedAvg, FedProx, and FedDPA-F on ScienceQA and IconQA.
-
Multimodal Federated Learning: A Survey through the Lens of Different FL Paradigms
A paradigm-based taxonomy of multimodal federated learning that assigns each branch a headline challenge: modality heterogeneity (horizontal), privacy leakage (vertical), and efficiency (hybrid).
-
DRAGD: A Federated Unlearning Data Reconstruction Attack Based on Gradient Differences
DRAGD and DRAGDP reconstruct erased federated-learning images by sequentially matching post-unlearning and pre-unlearning gradients, with DRAGDP adding a public-image prior.
Discussion (0). Continue with ORCID to comment.