Pith. sign in

REVIEW 6 cited by

Multimodal Federated Learning via Contrastive Representation Ensemble

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.08888 v3 pith:EXQAJIT6 submitted 2023-02-17 cs.LG cs.AI

classification cs.LGcs.AI
keywords multimodalmodeldataclientslearningmodalityensemblefederated
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the increasing amount of multimedia data on modern mobile systems and IoT infrastructures, harnessing these rich multimodal data without breaching user privacy becomes a critical issue. Federated learning (FL) serves as a privacy-conscious alternative to centralized machine learning. However, existing FL methods extended to multimodal data all rely on model aggregation on single modality level, which restrains the server and clients to have identical model architecture for each modality. This limits the global model in terms of both model complexity and data capacity, not to mention task diversity. In this work, we propose Contrastive Representation Ensemble and Aggregation for Multimodal FL (CreamFL), a multimodal federated learning framework that enables training larger server models from clients with heterogeneous model architectures and data modalities, while only communicating knowledge on public dataset. To achieve better multimodal representation fusion, we design a global-local cross-modal ensemble strategy to aggregate client representations. To mitigate local model drift caused by two unprecedented heterogeneous factors stemming from multimodal discrepancy (modality gap and task gap), we further propose two inter-modal and intra-modal contrasts to regularize local training, which complements information of the absent modality for uni-modal clients and regularizes local clients to head towards global consensus. Thorough evaluations and ablation studies on image-text retrieval and visual question answering tasks showcase the superiority of CreamFL over state-of-the-art FL methods and its practical value.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FedTaste: Topology-Aware Structural Transfer for Multimodal Federated Learning with Missing Modalities

    cs.MM 2026-07 reject novelty 6.0 of 10

    FedTaste aligns missing-modality clients to a server-built semantic graph distilled from full-modality CLIP clients by updating only lightweight prompts, reporting state-of-the-art federated retrieval numbers.

  2. Multimodal Federated Learning With Missing Modalities through Feature Imputation Network

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A federated feature imputation network that synthesizes missing modality bottleneck features improves multimodal federated learning accuracy over naive and generative baselines.

  3. Sheaf-Based Decentralized Multimodal Learning for Next-Generation Wireless Communication Systems

    cs.LG 2025-06 reject novelty 5.0 of 10

    A sheaf-Laplacian regularized decentralized multimodal federated learning algorithm with local attention is proposed and claimed to converge to a stationary point, with reported gains on two wireless tasks.

  4. FedNano: Toward Lightweight Federated Tuning for Pretrained Multimodal Large Language Models

    cs.LG 2025-06 reject novelty 5.0 of 10

    FedNano centralizes the frozen LLM on the server, trains lightweight NanoAdapters on clients, and reports higher federated VQA accuracy than FedAvg, FedProx, and FedDPA-F on ScienceQA and IconQA.

  5. Multimodal Federated Learning: A Survey through the Lens of Different FL Paradigms

    cs.LG 2025-05 conditional novelty 5.0 of 10

    A paradigm-based taxonomy of multimodal federated learning that assigns each branch a headline challenge: modality heterogeneity (horizontal), privacy leakage (vertical), and efficiency (hybrid).

  6. DRAGD: A Federated Unlearning Data Reconstruction Attack Based on Gradient Differences

    cs.LG 2025-07 reject novelty 3.0 of 10

    DRAGD and DRAGDP reconstruct erased federated-learning images by sequentially matching post-unlearning and pre-unlearning gradients, with DRAGDP adding a public-image prior.

Pith tools