Pith. sign in

REVIEW 3 cited by

FedMLLM: Federated Fine-tuning MLLM on Multimodal Heterogeneity Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.14717 v2 pith:KO4DSJH3 submitted 2024-11-22 cs.LG cs.CLcs.CV

classification cs.LGcs.CLcs.CV
keywords multimodaldatamllmsacrossfederatedfine-tuningheterogeneitybenchmark
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multimodal Large Language Models (MLLMs) have made significant advancements, demonstrating powerful capabilities in processing and understanding multimodal data. Fine-tuning MLLMs with Federated Learning (FL) allows for expanding the training data scope by including private data sources, thereby enhancing their practical applicability in privacy-sensitive domains. However, current research remains in the early stage, particularly in addressing the \textbf{multimodal heterogeneities} in real-world applications. In this paper, we introduce a benchmark to evaluate the performance of federated fine-tuning of MLLMs across various multimodal heterogeneous scenarios, laying the groundwork for future research in the field. Our benchmark includes two lightweight MLLMs, two downstream tasks, three evaluation metrics, and five datasets across three domains, along with six comparison baselines, covering over ten types of modality heterogeneities across four multimodal scenarios. To address the challenges posed by multimodal heterogeneity, we develop a general FedMLLM framework that integrates classic FL methods alongside two modality-agnostic strategies. Extensive experimental results show that our proposed FL paradigm improves the performance of MLLMs by broadening the range of training data and mitigating multimodal heterogeneity. Code is available in supplementary materials.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FedVLMBench: Benchmarking Federated Fine-Tuning of Vision-Language Models

    cs.LG 2025-06 conditional novelty 6.0 of 10

    FedVLMBench systematically benchmarks federated fine-tuning of vision-language models and finds that a 2-layer MLP connector with joint connector-LLM training is optimal for encoder-based models, while vision-centric ...

  2. FedNano: Toward Lightweight Federated Tuning for Pretrained Multimodal Large Language Models

    cs.LG 2025-06 reject novelty 5.0 of 10

    FedNano centralizes the frozen LLM on the server, trains lightweight NanoAdapters on clients, and reports higher federated VQA accuracy than FedAvg, FedProx, and FedDPA-F on ScienceQA and IconQA.

  3. Filter-And-Refine: A MLLM Based Cascade System for Industrial-Scale Video Content Moderation

    cs.LG 2025-07 conditional novelty 4.0 of 10

    A cascade of an embedding router and a fine-tuned multimodal LLM ranker is claimed to improve content moderation F1 by 66.5% while using 1.5% of the compute of direct LLM deployment.

Pith tools