REVIEW 4 major objections 6 minor 52 references
Pilot: Building the Federated Multimodal Instruction Tuning Framework
T0 review · 4 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Pilot lets devices with different visual tasks share knowledge without sharing data.
desk verdict Pilot defines a genuinely new federated multimodal instruction-tuning task and a sensible two-stage adapter architecture, but the main aggregation mechanism is unvalidated and the experiments are too thin to support the strong claims. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing components are (1) the two-stage 'adapter-on-adapter' connector: stage one's task-specific adapter $\psi_t$ and client-specific adapter $\psi_s$ with difference loss $L_d = \|\mathbf{x}_t^\top \mathbf{x}_s\|_F^2$, and stage two's Cross-task Mixture-of-Adapters (CT-MoA), where each non-local task-specific adapter has a cross-task adapter $\psi^c$ initialized from the local adapter, and a router $\phi$ with load-balancing loss $L_b$ and router z-loss $L_z$; and (2) the adaptive text-adapter aggregation, which computes Euclidean distances $d_{k,i}$ between LoRA text adapters, keeps the Top-M closest, and weights them by inverse distance. The CT-MoA lets a client activate knowledge from other tasks, while the adaptive aggregation prevents negative transfer by not averaging incompatible text adapters.
What would settle it
Run Pilot on the same two scenarios but replace the Euclidean-distance selection of the Top-M text adapters with a random selection of M clients; if the random selection matches the distance-based selection in average task accuracy, the distance assumption is not driving the reported gains.
Extended reading notes
Core claim
Pilot solves the FedMIT task, defined as collaboratively instruction-tuning an MLLM on K clients whose data cover T different multimodal task types, by making the visual connector itself the site of cross-task communication. The connector is trained in two stages: stage one uses a task-specific adapter and a client-specific adapter with a soft subspace orthogonality loss, so that shared task knowledge and private data patterns do not collide; stage two assembles a Cross-task Mixture-of-Adapters (CT-MoA) with one adapter per task, adds a cross-task adapter on top of each non-local task adapter to smooth the heterogeneity gap, and uses a router with load-balancing and z-losses. On the text side, each client trains a LoRA adapter, and the server aggregates every client's text adapter by selecting the M nearest peers in Euclidean parameter distance and inverse-distance weighting, instead of averaging all clients. Evaluated with LLaVA 1.5 on two scenarios (GQA/COCO/RefCOCO and ScienceQA/GQA/OCRVQA), Pilot outperforms FedAvg, FedProx, FedAdam, FedDPA, and Shepherd, and surpasses local training, which the ablations show is mainly due to the adaptive aggregation and the CT-MoA cross-task adapters.
Load-bearing premise
The method assumes that if two clients' text adapter parameters are close points in Euclidean parameter space, then merging their adapters will help both, and that the closest M are the best partners; the paper does not prove this correlation and tests the selection size M in only one scenario.
Editorial extensions
If this is right
- Federated multimodal instruction tuning can work across different visual task types without centralized data collection, as long as the connector is made the locus of cross-task adaptation.
- The two-stage connector gives a concrete recipe for separating personalized from shared visual knowledge under task heterogeneity.
- Adaptive, distance-based aggregation of text adapters provides a way to reduce negative transfer in heterogeneous federated fine-tuning.
- Pilot's state-of-the-art results in both evaluated scenarios suggest that task heterogeneity is not an obstacle to collaborative MLLM tuning, provided the right architecture is used.
Reading between the lines
- The distance-based aggregation rule is a heuristic that could be probed further: if parameter distance does not track merge compatibility, the method would be selecting partners almost arbitrarily, so a direct study of when parameter distance predicts successful merging would settle the mechanism.
- The CT-MoA design assumes one task per client; extending it to clients holding multiple tasks would require routing at the sample level rather than the client level.
- Because all task-specific adapters are broadcast to every client, communication cost grows linearly with the number of tasks; hierarchical or clustered aggregation could keep the cross-task benefit while scaling to more tasks.
- The same connector-level two-stage adapter idea could be tested on other modality pairs, such as audio-text or video-text, in federated settings.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces the Federated Multimodal Instruction Tuning (FedMIT) task and proposes Pilot, a framework for collaboratively fine-tuning multimodal large language models (MLLMs) on heterogeneous visual instruction tasks across clients. Pilot integrates a two-stage 'adapter-on-adapter' design into the vision-LLM connector: stage 1 trains a task-specific adapter and a client-specific adapter with an orthogonality loss, and stage 2 builds a Cross-task Mixture-of-Adapters (CT-MoA) module with cross-task adapters and a router. On the server side, task-specific visual adapters are aggregated per task, while text adapters (LoRA) are aggregated via an adaptive scheme that selects the Top-M closest clients by Euclidean distance (Eqs. 11-12). Experiments on two 9-client, 3-task scenarios (GQA/COCO/RefCOCO and ScienceQA/GQA/OCRVQA) report that Pilot outperforms FedAvg, FedProx, FedAdam, Shepherd, and FedDPA, with ablations supporting the contribution of each module.
Significance. If the reported results hold, the paper makes a useful contribution by identifying and formalizing the FedMIT problem and by demonstrating that a mix of task-specific adapters, cross-task MoA, and distance-based selective aggregation can mitigate task heterogeneity in federated MLLM instruction tuning. The two-stage connector design is intuitive, and the empirical gains over strong baselines (e.g., Tables 1-2) are consistent across both scenarios. The paper also provides a new evaluation scenario for cross-task federated learning with multimodal data. However, the novelty is incremental in its components (adapter-on-adapter, MoA, distance-weighted aggregation are known ideas), and the central claims rest on an experimental validation whose statistical robustness is not established.
major comments (4)
- [Experimental Results; Ablation Studies] The results in Tables 1-5 are reported without error bars, repeated seeds, or statistical significance tests. The improvements over the best baselines are often small (e.g., GQA Client 1: Pilot 51.4 vs FedAdam 50.8 in Table 1; Pilot 49.2 vs FedDPA 47.3 in Table 2). Given a single random data split and no variance information, the claim that Pilot 'achieves state-of-the-art results' is not yet supported. I recommend reporting mean and standard deviation over at least three seeds and, if feasible, a paired significance test across clients.
- [Adaptive Text-adapter Aggregation, Eqs. 11-12] The central novelty of the aggregation strategy is the assumption that Euclidean distance in LoRA parameter space identifies clients whose text-adapters are beneficial to merge. The paper offers no theoretical justification, no analysis of the parameter geometry, and no direct validation of this proxy. The ablation in Table 3 shows that removing ATA degrades COCO from 124.0 to 114.5 and RefCOCO from 51.0 to 47.5, so the method's main gain relies on this assumption. Table 5 compares M values and same/all-client aggregation, but it does not include a control such as random Top-M selection, cosine similarity, or an oracle that selects partners by measured task complementarity. Without such a control, the observed gains could be due to selecting any subset of clients rather than to the distance criterion.
- [Implementation Details; Further Remarks, Table 5] The hyperparameter Top-M is set to 6 in Implementation Details and is then evaluated at M=5,6,7 in Table 5, apparently on the same test sets used for the final results. There is no held-out validation split for hyperparameter selection, so the reported numbers may be optimistically biased. Similarly, the coefficients lambda0, lambda1, lambda2 and other hyperparameters are fixed without sensitivity analysis. The paper should either use a validation split or explicitly acknowledge that the reported configuration was selected on the test data, and provide robustness analysis for the key hyperparameters.
- [Introduction; Methodology; Conclusion] The abstract and conclusion state that Pilot learns 'without being affected by the task heterogeneity during instruction tuning.' This is too strong a claim given the experimental scope: only 9 clients and 3 tasks per scenario, 3 communication rounds, and no evaluation under different numbers of clients, tasks, or partial participation. The method deliberately preserves task-specific adapters, so it does not remove heterogeneity effects; it mitigates them. I recommend softening the claim and adding discussion of conditions under which the method might fail (e.g., highly imbalanced data, non-IID label distributions, or more tasks than adapters).
minor comments (6)
- [Abstract and Introduction] The phrase 'without being affected by the task heterogeneity' is repeated several times and overstates the mitigation; consider rewording to 'reduces the impact of task heterogeneity' or 'robust to task heterogeneity in the tested settings.'
- [Experiment, Baselines] In Table 1 and Table 2, the baseline is spelled 'FedA VG' and 'FedA VG' with a space and small caps; elsewhere it is 'FedAvg.' Please unify the notation.
- [Experimental Results] The evaluation metric is 'loU' (IoU) in the text; use 'IoU' consistently.
- [Equation (2)] The total loss formula has an awkward fraction layout: the inner sum over n_k and the outer weighting by n_k/n make the expression redundant; simplify it to a weighted sum over clients.
- [Stage 1: Task-specific Feature Mining] The orthogonality constraint in Eq. (3) is applied to the output features x_t and x_s, but the text says the client-specific adapter produces 'more refined' features. The connection between the Frobenius-norm condition and the semantic goal is not explained; please clarify what property of the features is being enforced.
- [Stage 2: Cross-task Visual Interaction] The acronym is inconsistently written as CT-MOA and CT-MoA; choose one and use it throughout.
Circularity Check
No significant circularity: Pilot's effectiveness claims rest on benchmark experiments, not on a derivation that reduces to its own inputs.
full rationale
The paper is an empirical systems paper. Its central claims are that the proposed Pilot framework improves federated multimodal instruction tuning on two constructed scenarios, supported by comparisons against FedAvg, FedProx, FedAdam, Shepherd, and FedDPA in Tables 1 and 2, plus ablations in Tables 3-6. No equation in the paper is derived from the result it is claimed to predict, and no fitted parameter is renamed as a prediction. The adaptive text-adapter aggregation (Eqs. 11-12) is a heuristic based on Euclidean distance in parameter space, but a heuristic that lacks theoretical justification is a correctness or robustness concern, not circularity: the aggregation weights are computed from trained parameters, not from the evaluation targets, and the method is tested on held-out test splits. The choice of Top-M=6 is a hyperparameter, studied in Table 5 and reported as a tuning choice, not as a derived prediction. Self-citations in the introduction and related work are contextual and do not carry the argument; no uniqueness theorem or load-bearing result is imported from the authors' prior work. The paper's limitations (e.g., single data split, no error bars) weaken the evidence but do not make the derivation circular. Accordingly, no circular step satisfying the required reduction standard is present.
Assumptions & free parameters
free parameters (6)
- Lambda0 (difference loss coefficient) =
0.1
- Lambda1 (load balancing loss coefficient) =
0.1
- Lambda2 (router z-loss coefficient) =
0.01
- Top-M (adaptive aggregation neighbor count) =
6
- LoRA rank =
64
- Communication rounds R =
3
assumptions (5)
- domain assumption Frozen CLIP visual encoder features are sufficient for cross-task knowledge transfer.
- ad hoc to paper The orthogonality constraint (Eq. 3) separates task-specific and client-specific features.
- domain assumption Clients with the same task type have similar task-specific adapters, justifying per-task averaging in Eq. 10.
- ad hoc to paper Euclidean distance between text-adapter parameters indicates beneficial aggregation.
- domain assumption The router with load-balancing losses improves cross-task learning.
invented entities (2)
-
Cross-task adapter (psi_c)
-
Client-specific adapter (psi_s)
Cite this review
Pith. "Pith review of Pilot: Building the Federated Multimodal Instruction Tuning Framework." pith.science (2026). https://pith.science/paper/UDPV3RHU
@misc{pith2026250113985,
author = {Pith},
title = {Pith review of: Pilot: Building the Federated Multimodal Instruction Tuning Framework},
year = {2026},
howpublished = {\url{https://pith.science/paper/UDPV3RHU}},
note = {Machine review of arXiv:2501.13985}
}
read the original abstract
In this paper, we explore a novel federated multimodal instruction tuning task(FedMIT), which is significant for collaboratively fine-tuning MLLMs on different types of multimodal instruction data on distributed devices. To solve the new task, we propose a federated multimodal instruction tuning framework(Pilot). Our framework integrates two stages of "adapter on adapter" into the connector of the vision encoder and the LLM. In stage 1, we extract task-specific features and client-specific features from visual information. In stage 2, we build the cross-task Mixture-of-Adapters(CT-MoA) module to perform cross-task interaction. Each client can not only capture personalized information of local data and learn task-related multimodal information, but also learn general knowledge from other tasks. In addition, we introduce an adaptive parameter aggregation strategy for text training parameters, which optimizes parameter aggregation by calculating weights based on the euclidean distance between parameters, so that parameter aggregation can benefit from positive effects to the greatest extent while effectively reducing negative effects. Our framework can collaboratively exploit distributed data from different local clients to learn cross-task knowledge without being affected by the task heterogeneity during instruction tuning. The effectiveness of our method is verified in two different cross-task scenarios.
Figures
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Arevalo, C. A.; Noorbakhsh, S. L.; Dong, Y.; Hong, Y.; and Wang, B. 2024. Task-Agnostic Privacy-Preserving Representation Learning for Federated Learning against Attribute Inference Attacks. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, 10909--10917
work page 2024
-
[4]
Bousmalis, K.; Trigeorgis, G.; Silberman, N.; Krishnan, D.; and Erhan, D. 2016. Domain separation networks. Advances in neural information processing systems, 29
work page 2016
-
[5]
Chen, D.; Liu, J.; Dai, W.; and Wang, B. 2024 a . Visual instruction tuning with polite flamingo. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, 17745--17753
work page 2024
-
[6]
Chen, H.; Zhang, Y.; Krompass, D.; Gu, J.; and Tresp, V. 2024 b . Feddat: An approach for foundation model finetuning in multi-modal heterogeneous federated learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, 11285--11293
work page 2024
-
[7]
Chen, J.; and Zhang, A. 2024. On Disentanglement of Asymmetrical Knowledge Transfer for Modality-Task Agnostic Federated Learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, 11311--11319
work page 2024
-
[8]
Chen, K.; Zhang, Z.; Zeng, W.; Zhang, R.; Zhu, F.; and Zhao, R. 2023. Shikra: Unleashing multimodal llm's referential dialogue magic. arXiv preprint arXiv:2306.15195
arXiv 2023
Show all 52 references
-
[9]
E.; et al
Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J. E.; et al. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90\ See https://vicuna. lmsys. org (accessed 14 April 2023), 2(3): 6
2023
-
[10]
Dai, W.; Li, J.; et al. 2023. InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning. arXiv:2305.06500
2023 arXiv
-
[11]
Dhariwal, P.; and Nichol, A. 2021. Diffusion models beat gans on image synthesis. Advances in neural information processing systems, 34: 8780--8794
2021
-
[12]
S.; et al
Driess, D.; Xia, F.; Sajjadi, M. S.; et al. 2023. Palm-e: An embodied multimodal language model. arXiv preprint arXiv:2303.03378
2023 arXiv
-
[13]
Fallah, A.; Mokhtari, A.; and Ozdaglar, A. 2020. Personalized federated learning: A meta-learning approach. arXiv preprint arXiv:2002.07948
2020 arXiv
-
[14]
Fan, T.; Kang, Y.; Ma, G.; Chen, W.; Wei, W.; Fan, L.; and Yang, Q. 2023. Fate-llm: A industrial grade federated learning framework for large language models. arXiv preprint arXiv:2310.10049
2023 arXiv
-
[15]
Finn, C.; Abbeel, P.; and Levine, S. 2017. Model-agnostic meta-learning for fast adaptation of deep networks. In International conference on machine learning, 1126--1135. PMLR
2017
-
[16]
Gorbunov, E.; Hanzely, F.; and Richt \'a rik, P. 2021. Local sgd: Unified theory and new efficient methods. In International Conference on Artificial Intelligence and Statistics, 3556--3564. PMLR
2021
-
[17]
Guo, T.; et al. 2023. Pfedprompt: Learning personalized prompt for vision-language models in federated learning. In Proceedings of the ACM Web Conference 2023, 1364--1374
2023
-
[18]
He, J.; Guo, H.; Tang, M.; and Wang, J. 2023. Continual instruction tuning for large multimodal models. arXiv preprint arXiv:2311.16206
2023 arXiv
-
[19]
J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2021. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685
2021 arXiv
-
[20]
Hudson. 2019. Gqa: A new dataset for real-world visual reasoning and compositional question answering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 6700--6709
2019
-
[21]
Jia, Y.; Zhang, X.; Beheshti, A.; and Dou, W. 2024. FedLPS: Heterogeneous Federated Learning for Multiple Tasks with Local Parameter Sharing. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, 12848--12856
2024
-
[22]
Kazemzadeh, S.; Ordonez, V.; Matten, M.; and Berg, T. 2014. Referitgame: Referring to objects in photographs of natural scenes. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), 787--798
2014
-
[23]
Li, J.; Li, D.; Savarese, S.; and Hoi, S. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning, 19730--19742. PMLR
2023
-
[24]
K.; Zaheer, M.; Sanjabi, M.; Talwalkar, A.; and Smith, V
Li, T.; Sahu, A. K.; Zaheer, M.; Sanjabi, M.; Talwalkar, A.; and Smith, V. 2020. Federated optimization in heterogeneous networks. Proceedings of Machine Learning and Systems, 2: 429--450
2020
-
[25]
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Doll \'a r, P.; and Zitnick, C. L. 2014. Microsoft coco: Common objects in context. In Computer Vision--ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 1...
2014
-
[26]
Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2024. Visual instruction tuning. Advances in neural information processing systems, 36
2024
-
[27]
Lu, P.; Mishra, S.; Xia, T.; Qiu, L.; Chang, K.-W.; Zhu, S.-C.; Tafjord, O.; Clark, P.; and Kalyan, A. 2022. Learn to explain: Multimodal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems, 35: 2507--2521
2022
-
[28]
Lu, W.; Hu, X.; Wang, J.; and Xie, X. 2023. Fedclip: Fast generalization and personalization for clip in federated learning. arXiv preprint arXiv:2302.13485
2023 arXiv
-
[29]
McMahan, B.; Moore, E.; Ramage, D.; Hampson, S.; and y Arcas, B. A. 2017. Communication-efficient learning of deep networks from decentralized data. In Artificial intelligence and statistics, 1273--1282. PMLR
2017
-
[30]
K.; and Chakraborty, A
Mishra, A.; Shekhar, S.; Singh, A. K.; and Chakraborty, A. 2019. Ocr-vqa: Visual question answering by reading text in images. In 2019 international conference on document analysis and recognition (ICDAR), 947--952. IEEE
2019
-
[31]
W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learning, 8748--8763. PMLR
2021
-
[32]
Reddi, S.; Charles, Z.; Zaheer, M.; Garrett, Z.; Rush, K.; Kone c n \`y , J.; Kumar, S.; and McMahan, H. B. 2020. Adaptive federated optimization. arXiv preprint arXiv:2003.00295
2020 arXiv
-
[33]
Shen, Y.; Xu, Z.; Wang, Q.; Cheng, Y.; Yin, W.; and Huang, L. 2024. Multimodal Instruction Tuning with Conditional Mixture of LoRA. arXiv preprint arXiv:2402.15896
2024 arXiv
-
[34]
Shi, J.; Zheng, S.; Yin, X.; Lu, Y.; Xie, Y.; and Qu, Y. 2023. Clip-guided federated learning on heterogeneous and long-tailed data. arXiv preprint arXiv:2312.08648
2023 arXiv
-
[35]
Su, S.; Yang, M.; Li, B.; and Xue, X. 2024. Federated adaptive prompt tuning for multi-domain collaborative learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, 15117--15125
2024
-
[36]
Sun, L.; Zhang, K.; Li, Q.; and Lou, R. 2024. Umie: Unified multimodal information extraction with instruction tuning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, 19062--19070
2024
-
[37]
Y.; Guu, K.; Yu, A
Wei, J.; Bosma, M.; Zhao, V. Y.; Guu, K.; Yu, A. W.; Lester, B.; Du, N.; Dai, A. M.; and Le, Q. V. 2021. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652
2021 arXiv
-
[38]
Xiao, L.; Yang, X.; Peng, F.; Wang, Y.; and Xu, C. 2024 a . Hivg: Hierarchical multimodal fine-grained modulation for visual grounding. In Proceedings of the 32nd ACM International Conference on Multimedia, 5460--5469
2024
-
[39]
Xiao, L.; Yang, X.; Peng, F.; Wang, Y.; and Xu, C. 2024 b . OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring Modeling. arXiv preprint arXiv:2410.08021
2024 arXiv
-
[40]
Xiong, B.; Yang, X.; Song, Y.; Wang, Y.; and Xu, C. 2023. Client-Adaptive Cross-Model Reconstruction Network for Modality-Incomplete Multimodal Federated Learning. In Proceedings of the 31st ACM International Conference on Multimedia, 1241--1249
2023
-
[41]
Xiong, B.; Yang, X.; Song, Y.; Wang, Y.; and Xu, C. 2024. Modality-Collaborative Test-Time Adaptation for Action Recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 26732--26741
2024
-
[42]
Xu, Y.; Yang, X.; Song, Y.; and Xu, C. 2024. Libra: Building Decoupled Vision System on Large Language Models. arXiv preprint arXiv:2405.10140
2024 arXiv
-
[43]
Xu, Z.; et al. 2022. Multiinstruct: Improving multi-modal zero-shot learning via instruction tuning. arXiv preprint arXiv:2212.10773
2022 arXiv
-
[44]
Yang, M.; Su, S.; Li, B.; and Xue, X. 2024 a . Exploring One-Shot Semi-supervised Federated Learning with Pre-trained Diffusion Models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, 16325--16333
2024
-
[45]
Yang, X.; Xiong, B.; Huang, Y.; and Xu, C. 2024 b . Cross-Modal Federated Human Activity Recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence
2024
-
[46]
Yang, Y.; Long, G.; Shen, T.; Jiang, J.; and Blumenstein, M. 2024 c . Dual-Personalizing Adapter for Federated Foundation Models. arXiv preprint arXiv:2403.19211
2024 arXiv
-
[47]
Ye, Q.; Xu, H.; Xu, G.; Ye, J.; Yan, M.; Zhou, Y.; Wang, J.; Hu, A.; Shi, P.; Shi, Y.; et al. 2023. mplug-owl: Modularization empowers large language models with multimodality. arXiv preprint arXiv:2304.14178
2023 arXiv
-
[48]
Ye, R.; Wang, W.; Chai, J.; Li, D.; Li, Z.; Xu, Y.; Du, Y.; Wang, Y.; and Chen, S. 2024. Openfedllm: Training large language models on decentralized private data via federated learning. arXiv preprint arXiv:2402.06954
2024 arXiv
-
[49]
Zhang, J.; Liu, Y.; Hua, Y.; and Cao, J. 2024 a . Fedtgp: Trainable global prototypes with adaptive-margin-enhanced contrastive learning for data and model heterogeneity in federated learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, 16768--16776
2024
-
[50]
Zhang, J.; Vahidian, S.; Kuo, M.; Li, C.; Zhang, R.; Yu, T.; Wang, G.; and Chen, Y. 2024 b . Towards building the federatedGPT: Federated instruction tuning. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 6915--6919. IEEE
2024
-
[51]
Zhu, D.; Chen, J.; Shen, X.; Li, X.; and Elhoseiny, M. 2023. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592
2023 arXiv
-
[52]
Zoph, B.; Bello, I.; Kumar, S.; Du, N.; Huang, Y.; Dean, J.; Shazeer, N.; and Fedus, W. 2022. St-moe: Designing stable and transferable sparse expert models. arXiv preprint arXiv:2202.08906
2022 arXiv
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.