{"id":"8af771ce-dd72-466e-a665-78441e3f0e74","arxiv_id":"2504.21099","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A review that sorts recent federated-learning PEFT approaches into additive, selective, and reparameterized (LoRA-style) families and maps them onto NLP and vision applications.","lead":"This paper surveys research that combines parameter-efficient fine-tuning (PEFT) with federated learning, grouping methods into additive, selective, and reparameterized approaches. It is a reference map of recent FL-PEFT methods and their application to language and vision tasks, useful for researchers entering the area.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Taxonomy is internally inconsistent: FedAdapter is listed as additive though it trains existing layers, FedRA as selective though it uses adapters, and CELL/HeteroFL as selective PEFT though the text admits they are not PEFT.","rationale":"The reader's weakest_assumption concerns comprehensiveness and the lack of a search protocol. My concern is more internal and more specific: the organizing taxonomy itself is not applied consistently under the survey's own definitions. This goes beyond the mechanical citation errors the reader noted (PromptFL/PromptFolio both assigned to [86], duplicate references [17]/[37]) and directly affects the central 'systematic and comprehensive review' claim. The survey is still useful, and the LoRA aggregation bias discussion (Eqs. 3-6) is a correct technical contribution, so I agree with CONDITIONAL as a verdict. I set agreement_with_reader to 'partial' because the reader identified the missing search methodology as the weakest assumption, whereas my analysis shows category errors that can be demonstrated from the survey's own text and tables. The concrete test is inexpensive: re-read the primary sources listed in Tables 1-3 and re-check assignments against Section 2.2 definitions. If only a handful of entries are misassigned, the CONDITIONAL verdict stands with minor revisions. If many entries move, the survey's map-like utility is genuinely compromised and a major revision or reject would be warranted. No author-conduct concern is raised; the issue is purely about the accuracy of the taxonomy and coverage as stated.","tokens_in":23542,"tokens_out":3242,"duration_ms":30048,"concrete_test":"For each method in Tables 1-3, read the original paper and record: (a) is the backbone frozen with only new small modules trained (additive), (b) is only a subset of existing backbone parameters trained (selective), or (c) are weight updates reparameterized in low-rank form (reparameterized)? Then check the table assignment against the survey's own definitions in Section 2.2. At minimum, verify FedAdapter (does it train full backbone layers or only adapters?), FedRA (does it rely on adapters or only selects existing layers?), and CELL/HeteroFL (are they PEFT under the survey's definition at all?). Report how many entries change category or are dropped from Tables 1-3. If a substantial fraction of entries move, the 'systematic categorization' claim fails; if only 2-3 entries change, the CONDITIONAL verdict stands with minor revisions.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is a systematic, comprehensive review organized by a clean Additive/Selective/Reparameterized taxonomy. The taxonomy is applied inconsistently even within the survey's own definitions. (1) Section 3.1.1 and Table 1 classify FedAdapter [13] as Additive PEFT, but FedAdapter (Cai et al., 2022) progressively unfreezes and trains large subsets of the original network while also using adapters, so it does not satisfy the survey's own defining criterion for additive PEFT: keeping the backbone frozen and adding only small trainable modules. (2) Section 3.2 and Table 2 classify FedRA [97] as Selective PEFT, yet the survey's own text says FedRA randomly allocates model layers to clients and applies adapter-based fine-tuning; the FedRA paper combines layer allocation with adapters, so it is not purely selective under the definition given. The paper even concedes FedRA spans categories, but tables force it into one column. (3) Table 2 includes CELL [94] and HeteroFL [29], which the text itself describes as traditional FL not focused on foundation models and which are not PEFT under the survey's own definition (no frozen pre-trained backbone). This inflates the selective-PEFT count and makes the coverage claim harder to verify. These are not purely cosmetic citation issues; the load-bearing assumption that the survey is a reliable map organized by a consistent PEFT taxonomy is weakened. A reader using this survey to select 'selective PEFT' methods in FL would be pointed to methods that are neither selective PEFT nor foundation-model fine-tuning. The CONDITIONAL verdict is appropriate, but the condition must include checking all category assignments, not just fixing citation formatting.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of parameter-efficient fine-tuning (PEFT) methods in federated learning (FL). It organizes the literature into three categories—additive, selective, and reparameterized PEFT—following the taxonomy of Han et al. [42], and it further groups works by application domain (NLP and vision). The survey also reviews challenges in FL (data heterogeneity, communication efficiency, computational constraints, privacy), discusses LoRA-specific issues such as server-side aggregation bias, and outlines future directions. The paper does not report a formal search protocol or inclusion criteria.","tokens_in":23916,"tokens_out":6349,"duration_ms":64473,"significance":"If the survey were accurate and comprehensive, it would be a useful reference map for the rapidly growing FL-PEFT area. The paper covers a broad set of recent methods, and the discussion of LoRA aggregation bias around Eq. (5)-(6) is a concrete technical contribution to the survey's conceptual framing. The application-oriented tables (Tables 4 and 5) also provide practical guidance. However, the manuscript's central claim is that it is a 'systematic and comprehensive review,' and that claim is weakened by internal taxonomy inconsistencies and citation/naming errors. The misclassification of FedRA in particular cuts against the survey's stated organizing principle, and the absence of a search protocol makes the comprehensiveness claim hard to verify.","major_comments":[{"comment":"FedRA [97] is listed under 'Selective PEFT' in Table 2, but the text at the end of Section 3.2 explicitly states that FedRA 'applies adapter-based fine-tuning,' and Section 4.2.1 / Table 5 describe FedRA as 'Selective and Additive.' Since the entire survey is organized around the Additive/Selective/Reparameterized taxonomy, this internal contradiction is load-bearing: a reader consulting Table 2 to identify selective-PEFT methods will be pointed to a method that the paper itself treats as hybrid. The taxonomy either needs a defined hybrid category or the tables and prose must be aligned.","section":"Section 3.2 and Table 2"},{"comment":"DepthFL [51] is included as a selective-PEFT method in Table 2, but the description in Section 3.2 (depth-wise pruning of a global model with mutual self-distillation) does not match the survey's definition of selective PEFT as fine-tuning a subset of the parameters of a frozen pre-trained backbone. The text correctly distinguishes 'traditional federated learning' methods such as CELL and HeteroFL from FL-PEFT methods, yet DepthFL is placed in the FL-PEFT portion and counted in Table 2. This blurs the boundary the survey needs to maintain to support its comprehensive-taxonomy claim. I do not find a comparable problem with CELL and HeteroFL, which the text frames only as background and which do not appear in Table 2.","section":"Section 3.2 and Table 2"},{"comment":"The abstract and introduction claim a 'systematic and comprehensive review,' but the manuscript does not provide a search protocol, database list, inclusion/exclusion criteria, or coverage cutoff date. Without such information, the reader cannot verify whether the reviewed set of papers is representative or whether omissions are intentional. Given that the central value of a survey is its reliability as a map of the literature, this lack of methodological transparency is a substantive weakness, not merely a presentation issue. The authors should either add a short methodology paragraph or soften the comprehensiveness claim.","section":"Section 1 and Section 3"}],"minor_comments":[{"comment":"In Section 3.1.2, both PromptFL and PromptFolio are cited as reference [86], but Table 1 lists PromptFL as [41]. Please correct the in-text citation for PromptFL.","section":"Section 3.1.2 and Table 1"},{"comment":"The method referenced as 'FedBF [127]' in Section 3.2 and Table 2 is called 'FedPETuning' in Table 4 and in the reference list entry itself. Please unify the method name and the table labels.","section":"Section 3.2, Table 2, and Table 4"},{"comment":"References [17] and [37] share the same title ('Prompt-enhanced federated content representation learning for cross-domain recommendation'); these appear to be the arXiv and published versions of the same work. The manuscript should cite one version consistently and distinguish them.","section":"References"},{"comment":"The row for FedBiOT contains the typo 'AReparameterized'; it should read 'Reparameterized.'","section":"Table 4"},{"comment":"The spelling 'Reparametrized' is used in the Section 3.3 heading while the rest of the paper uses 'Reparameterized.' Please make the spelling consistent.","section":"Section 3.3 heading and elsewhere"},{"comment":"The text in Section 3.1.1 refers to 'ADAFEDSELECKD [32]' while Table 1 and the reference entry use 'ADAFED.' Please align the naming.","section":"Section 3.1.1 and Table 1"},{"comment":"The entry 'LLaV A 1.5' should be 'LLaVA 1.5'.","section":"Table 5"}],"recommendation":"major_revision","confidential_remarks":"The paper is an arXiv preprint and the issues identified are fixable, but the FedRA/DepthFL classification problems directly affect the survey's organizing taxonomy and should be resolved before publication. I would also encourage the editor to ask for a brief statement of scope or methodology, since the journal's standards for surveys typically require more transparency about coverage."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Zoe — quick take on arXiv:2504.21099. It's a survey of parameter-efficient fine-tuning in federated learning, and honestly the field needs one. The paper is organized by additive/selective/reparameterized PEFT and then by application. The LoRA aggregation bias discussion (Eqs. 4-6) is correct and genuinely useful: the averaged product of low-rank factors is not the average of the full updates, and that's the key technical insight that motivates half the papers in the reparameterized section. If you're new to the area, this gives you the lay of the land.\n\nWhat's new is mostly the FL-specific arrangement and the recent 2023-2025 coverage. The novelty is organizational, not technical, which is fine for a survey.\n\nNow the soft spots. The reader's conditional verdict is right, but I'd push the condition further than citation formatting. The taxonomy as applied in the tables is internally inconsistent. FedAdapter is listed as additive in Table 1 even though its whole point is progressively unfreezing and training parts of the backbone; the survey's own definition of additive says the backbone stays frozen. FedRA sits in Table 2 under selective, while the text and Table 5 both acknowledge it combines layer selection with adapters. And Table 2 includes CELL and HeteroFL, which the text itself admits are traditional FL methods, not PEFT at all. That's not cosmetic: a reader using this survey to pick a 'selective PEFT' method would be pointed at things that don't fit the category.\n\nThe citation issues are smaller but real. PromptFL and PromptFolio both get reference [86] in the text, though Table 1 correctly distinguishes them ([41] vs [86]). And references [17] and [37] are the same paper, one arXiv and one WWW version. Also, the abstract claims a 'comprehensive review' but there's no stated search protocol, inclusion criteria, or cutoff date, so comprehensiveness is unverifiable.\n\nNone of this kills the paper. The central map is usable, the LoRA bias section is solid, and the self-citations to the authors' own LoRA-FAIR and FedALT don't load-bear on the organization. But the category assignments need checking method-by-method, and the citations need a cleanup.\n\nRecommendation: send to peer review. It deserves a serious referee, but I'd make the revision condition explicitly include re-auditing every table entry against the survey's own definitions, not just fixing reference numbers. This is a paper worth having in the literature once it's reliable.","headline":"A useful but sloppy survey of PEFT in federated learning; the organization is sound, but the category assignments and citations need a careful pass before it can serve as a dependable reference.","tokens_in":24418,"tokens_out":2593,"would_cite":true,"duration_ms":25172,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A survey maps parameter-efficient fine-tuning inside federated learning onto three families—additive, selective, and reparameterized—and identifies a server-side aggregation bias that makes naive LoRA averaging deviate from the ideal…","keywords":["federated learning","parameter-efficient fine-tuning","low-rank adaptation","LoRA","prompt tuning","adapters","foundation models","data heterogeneity"],"falsifier":"Compile the set of federated PEFT methods published before the newest reference in the survey and check whether every one appears in Tables 1–3 under the correct category; one missing or misclassified method that is demonstrably part of the literature would show the comprehensiveness claim to be overstated. A second check: identify a method whose update rule mixes categories (such as FedRA, which the survey itself notes spans selective and additive) and show the taxonomy cannot accommodate it without double-labeling.","tokens_in":23265,"feed_emoji":"🧩","tokens_out":5670,"duration_ms":56008,"temperature":0.7,"pith_summary":"This survey tries to establish that the scattered literature on parameter-efficient fine-tuning inside federated learning can be organized into a single map. It groups every method it covers into three families—additive (new trainable modules such as adapters and soft prompts), selective (updates only a subset of existing parameters such as bias terms), and reparameterized (low-rank decompositions such as LoRA)—and attaches each method to the federated challenge it solves: data heterogeneity, communication cost, computational limits, or privacy. A sympathetic reader would care because the map turns a pile of method papers into guidance: which PEFT class to pick for which federated bottleneck, and which foundation model and dataset pairs have already been tested. The survey also singles out a concrete technical obstacle unique to federated LoRA: averaging local low-rank updates at the server differs from the ideal average of the full updates, which it calls server-side aggregation bias.","feed_headline":"Three families of tricks make foundation-model fine-tuning federated","feed_subtitle":"A survey sorts FL-PEFT methods into additive, selective, and reparameterized, and names the LoRA aggregation bias","key_machinery":"The load-bearing object is the PEFT taxonomy of [42]—additive, selective, and reparameterized—transplanted into the federated setting. Additive methods insert trainable adapters or soft prompts into a frozen foundation model; selective methods freeze most weights and update only distinguished subsets such as bias terms; reparameterized methods express weight updates as products of low-rank matrices, LoRA being the canonical case. The survey's analytical device is the comparison between the ideal federated update, a weighted sum of local full updates, and what naive LoRA aggregation actually computes, a product of weighted averages of low-rank factors; the mismatch, called server-side aggregation bias, organizes the reparameterized section and motivates the correction methods reviewed there.","core_discovery":"The paper's central claim is that every current approach to parameter-efficient fine-tuning in federated learning falls into one of three PEFT families and can be reviewed against a standard set of FL challenges. Within additive tuning, adapters and soft prompts are the workhorses, with methods like FedPrompt reducing communication to prompt parameters and adapter methods personalizing per client. Within selective tuning, bias-only updates and parameter-selection strategies provide cheap adaptation. Within reparameterized tuning, LoRA dominates, and the survey identifies a distinctive federated failure mode: because clients send low-rank factors $A_k$ and $B_k$ rather than full updates, the server's weighted average of products is not the product of averages, so naive FedIT aggregation deviates from the ideal global update. Several methods (FFA-LoRA, RoLoRA, FLoRA, FedEx-LoRA, LoRA-FAIR) are reviewed as attempts to remove that bias. On the application side, the survey maps these methods onto NLP tasks (text classification, text generation, machine translation, recommendation) and CV tasks (image classification, domain adaptation, multimodal learning), with recommended datasets and foundation models for each.","pith_inferences":["If the three-way taxonomy is as exhaustive as the survey claims, then any future FL-PEFT method that cannot be filed as additive, selective, or reparameterized would be a genuinely new category; the survey does not predict what that would look like.","The server-side aggregation bias identified for LoRA should in principle afflict any reparameterized method whose local updates are averaged before being composed, so DoRA- or VeRA-style decompositions in federated settings would likely need the same correction machinery; the paper does not draw this extension.","A testable next step would be a standardized benchmark that reuses the exact model-dataset pairs in Tables 4 and 5, so that additive, selective, and reparameterized methods could be compared head-to-head under identical non-IID partitions.","Because the survey is a snapshot, its map will age quickly; the same taxonomy could be maintained as a living document that tracks new FL-PEFT methods as they appear."],"forward_implications":["A practitioner choosing a federated fine-tuning method can use the tables to shortlist by the bottleneck they face: adapters and prompts for communication savings, selective bias tuning for extreme resource limits, and LoRA variants for accuracy near full fine-tuning.","FedIT-style naive LoRA aggregation is expected to be biased, so any new federated LoRA method should compare against a correction (fixed A, alternating freeze, stacking, or residual term) rather than treating FedIT as an unbiased baseline.","The survey's application tables imply that BERT-family models and GLUE tasks are the default testbed for text classification, while CLIP and image-classification datasets are the default testbed for vision, giving new work a ready-made evaluation template.","Scaling to trillion-parameter models will require federated PEFT to solve communication and memory bottlenecks even for the small trainable parts, a direction the survey flags as open."],"supporting_citations":[{"why":"Supplies the additive, selective, and reparameterized PEFT taxonomy that organizes the entire survey.","marker":"[42]"},{"why":"Defines the federated learning setup and the weighted aggregation rule that the survey uses as the ideal global update.","marker":"[79]"},{"why":"Introduces LoRA, the low-rank reparameterization on which the reparameterized section and the aggregation-bias analysis depend.","marker":"[47]"},{"why":"FedIT is the baseline that directly combines LoRA with FL, and its averaging rule is what exposes server-side aggregation bias.","marker":"[121]"},{"why":"FedPrompt is the canonical prompt-tuning-in-FL method that motivates the additive prompt-tuning discussion.","marker":"[128]"},{"why":"PromptFL is an early framework for collaborative soft-prompt training over frozen CLIP, grounding communication-efficiency claims for additive PEFT.","marker":"[41]"},{"why":"FedPETuning combines FedAvg with BitFit-style selective tuning and anchors the selective PEFT category for NLP.","marker":"[127]"},{"why":"FedBiOT illustrates the closed-source foundation model setting where clients fine-tune only lightweight LoRA adapters.","marker":"[112]"},{"why":"Identified as the prior survey whose coverage ends before recent FL-PEFT methods, defining the gap this survey aims to close.","marker":"[64]"}],"fun_headline_variants":["Survey: three PEFT families for federated fine-tuning","LoRA aggregation bias named in federated PEFT survey","Federated PEFT: additive, selective, reparameterized","How PEFT adapts to federated learning: a survey","Survey exposes LoRA's federated averaging bias"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's practical value rests on its coverage being complete and representative, yet it provides no search protocol, database list, inclusion criteria, or cutoff date, so any important FL-PEFT method missing from the three categories would undermine the map it offers.","fun_headline_variants_meta":{"raw":{"variants":["Survey: three PEFT families for federated fine-tuning","LoRA aggregation bias named in federated PEFT survey","Federated PEFT: additive, selective, reparameterized","How PEFT adapts to federated learning: a survey","Survey exposes LoRA's federated averaging bias"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000229,"raw_usage":{"total_tokens":1509,"prompt_tokens":1008,"completion_tokens":501,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":624,"completion_tokens_details":{"reasoning_tokens":419}},"tokens_in":624,"tokens_out":501,"duration_ms":5278,"temperature":1.0,"reasoning_tokens":419,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:12:39.553989+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compile the set of federated PEFT methods published before the newest reference in the survey and check whether every one appears in Tables 1–3 under the correct category; one missing or misclassified method that is demonstrably part of the literature would show the comprehensiveness claim to be overstated. A second check: identify a method whose update rule mixes categories (such as FedRA, which the survey itself notes spans selective and additive) and show the taxonomy cannot accommodate it without double-labeling.","supporting_citations":[{"cited_title":"In: ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","cited_arxiv_id":null,"evidence_quote":"FedIT is the baseline that directly combines LoRA with FL, and its averaging rule is what exposes server-side aggregation bias."},{"cited_title":"In: ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","cited_arxiv_id":null,"evidence_quote":"FedPrompt is the canonical prompt-tuning-in-FL method that motivates the additive prompt-tuning discussion."},{"cited_title":"In: Annual Meeting of the Association of Computational Linguistics 2023","cited_arxiv_id":null,"evidence_quote":"FedPETuning combines FedAvg with BitFit-style selective tuning and anchors the selective PEFT category for NLP."},{"cited_title":"In: Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining","cited_arxiv_id":null,"evidence_quote":"FedBiOT illustrates the closed-source foundation model setting where clients fine-tune only lightweight LoRA adapters."}],"review_version":1}