{"id":"7a740b5b-e8a3-4bae-b512-96cbbdd9927e","arxiv_id":"2411.12377","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A comprehensive survey that organizes non-IID data in federated learning into taxonomies of skew types, partition protocols, and metrics, with a meta-analysis of 235 selected papers.","lead":"This paper surveys how non-IID data is defined, simulated, and quantified in federated learning, organizing prior work into taxonomies of data skew, partition protocols, metrics, solutions, and frameworks. A smart generalist would read it to see how the field is standardizing benchmarks and which open problems remain in handling heterogeneous data.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The PRISMA flow in Section III-A/Fig. 3 is internally inconsistent (5489 vs 5580 unique papers; 235 vs 219+36 included), so the meta-analysis percentages and the 'complete coverage' claim rest on an unreproducible paper set.","rationale":"Reading the manuscript in good faith, the qualitative contributions are real: the taxonomy of skew types (label, attribute, quantity, spatiotemporal, participation, modality) is clearly organized, the partition protocol and metric taxonomies are useful, and the inclusion of modality skew is a genuine addition. The derivation in Section IV connecting marginal and conditional label skew via Bayes' theorem is a sound mathematical observation. My concern is not with the conceptual framework but with the empirical engine that powers the survey's most distinctive quantitative claims. The paper repeatedly reports precise prevalence statistics as 'key findings' (e.g., '13.1% of them employed metrics,' '14.2% of them use FL frameworks,' '60.3% of the papers that include label skew approaches do not mention anything about quantity skew'), and the claim of 'complete coverage' in Table I is a comparative statement against prior surveys. All of these rest on the 235-paper corpus, yet the corpus is described inconsistently within the same section: 5489 vs 5580 unique papers after deduplication, 202+36 vs 235 vs 253 vs 219+36 included papers, with no search scripts, query strings, or paper list provided. This means the reported percentages are not auditable, and the central claim of a definitive, quantified map of the field is fragile. The issue is correctable: releasing the exact query set and included-paper list, and reconciling the flow diagram, would let an independent reviewer verify or correct the statistics. I therefore agree with the reader's assessment that the appropriate verdict is conditional on this revision, not a rejection.","tokens_in":41649,"tokens_out":4418,"duration_ms":38260,"concrete_test":"Obtain from the authors the complete list of included papers (with the final included count reconciled to a single number: 235, 238, 253, or 255) and the 83 search queries used for each of the six databases. Independently re-classify every paper in that list into the Fig. 7 skew subtypes using the paper's own definitions, and recompute the prevalence percentages reported in Section IV-A and Section VIII (label skew, attribute skew, quantity skew, modality skew, metric usage, framework usage). If the recomputed percentages differ by more than one percentage point from the published values, or if the PRISMA flow numbers cannot be reconciled to exactly one included set, the quantitative meta-analysis and the 'complete coverage' claim are not supported and must be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The survey's central claim of 'complete coverage of all the aspects related to non-IID data' (Section I-B, Table I) and its quantified meta-analysis — e.g., '13.1% of papers use metrics' and '14.2% use frameworks' (Section VIII) — depend directly on the corpus selected by the PRISMA-style screening in Section III-A. That corpus is not internally consistent. The text states that after duplicate removal 'we compiled a collection of 5489 unique papers' (Section III-A), while Fig. 3 reports 'Remaining after duplicates removal: 5580 papers.' The eligibility stage says '202 papers were marked as useful' and '36 additional papers are marked as useful,' which sums to 238, but the text then says 'we included 235 papers.' Fig. 3 shows 'Useful: 253 papers,' then '219 papers' from top-tier venues plus '36 papers' from unpublished/high-citation sources, summing to 255. These discrepancies change the denominator of every reported percentage: the difference between 235 and 255 is 8.5% of the corpus, and the prevalence of label skew subtypes, metric usage, framework adoption, and the '2% modality skew' figure would shift accordingly. The selection criteria also include subjective elements — 'top-tier conferences/journals,' QS top-100 university affiliation, and citation-count thresholds — without listing the 83 search queries or the included paper list, so an independent reviewer cannot reconstruct the set or verify that the taxonomy percentages are stable. This is a correctness risk in the empirical foundation of the survey's headline contribution.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey aims to organize and synthesize the literature on non-IID data in federated learning. It introduces a taxonomy of data heterogeneity types (label skew, attribute skew, quantity skew, spatiotemporal heterogeneity, participation skew, and modality skew), catalogs partition protocols and non-IID metrics, reviews popular solution methods, compares standardized FL frameworks, and presents a PRISMA-style meta-analysis of a selected corpus of papers. The paper claims to be the first technical survey with complete coverage of all aspects of non-IIDness in FL, including a novel treatment of modality skew and a quantitative map of current research practice.","tokens_in":41945,"tokens_out":3922,"duration_ms":40882,"significance":"If the quantitative claims were reproducible, this would be a valuable reference survey for the FL community: the proposed taxonomy is broad and thoughtfully structured, the formal definitions of skew types in Section IV are a useful starting point for standardization, and the comparison of frameworks in Table II addresses a real gap. The paper also responsibly highlights understudied areas such as modality skew and combined-skew metrics. However, the quantified meta-analysis is currently the weakest part of the paper: the PRISMA flow in Section III-A/Fig. 3 has internal arithmetic inconsistencies, and the classification of papers into taxonomy subtypes appears to rest on undocumented, subjective judgments. Because the reported prevalence percentages and the \"complete coverage\" claim depend directly on this corpus, the numerical findings and the overarching contribution claim are not yet fully supported.","major_comments":[{"comment":"The PRISMA flow is internally inconsistent, and this directly affects the denominators of every reported percentage in the meta-analysis. The text states that after duplicate removal \"we compiled a collection of 5489 unique papers,\" while Fig. 3 reports \"Remaining after duplicates removal: 5580 papers.\" In the eligibility stage, the text says \"202 papers were marked as useful\" and \"36 additional papers are marked as useful,\" which sums to 238, but the next sentence says \"we included 235 papers.\" Fig. 3 shows \"Useful: 253 papers\" and then \"219 papers\" from top-tier venues plus \"36 papers\" from unpublished/high-citation sources, totaling 255. These discrepancies change the denominator for the reported rates (e.g., 13.1% using metrics, 14.2% using frameworks, the 2% modality-skew figure, and the subtype percentages in Fig. 7). Please provide a single consistent PRISMA chart and, ideally, release the 83 search queries and the list of included papers so that the corpus can be reconstructed independently.","section":"Section III-A, Fig. 3"},{"comment":"The meta-analysis percentages rest on the authors' subjective classification of each selected paper into skew subtypes, framework usage, and metric usage, but no inter-annotator agreement, second-annotator validation, or coding rubric is reported. The selection criteria themselves include subjective components (e.g., \"best top-tier conferences/journals,\" QS top-100 university affiliation, citation-count thresholds). Without an included-paper list and a transparent coding protocol, an independent reader cannot verify that the reported prevalences (e.g., 60.3% of label-skew papers not mentioning quantity skew, 13.1% using metrics, 14.2% using frameworks) are stable or representative. This is a load-bearing issue for the paper's quantitative contribution, so please document the classification process and provide reproducibility artifacts.","section":"Section III-B and Section VIII"},{"comment":"The claims of \"complete coverage of all the aspects related to non-IID data\" and of being \"the first technical survey dedicated to organizing and synthesizing the existing knowledge regarding distribution skewness in FL\" are not backed by a transparent comparison protocol. Table I assigns binary or partial checkmarks to prior surveys without specifying the rubric used to determine whether a topic is \"included,\" \"partially included,\" or \"not included.\" Given the corpus inconsistencies described above, these claims are stronger than the current evidence supports. Please either provide a detailed comparison rubric and a reproducible corpus, or soften the claims to reflect coverage of the selected literature rather than complete coverage of the field.","section":"Section I-B, I-C, and Table I"}],"minor_comments":[{"comment":"In the paragraph defining label skew, the text says \"the marginal f(i)X or conditional f(i)Y|X distributions of the labels,\" but the equations and surrounding discussion refer to f(i)Y and f(i)Y|X; the marginal label distribution should be f(i)Y, not f(i)X. Please correct this typo in the formal definition.","section":"Section IV, Eq. (3)"},{"comment":"Equation (1) defines the optimization objective as min_w l(w) := h(L_k(w)), but h is later described as an aggregation function over client objectives. As written, h is applied to a single scalar L_k(w), which is a type mismatch. Please clarify the notation, for example by writing l(w) = h(L_1(w), ..., L_K(w)).","section":"Section II, Eq. (1)"},{"comment":"The Client-Wise Non-IID Index formula appears garbled with OCR artifacts (e.g., the averaging sets and the normalization term are unclear). Please render the equation cleanly and verify the indices in the numerator and the definition of |C_i| versus |C_j|.","section":"Section V-B, Eq. (10)"},{"comment":"Several protocol popularity percentages (e.g., Dirichlet 27%, Sharding 20%, Percentage-of-non-IID-ness 7%) are reported without stating the denominator. Please clarify whether these are percentages of the 235 included papers or of the subset of papers that use partition protocols, and add the corresponding counts.","section":"Section V-A"},{"comment":"The paragraph on FedDyn describes a knowledge-distillation approach with \"focus distillation\" and local differential privacy, but the cited reference [198] appears to be about a federated distillation method for recommender systems. Please verify that the description matches the cited work or replace the reference.","section":"Section VI, FedDyn paragraph"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the scope of the journal and addresses a timely topic. However, the meta-analysis must be made internally consistent and reproducible before publication. I would also ask the editor to consider whether the authors' own tool, FedArtML [128], is given favorable placement in Table II and elsewhere without an explicit conflict-of-interest disclosure; this is not disqualifying, but it should be transparent in a survey that compares frameworks and tools. Finally, before a resubmission is evaluated, the authors should clarify the relationship between the claimed \"complete coverage\" and the relatively small, subjectively selected corpus of 235 papers."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou should know two things about this survey: the taxonomy work is genuinely useful, and the meta-analysis has internal inconsistencies that need fixing before the numbers can be trusted.\n\nWhat's actually new: they add modality skew as a category, give dedicated taxonomies for partition protocols and non-IID metrics, and provide a prevalence meta-analysis across the selected corpus. That's a real organizational contribution over previous surveys. The qualitative discussion of skew interdependencies and the gap analysis (e.g., lack of combined skew metrics) is thoughtful. The paper does well at mapping the solution landscape and comparing frameworks, and the writing is clear.\n\nWhere it gets soft: the PRISMA flow in Section III-A/Fig. 3 doesn't reconcile. The text says 5489 unique papers, the figure says 5580. Eligibility sums to 238 (202+36) but the text says 235 papers included; the figure shows 253 useful and 219+36=255. These discrepancies change the denominator of every percentage they report, including headline numbers like \"13.1% use metrics\" and \"14.2% use frameworks.\" Because the survey's empirical claim rests on this corpus, that's a real, correctable flaw. Also, they don't list the 83 search queries or the included paper list, so an independent reader can't reconstruct the set. The selection criteria (venue ranking, QS university, citation counts) are subjective and could bias the prevalence estimates. Minor point: the authors' own FedArtML is featured in the framework comparison, which isn't disqualifying but deserves a note about independence.\n\nNone of this sinks the taxonomy. The categories and their definitions are grounded in the cited literature, and the interdependency discussion holds up. But the \"complete coverage\" claim in Section I-B is overstated, and the quantitative claims should be softened until the corpus is reproducible.\n\nWho is this for? Anyone working on non-IID FL who wants a map of the field, partition protocols, and metrics. It would be a solid reference once the numbers are fixed.\n\nMy recommendation: send it to peer review. The taxonomy is a contribution worth refereeing, but the referees should require the authors to fix the PRISMA counts, release the paper list and queries, and temper the completeness claim.","headline":"A solid taxonomic survey whose meta-analysis numbers don't add up; fix the PRISMA flow and it's a useful standard reference.","tokens_in":42498,"tokens_out":2370,"would_cite":true,"duration_ms":22447,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey aims to give federated learning a standard taxonomy for non-IID data, and reports that the field rarely quantifies or benchmarks it.","keywords":["federated learning","non-IID data","data heterogeneity","label skew","partition protocols","heterogeneity metrics","modality skew","survey taxonomy"],"falsifier":"Re-running the documented search across the same literature repositories and applying the stated inclusion criteria should reproduce the same 235-paper corpus and the same subtype percentages; the flow diagram already shows a count mismatch (219 + 36 = 255 versus the stated 235), so a corrected count or an independent re-classification that yields materially different percentages would falsify the meta-analytic claims.","tokens_in":41460,"feed_emoji":"🔀","tokens_out":7135,"duration_ms":63832,"temperature":0.7,"pith_summary":"The paper sets out to give the field a standard way to name, quantify, and simulate non-IID data in federated learning. It argues that current research lacks consensus on how to classify or measure data heterogeneity across clients, and that this blocks fair comparison between methods. The survey builds a taxonomy of six skew types, a taxonomy of partition protocols used to create non-IID datasets, and a taxonomy of metrics for measuring non-IIDness, and it reports prevalence statistics over 235 selected papers. If the field adopts these categories, experiments could become more reproducible and benchmark comparisons more meaningful.","feed_headline":"Survey maps every flavor of non-IID data in federated learning","feed_subtitle":"Taxonomy of six skew types shows label skew dominates while only 13% of studies quantify heterogeneity.","key_machinery":"The organizing device is a formal density-based definition of skew: local data distributions differ when $f^{(i)} \\neq f^{(j)}$, with label skew and attribute skew separated through marginal and conditional densities, linked by Bayes' theorem. On top of this, the survey classifies partition protocols (e.g., Dirichlet-based, sharding, noise-based) and heterogeneity metrics (distance/divergence-based, statistical tests, class-based, model-based, encoder-based, performance-based). The quantitative findings come from a 235-paper corpus assembled through a structured screening of six literature databases; prevalence percentages such as label skew appearing in 48-55% of papers and only 13.1% of papers using metrics are the load-bearing outputs that the taxonomy explains.","core_discovery":"The central claim is that non-IIDness in federated learning decomposes into label skew, attribute skew, quantity skew, spatiotemporal heterogeneity, participation skew, and modality skew, and that prior surveys do not cover all of them. The authors assert this is the first technical survey to provide complete coverage of these aspects. They further claim that the literature is heavily skewed toward label skew, that only 13.1% of the analyzed papers quantify non-IIDness with metrics, only 14.2% use standardized FL frameworks, and that no partition protocol or metric simultaneously combines multiple skew types.","pith_inferences":["An implicit corollary is that benchmark suites should be designed around the taxonomy, e.g., a matrix of skew-type combinations, rather than single-skew perturbations.","The taxonomy suggests testable hypotheses, e.g., that methods validated only on label skew will degrade more when combined with quantity skew than methods trained on multi-skew partitions.","If combined metrics were developed, they could also serve as calibration tools for setting the concentration parameters in Dirichlet-based partition protocols, linking simulation degree to measured heterogeneity.","The prevalence statistics likely overstate label skew's practical importance because label-skew partitions are the cheapest to generate; a corpus-level test would be to check whether papers that report accuracy gains under label skew retain those gains under attribute or modality skew."],"forward_implications":["Standardized partition protocols would let researchers state exactly which skew type and degree their experiments realize, making results comparable across papers.","If the 13.1% metric-usage finding is accurate, the field's central claims about non-IIDness are mostly unsupported by quantitative heterogeneity measures.","The absence of combined-skew metrics and combined-skew partition protocols means real-world scenarios with simultaneous label, attribute, and quantity skew are currently under-served.","Modality skew, present in only 2% of papers, is a growth area; frameworks and benchmarks that support multimodal non-IID data could lead practical deployments.","Higher adoption of standardized frameworks (currently 14.2%) would likely improve reproducibility of non-IID federated learning experiments."],"supporting_citations":[{"why":"Supplies the weighted-averaging baseline objective and the decentralized setting that non-IID data perturbs.","marker":"[1]"},{"why":"Earlier survey of non-IID data in federated learning that this paper's taxonomy and coverage claims are positioned against.","marker":"[17]"},{"why":"Foundational study of federated learning with non-IID data; provides the data-sharing solution and the definitional starting point.","marker":"[27]"},{"why":"Experimental study of non-IID data silos that supplies the Dirichlet partition protocol and evidence on label/quantity skew effects.","marker":"[34]"},{"why":"Structured systematic-review methodology that the paper adapts to build its 235-paper corpus.","marker":"[44]"},{"why":"Defines the Dirichlet distribution used by the most common partition protocol.","marker":"[125]"},{"why":"Source for multiple partition protocols and some non-IID metrics, including Hist-Dirichlet and Min-Size-Dirichlet.","marker":"[128]"},{"why":"Federated learning framework whose dataset partitioning utilities ground the framework taxonomy.","marker":"[154]"},{"why":"Proximal-regularized optimization method reported as the most-used solution, anchoring the solution-prevalence analysis.","marker":"[177]"},{"why":"Control-variate method that counters client drift, discussed as a leading alternative solution.","marker":"[178]"}],"fun_headline_variants":["First full taxonomy of non-IID data in federated learning","Label skew dominates federated learning's non-IID challenge","Only 13% of federated learning studies quantify non-IIDness","Six skew types, one taxonomy: non-IID data decoded"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The quantitative findings stand on the assumption that the 235 selected papers fairly represent the non-IID federated learning literature and that the authors' manual classification of each paper into skew subtypes is accurate and consistent.","fun_headline_variants_meta":{"raw":{"variants":["First full taxonomy of non-IID data in federated learning","Label skew dominates federated learning's non-IID challenge","Only 13% of federated learning studies quantify non-IIDness","Six skew types, one taxonomy: non-IID data decoded"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000631,"raw_usage":{"total_tokens":2852,"prompt_tokens":823,"completion_tokens":2029,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":439,"completion_tokens_details":{"reasoning_tokens":1957}},"tokens_in":439,"tokens_out":2029,"duration_ms":15019,"temperature":1.0,"reasoning_tokens":1957,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T17:35:38.575458+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-running the documented search across the same literature repositories and applying the stated inclusion criteria should reproduce the same 235-paper corpus and the same subtype percentages; the flow diagram already shows a count mismatch (219 + 36 = 255 versus the stated 235), so a corrected count or an independent re-classification that yields materially different percentages would falsify the meta-analytic claims.","supporting_citations":[],"review_version":1}