{"id":"3479d34f-6689-4ce3-ac3c-bc70771d8ae1","arxiv_id":"2501.04000","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper reviews 211 federated learning studies across six human sensing domains, assesses them along eight dimensions, and identifies five areas needing urgent research.","lead":"This preprint is a systematic literature review of 211 papers on federated learning applied to human sensing, with a proposed taxonomy and an eight-dimensional assessment of each study. A smart generalist would read it to learn where privacy-preserving machine learning for wearables, activity recognition, and other human-centric sensing actually stands, and which technical gaps remain unresolved.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The corpus-level rankings depend on unvalidated binary codings; without inter-rater reliability or a sensitivity analysis, the 'most analyzed gaps' conclusion is not yet supported.","rationale":"The survey is well structured and the PRISMA-style corpus construction is transparent, which is real supporting evidence. However, the paper's analytical output is quantitative: it ranks FL characteristics by how frequently they are addressed and uses that ranking to motivate five research priorities. That ranking is produced by binary judgments that are never validated. The reader's verdict treated coding reliability as a typical, non-load-bearing survey weakness, but here it is load-bearing because the paper's headline finding is exactly a ranking of those judgments. The risk is concrete: the 'Server-optimized FL' definition is broad enough that default server aggregation could be counted as optimization, and 'Statistical Heterogeneity' could be counted from dataset properties rather than from explicit analysis. A sensitivity analysis or inter-rater reliability study would settle whether the top-two ranking survives reasonable coding variation. If it does, the survey's conclusions stand; if it does not, the conclusion should be rephrased as provisional. This warrants conditional acceptance rather than rejection, since the corpus, taxonomy, and qualitative domain summaries remain valuable regardless of the exact ranking.","tokens_in":49712,"tokens_out":7356,"duration_ms":78221,"concrete_test":"Draw a stratified random sample of 40 papers from the 211, covering all six application domains. Have two independent raters, blinded to the authors' checkmarks, recode all eight dimensions using the Section 2 definitions plus a written rule for N/A, and compute Cohen's kappa per dimension. Then bootstrap the full corpus: for each dimension, flip a checkmark with probability equal to the observed disagreement rate (bounded by the per-dimension 1 - kappa) and recompute the Figure 5 rankings 10,000 times. Additionally recompute the percentages twice, once with N/A counted in the denominator and once with N/A excluded. If the pair {Statistical Heterogeneity, Server-optimized FL} is not the top-two in at least 95% of replicates under both denominator conventions, the headline ranking is not robust to coding subjectivity and should be reported as provisional.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central quantitative claims—that Statistical Heterogeneity and Server-optimized FL are the most frequently addressed characteristics, and that system heterogeneity and privacy are under-addressed—are computed from binary checkmarks in Tables 2–9. These codings are the load-bearing evidence, and their reliability is not established. The paper provides only informal one-sentence dimension definitions, no coding manual, no inter-rater reliability score, and no statement of how no-check versus N/A entries are treated when the Figure 5 percentages are computed. The 'Server-optimized FL' dimension is especially fragile: every FL deployment has a server aggregation step, so without an explicit rule distinguishing default FedAvg from deliberate server-side optimization, the checkmark is easily over-assigned. 'Statistical Heterogeneity' is similarly at risk if the mere presence of non-IID data is coded as addressing the challenge rather than as an explicit robustness analysis. If the two 'most analyzed' dimensions are inflated by permissive coding, the paper's prioritized list of research gaps is not trustworthy. This is a validation gap, not a claim about author intent.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a systematic survey of federated learning (FL) in human sensing, following the PRISMA methodology. It constructs a corpus of 211 papers from three digital libraries and thirteen selected conferences, organizes them into a taxonomy of six application domains (audio/speech processing, well-being, user identification, human mobility and localization, activity recognition, and interface development), and evaluates each study against eight dimensions: privacy and security, communication cost, system heterogeneity, statistical heterogeneity, unlabeled data usage, simplified setup, server-optimized FL, and client-optimized FL. Based on the aggregated eight-dimensional assessment, the paper claims that Statistical Heterogeneity and Server-optimized FL are the most frequently analyzed characteristics, that system heterogeneity and privacy are under-addressed, and it proposes five research areas requiring urgent attention.","tokens_in":49869,"tokens_out":5648,"duration_ms":51850,"significance":"If the corpus and the eight-dimensional codings are reliable, this survey is a valuable contribution: it is the first PRISMA-based systematic mapping of FL in human sensing, provides a reusable taxonomy and per-domain tables (Tables 2–9) that are a useful reference resource, and makes concrete, falsifiable claims about research gaps. The search strategy and exclusion rules are described transparently, which is a notable strength for a survey. However, the central quantitative claims about which FL characteristics are most or least analyzed rest entirely on binary checkmarks that are not validated, and the paper provides no inter-rater reliability, no coding manual, and no sensitivity analysis. The ranking of research gaps and the related recommendations are therefore not yet supported by the evidence presented.","major_comments":[{"comment":"The binary checkmarks for Statistical Heterogeneity and Server-optimized FL are assigned using definitions that do not reliably distinguish explicit consideration from incidental features of the setup. For Statistical Heterogeneity, the text states both that a system 'addresses' the challenge if it 'applies strategies to ensure robustness against non-IID-related attacks' and that the authors 'expect systems using datasets that are inherently heterogeneous to explicitly analyze the impact of such heterogeneity on FL performance.' These criteria can yield different codings for the same paper. For Server-optimized FL, the definition is that 'the server employs methods to enhance model convergence speed and overall performance,' which is satisfied by virtually any FL deployment, including vanilla FedAvg. Without an explicit coding rule (for example, requiring a server-side algorithm beyond default aggregation or an ablation demonstrating the server-side contribution), the checkmarks are easily over-assigned. Since the paper's headline conclusion that these two dimensions are the 'most analyzed' is computed from these checkmarks, the conclusion is not yet supported. I recommend adding a detailed coding manual, reporting inter-rater reliability on a random subset (e.g., Cohen's kappa), and conducting a sensitivity analysis in which the ambiguous dimensions are re-coded under stricter criteria to demonstrate that the ranking in Figure 5 is robust.","section":"Section 2 (dimensions 4 and 7); Tables 2–9; Figure 5"},{"comment":"The treatment of N/A entries in the percentage computations is unspecified. The tables contain N/A in many rows (e.g., Table 2, row [293]; Table 3, row [67]; Table 5, row [223]), and the histograms in Figure 5 show three categories: Consideration, Not Applicable, and No Consideration. The paper does not state whether the denominator for each dimension excludes N/A entries or treats them as 'No Consideration.' If N/A entries are excluded, the percentages for different dimensions are not directly comparable because the denominators differ; if they are included as 'No Consideration,' the percentage is distorted. Please specify the exact formula used and justify the treatment; this is essential for interpreting the central claim about which dimensions are most frequently analyzed.","section":"Section 10, Figure 5"},{"comment":"The exclusion criteria for irrelevance are described only in prose, with no operational definitions or reliability check. The categories 'non-application papers' and 'application papers in other fields' require subjective judgments, particularly the boundary between human sensing and adjacent areas such as IoT and medical research. The manuscript does not report how many papers were excluded under each criterion, nor whether screening was performed independently by multiple reviewers. Because the final corpus of 211 papers is the evidence base for all meta-conclusions, the reproducibility of the screening should be documented—for example, with a PRISMA flow diagram that includes exact counts per exclusion reason and a dual-screening protocol with disagreement resolution, or at least a sensitivity analysis with alternative inclusion rules.","section":"Section 3.3, Exclusion for Irrelevance"},{"comment":"The statement that 'only a minority achieve the reliability necessary for deployment in practical settings' is not derived from the eight-dimensional assessment in a transparent way. No composite metric or threshold is defined that maps the binary checkmarks onto a reliability judgment, and the paper does not say how the eight dimensions are combined (e.g., whether all eight must be satisfied, or a subset). Please specify how the dimensions are combined to determine deployment readiness, or soften the conclusion to match what the presented data can support.","section":"Section 11, Conclusion"}],"minor_comments":[{"comment":"Reference [38], cited for the ACM Computing Classification System, is listed as 'Generate Code. 1998'; this appears to be an artifact, and it should be corrected to the proper ACM CCS citation.","section":"Section 3.1, reference list"},{"comment":"The application label 'S,FER' is unclear; it should be written as 'SER/FER' or 'Multi' to match the study's use of both speech and facial inputs.","section":"Table 3, row for [36]"},{"comment":"The PRISMA flow diagram is referenced but the actual numbers at each stage (identification, screening, exclusion, inclusion) are not visible in the manuscript text; the final version should include the full flow diagram with counts per exclusion reason to support the reproducibility claim.","section":"Section 3.1, Figure 3"},{"comment":"The statement that 'system heterogeneity is the least addressed attribute in the corpus' is presented without an explicit comparison to other under-addressed dimensions such as privacy and unlabeled data usage; please clarify whether this refers to the overall corpus and show the underlying percentages from Figure 5.","section":"Section 10.2"},{"comment":"The discussion of statistical heterogeneity would benefit from distinguishing data skewness in features, labels, and temporal distributions, since the coding rule currently conflates these forms of non-IID data.","section":"Section 2, dimension 4"},{"comment":"There are several minor typographical errors, including 'collaoborative' in Section 7.2 and 'ZHuang' in reference [312]; these should be corrected in a final pass.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The survey is timely and well-structured, and the authors have made a serious effort at methodological transparency. However, the central quantitative claims about research gaps depend on unvalidated binary codings, and the manuscript currently lacks the validation evidence (inter-rater reliability, coding manual, sensitivity analysis) needed to support them. I would be willing to accept after the authors address the validation gap and clarify the N/A handling in Figure 5. A minor concern for the editor: the claim of a 'comprehensive corpus' may draw criticism because the search is limited to three digital libraries and a selected set of conferences; the authors should either temper this claim or provide stronger evidence of coverage."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a genuinely useful survey, and the main caveat is isolated to one methodological step. It fills a real gap: prior FL surveys cover IoT, healthcare, and recommender systems, but not human sensing as defined here. The corpus of 211 papers, the raw-data taxonomy, and the six-domain organization are new organizational tools. The PRISMA-style search is described in detail, with explicit venues and exclusion rules; that is more transparent than most surveys in this space.\n\nThe authors are honest about the state of the field. They explicitly note that most studies use simplified setups and that only a minority approach deployment-level reliability. The discussion of open challenges is sensible, and the five research directions follow from the analysis rather than being bolted on.\n\nThe soft spot is the eight-dimensional assessment. The binary checkmarks in Tables 2–9 are the basis for the quantitative claims about which FL characteristics are most and least analyzed. There is no coding manual, no inter-rater reliability score, and no statement of how N/A versus no-check entries are treated when computing the Figure 5 percentages. The 'Server-optimized FL' dimension is particularly fragile: every FL deployment has a server aggregation step, so without an explicit rule to distinguish deliberate server-side optimization from default FedAvg, those checkmarks could be over-assigned. 'Statistical Heterogeneity' carries a similar risk if the mere presence of non-IID data is coded as addressing the challenge rather than as an explicit robustness analysis.\n\nThat weakens the specific ranking—that Statistical Heterogeneity and Server-optimized FL are the most frequently analyzed characteristics. The qualitative pattern, though, is consistent with what I know of the field: privacy and system heterogeneity are under-addressed, while non-IID data gets routine attention. So the overall conclusion is plausible; the issue is a validation gap, not evidence of carelessness.\n\nWho is this for: anyone working on FL for wearable, audio, video, or location-based sensing, plus researchers looking for open problems in this intersection. It deserves peer review. A referee should ask for coding reliability or a sensitivity analysis, but that is a revision request, not a rejection.\n\nI would send it to review and likely accept after minor revisions. I would also cite it if I were writing in this area.","headline":"A useful, transparent survey of FL in human sensing whose quantitative gap analysis needs a coding reliability check before the specific rankings are cited.","tokens_in":50429,"tokens_out":2530,"would_cite":true,"duration_ms":24227,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A systematic review of 211 federated-learning studies in human sensing finds that most papers ignore the field's hardest real-world problems, and only a minority demonstrate the reliability needed for practical deployment.","keywords":["federated learning","human sensing","survey","taxonomy","eight-dimensional assessment","privacy","system heterogeneity","activity recognition"],"falsifier":"Re-code the same 211 papers with an independent coder (or with the authors plus a second coder) using a written coding manual; if the per-dimension consideration rates differ materially for privacy/security or system heterogeneity, the survey's ranking of research gaps is not reproducible. Alternatively, compile all FL-human-sensing papers published before November 2023 that the search missed and check whether they predominantly address system heterogeneity.","tokens_in":49533,"feed_emoji":"📡","tokens_out":4068,"duration_ms":37069,"temperature":0.7,"pith_summary":"Federated learning promises to train machine-learning models on privacy-sensitive human-sensing data—activity, speech, physiological signals, location—without moving raw data off users' devices. This survey asks how far that promise has actually been kept. It assembles a corpus of 211 studies across six application domains and scores each against eight challenge dimensions (privacy and security, communication cost, system versus statistical heterogeneity, use of unlabeled data, simplified setups, and server- versus client-oriented optimization). The core finding is that the field has concentrated on the easiest-to-simulate problems—statistical heterogeneity and server-side model quality—while privacy defenses, heterogeneous device participation, and learning from unlabeled data remain comparatively neglected, so that only a minority of studies meet the reliability bar for real deployment. If the survey's map is right, it gives practitioners a prioritized list of where to put effort next.","feed_headline":"Most federated sensing papers skip real-world hurdles","feed_subtitle":"Review of 211 studies finds privacy, device diversity, and unlabeled data overlooked.","key_machinery":"The carrying object is the eight-dimensional assessment: each surveyed paper is coded with a binary checkmark for whether it addresses privacy and security, communication cost, system heterogeneity, statistical heterogeneity, unlabeled data usage, simplified setup, server-optimized FL, or client-optimized FL. The dimension definitions operationalize each challenge (e.g., privacy counts only if the paper adds defenses such as differential privacy or secure aggregation or runs a vulnerability analysis; system heterogeneity counts only if training or deployment involves heterogeneous devices). Layered on top of a taxonomy of six application domains and nine raw-data types, this coding lets the survey turn a heterogeneous literature into comparable percentages and a ranking of research gaps.","core_discovery":"The paper's central claim is that, assessed across eight dimensions, current federated-learning research in human sensing is lopsided: statistical heterogeneity and server-optimized federated learning are the characteristics analyzed most frequently, whereas privacy and security, communication cost, system heterogeneity, and unlabeled-data usage receive far less attention. Across the six application domains—audio and speech processing, well-being, user identification, human mobility and localization, activity recognition, and interface development—activity recognition and well-being dominate the corpus. The survey concludes that only a minority of the reviewed studies achieve the reliability necessary for deployment in practical settings, and it identifies five aspects needing urgent research: privacy and security under active attacks, unlimited participation across heterogeneous devices, exploiting unlabeled data in the wild, clarifying whether the primary target is the server or the clients, and moving beyond simplified experimental setups.","pith_inferences":["The survey's own evidence suggests its five priority directions could be re-derived as a coverage gap: the three dimensions with the lowest checkmark rates (system heterogeneity, privacy defenses, unlabeled-data usage) map almost one-to-one onto the first three urgent research aspects; one testable extension is to recompute the dimension-by-dimension percentages after a re-coding with a formal cod","A second extension: a prospective author could use the eight dimensions as a submission checklist, turning the survey into a de facto evaluation rubric for new FL-human-sensing papers.","If the claim that only a minority are deployment-ready is taken literally, a natural next study is to define an explicit deployment-readiness threshold (e.g., a minimum number of satisfied dimensions plus a real-device evaluation) and measure what fraction of the 211 papers pass; the survey does not itself formalize that threshold."],"forward_implications":["Practitioners selecting an FL approach for a human-sensing product should expect that most published baselines have not hardened privacy against inference or poisoning attacks, so production systems need additional defenses.","“Unlimited participation” is the paper's label for a concrete open problem: most studies assume homogeneous, well-provisioned clients, so stragglers and low-end devices are implicitly excluded; the survey argues this must change for worldwide deployment.","Because the corpus is weighted toward activity recognition and well-being, the five urgent research directions apply with different force per domain; interface development, with the fewest studies, is the least explored.","The simplified-setup dimension is treated as an undesired feature, and its prevalence implies reported accuracies may overstate real-world performance.","Researchers should treat the paper's priority list as a portfolio recommendation: privacy under attacks, system heterogeneity, unlabeled data, and server/client target clarity."],"supporting_citations":[{"why":"McMahan et al. define federated learning and FedAvg, the object the survey tracks across human-sensing applications.","marker":"[160]"},{"why":"Moher et al. supply the PRISMA methodology that structures the survey's corpus construction and screening steps.","marker":"[171]"},{"why":"Li et al. provide the general framing of FL challenges that motivates the eight assessment dimensions.","marker":"[139]"},{"why":"Zhao et al. establish the non-IID data problem that the survey uses to define the statistical heterogeneity dimension.","marker":"[300]"},{"why":"Jin et al. survey FL without full labels, the basis for the unlabeled-data usage dimension and its under-exploration claim.","marker":"[118]"},{"why":"Bonawitz et al. describe system design constraints and device heterogeneity, the reference behind the system heterogeneity dimension.","marker":"[18]"}],"fun_headline_variants":["Federated sensing studies overlook privacy and device diversity","FL for human sensing: statistical heterogeneity overstudied, security ignored","Survey finds federated sensing research lopsided on real-world hurdles","Federated learning in sensing: privacy and unlabeled data gaps persist"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The percentages and gap rankings rest on the assumption that the search and exclusion criteria produced a representative corpus of FL-in-human-sensing papers and that the binary checkmarks in the eight-dimensional assessment were applied consistently, and the paper does not report an inter-rater reliability check or a detailed coding manual.","fun_headline_variants_meta":{"raw":{"variants":["Federated sensing studies overlook privacy and device diversity","FL for human sensing: statistical heterogeneity overstudied, security ignored","Survey finds federated sensing research lopsided on real-world hurdles","Federated learning in sensing: privacy and unlabeled data gaps persist"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000199,"raw_usage":{"total_tokens":1370,"prompt_tokens":941,"completion_tokens":429,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":557,"completion_tokens_details":{"reasoning_tokens":356}},"tokens_in":557,"tokens_out":429,"duration_ms":5129,"temperature":1.0,"reasoning_tokens":356,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:40:25.366904+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-code the same 211 papers with an independent coder (or with the authors plus a second coder) using a written coding manual; if the per-dimension consideration rates differ materially for privacy/security or system heterogeneity, the survey's ranking of research gaps is not reproducible. Alternatively, compile all FL-human-sensing papers published before November 2023 that the search missed and check whether they predominantly address system heterogeneity.","supporting_citations":[{"cited_title":"Salman Avestimehr","cited_arxiv_id":null,"evidence_quote":"Zhao et al. establish the non-IID data problem that the survey uses to define the statistical heterogeneity dimension."}],"review_version":1}