{"id":"09b1e7a1-2607-4c15-b545-00e9d606fabf","arxiv_id":"2412.10919","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"Federated survival models match or beat locally trained models for hemodialysis patient survival prediction in most of NephroPlus's six Indian regions.","lead":"This paper tests whether federated learning, a privacy-preserving training method, can predict survival of kidney dialysis patients as well as or better than models trained at each center alone. Using data from over 24,000 patients across six regions of India's largest dialysis network, it finds federated models win for most regions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"FedSurf's tree-selection step (§III.C.c) may use the same test split used to report C-indices; if so, its 4/6-zone wins are circular and the central FL-outperforms-local claim is unsupported.","rationale":"I read the manuscript in good faith. The intended contribution is empirical: show that FL, and specifically FedSurf, yields survival models at least as accurate as local models on NephroPlus data. The reader's conditional verdict is reasonable: Table III is absent, inclusion criteria are undefined, and no error bars or central baseline are supplied. My stress-test focuses on a different, more internal threat. The FedSurf aggregation description in §III.C.c is the only place where the winning model's construction is described, and it is ambiguous about the data used for tree selection. If the 20% test split defined in §IV is used to rank 'top-performing' trees, the global model is effectively tuned to the test labels; the subsequent C-index comparison is circular. This is load-bearing because FedSurf is the best model in four of six zones; without those wins, the central claim reduces to 'FL is comparable,' a much weaker statement that would not justify the paper's recommendation. The check I propose directly tests this by rerunning selection restricted to training/OOB data versus test-contaminated selection. I do not claim the authors cheated; the text is simply underspecified. Because the flaw is correctable by re-running the experiments and reporting the selection protocol, the appropriate verdict remains CONDITIONAL, matching the reader's verdict; I would add the FedSurf selection protocol as an explicit condition for acceptance.","tokens_in":7409,"tokens_out":8794,"duration_ms":81569,"concrete_test":"Re-run the FedSurf pipeline with the same 80/20 splits under two variants: (A) select and weight trees using only each client's training split or OOB estimates, then freeze the global forest and evaluate once on the untouched test split; (B) select trees using the test split, i.e., the suspicious variant. Compare FedSurf versus local RSF per client across 10 random split seeds. If variant A no longer yields FedSurf wins in as many zones as variant B, the reported 4/6 wins are an artifact of test-set-contaminated tree selection. Report per-client tree counts and the exact 'performance' metric used.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that federated survival models, especially FedSurf, achieve comparable or better concordance than local models—rests on the fairness of the C-index comparison in Table III. The weakest load-bearing condition is that model construction never touches the held-out test split. Section III.C.c describes FedSurf aggregation as: 'Each client's decision trees were sorted based on performance. A central server determined the number of trees it required from each client based on the importance assigned to each of them. The top-performing trees were taken from each client's forest.' No dataset is named for this 'performance' ranking. Since Section IV defines the test split as a local 20% held-out set, if that split is used to sort or select trees, the global forest is constructed with knowledge of test outcomes and the subsequent C-index is optimistically biased. FedSurf is the model reported to win in North, East, West, and Andhra Pradesh; a selection artifact there would remove the main empirical support for FL's advantage. Even if the ranking is training-only, the number of trees per client and the 'importance' weighting are unspecified, so the reported win pattern cannot be independently reproduced or distinguished from configuration choice.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper applies federated learning (FL) to survival analysis for hemodialysis patients using data from NephroPlus, India's largest private dialysis network, with six geographic zones treated as clients. Four survival models (CoxPH, DeepSurv, Cox-nnet, Random Survival Forests) are trained locally and in federated variants (coefficient averaging, FedAvg, and a Federated Survival Forest, FedSurf), and compared by Harrell's concordance index on an 80/20 train/test split. The central claim is that federated survival models achieve comparable or better performance than locally trained models, with FedSurf being the best-performing model in four of six zones, and that this is the first application of FL to a large real-life hemodialysis dataset.","tokens_in":7608,"tokens_out":2821,"duration_ms":26371,"significance":"If the results are fully supported, the paper would provide a useful empirical demonstration that privacy-preserving federated training can match or exceed local survival models on a large, real-world, multi-center clinical dataset. The use of 183,063 patient records from 244 centers and the realistic client structure (six regional zones) are strengths, as is the comparison of four distinct model families with standard federated aggregation approaches. The study is directly relevant to applied ML for healthcare and to the operational question of whether dialysis networks should adopt FL. However, the significance is currently limited because the quantitative evidence for the central claim is incomplete: the actual C-index results are not visible in the manuscript, no uncertainty quantification or significance testing is reported, and a key methodological detail in the FedSurf tree-selection procedure could permit test-set leakage if not clarified.","major_comments":[{"comment":"The inclusion criteria that reduced the cohort from 183,063 patients to 24,052 patients are never described. This is load-bearing because every subsequent client-level comparison depends on the representativeness and pre-specification of this subsample. The authors should state the exact eligibility criteria, the number of patients excluded at each step, and whether the criteria were defined before any outcome analysis. If the criteria were chosen post hoc or are correlated with survival, the entire comparison between local and federated models could be invalidated.","section":"III.A"},{"comment":"The actual concordance index values are not present in the version under review: Table III is declared but its contents are missing. Since the paper's central claim—that federated models, especially FedSurf, match or outperform local models—rests entirely on this table, the quantitative basis for the claim cannot be assessed. In addition, even when the values are supplied, the authors should report error bars, confidence intervals, or significance tests (e.g., permutation or bootstrap tests on the C-index), because without such measures the observed differences may be within noise.","section":"IV, Table III"},{"comment":"The FedSurf tree-selection step is not specified with respect to which data are used for the 'performance' sorting. Section III.C.c says each client's decision trees were 'sorted based on performance' and the top-performing trees were taken, but it does not state whether this performance is evaluated on the local training split or on the 20% held-out test split. If the test split is used at any point in selecting or weighting trees, the reported C-index is circular and the 4/6-zone advantage of FedSurf is unsupported. The authors must clarify that only training data are used for tree selection and provide the exact importance-weighted selection procedure, including how the number of trees per client is determined.","section":"III.C.c"},{"comment":"The claim that 'each client has at least 1 Federated Learning model outperforming its corresponding local model' and the subsequent statement that FedSurf performs best in North, East, Andhra Pradesh, and West are not accompanied by the underlying per-client, per-model C-index values. Moreover, Table IV labels Andhra Pradesh as 'Federated & Local' while the text says FedSurf was the best model for Andhra Pradesh; these statements should be reconciled. The reader currently cannot verify any of the aggregate claims without the missing results table.","section":"IV"}],"minor_comments":[{"comment":"The abstract contains informal language ('don't perform as well') that should be corrected to formal academic style.","section":"Abstract"},{"comment":"The affiliation 'Massachusets Institute of Technology' is misspelled; it should be 'Massachusetts Institute of Technology'.","section":"Author block"},{"comment":"The text says weights were aggregated using 'equation (1)', but the federated averaging formula is equation (2). The equation reference should be corrected.","section":"III.C.b"},{"comment":"Equation (2) has a notational issue: 'N ∑ n_k' is awkwardly typeset and should be written as N = Σ_k n_k, with the aggregation formula clearly defined.","section":"II.B"},{"comment":"The phrase 'the baseline hazard λ0(t) has been parametrized by the cox-proportional hazards model' is inaccurate; in the standard Cox model the baseline hazard is left unspecified and only the log-risk function is parameterized. The text should be reworded.","section":"III.B"},{"comment":"Figure 2 is referenced but not described in the text; the authors should explain what the bar charts show (e.g., exact C-index values) and why RSF/FedSurf were selected for this display.","section":"IV"},{"comment":"The conclusion states that 'feature importances generated by RSF models differed from client-to-client' and that DeepSurv's weaker federated performance 'was likely due to the fact that the data was very heterogeneous', but no supporting results or analyses are provided; these statements should either be removed or supported with evidence.","section":"V"}],"recommendation":"major_revision","confidential_remarks":"The manuscript has a useful real-world dataset and a clear practical motivation, but in its current form the central quantitative evidence is missing from the visible text (Table III), and the FedSurf selection procedure could introduce leakage. These issues are fixable in a revision, so I recommend major revision rather than rejection. The authors should be asked to include the complete results table, add uncertainty quantification, document the inclusion criteria, and specify the data used in the FedSurf tree-selection step."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick read on arXiv:2412.10919. The genuinely new thing here is the dataset and setting: a nationwide Indian dialysis network (NephroPlus), 24k patients across six zones treated as FL clients, with four survival models compared local vs federated. That is a useful first step for FL in Indian healthcare, and the authors are upfront about the sparsity of prior work. What they do well: they pick standard models (CoxPH, DeepSurv, Cox-nnet, RSF), use a sensible client construction, and report per-client results. The claim of 'first FL on a large real-life hemodialysis dataset' is plausible.\n\nThe paper, as I have it, is not close to publishable. The central evidence, Table III with the actual C-indices, is missing from the manuscript. The text cites it but doesn't render it, so I cannot check whether FL indeed matches or beats local models. The second issue is the FedSurf aggregation step: §III.C.c says trees are 'sorted based on performance' with no dataset specified. If that ranking uses the same held-out test split used for the reported C-indices, the 4/6-zone wins for FedSurf are circular. I can't prove it from the text, but the paper must state whether the tree ranking uses training, validation, or test data. Third, the cohort drops from 183,063 to 24,052 patients with no description of inclusion criteria; if those exclusions are post hoc or outcome-related, the whole comparison is biased. Fourth, no confidence intervals, error bars, or significance tests around any C-index; the differences they discuss could be noise.\n\nThe reader's take is about right, and the stress-test note is a real concern, though not proven. I'd soften it to: the missing specification makes the FedSurf result unverifiable, not necessarily circular.\n\nWho is this for? Readers interested in practical FL deployments in healthcare, and anyone teaching common pitfalls in applied ML reporting. It deserves a serious referee if the authors complete the manuscript: add Table III, define inclusion criteria, specify the FedSurf tree selection rank data, add a centralized baseline, and report uncertainty. With those, the empirical question is worth answering. As it stands, I would not cite it, but I would not desk-reject it either; I'd send it back for major revision and a full re-review.","headline":"A real-world FL survival application with an unverifiable core result: Table III is missing and FedSurf's tree-selection step may leak the test split.","tokens_in":8125,"tokens_out":2445,"would_cite":false,"duration_ms":22602,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Federated learning matches or beats local survival models for most dialysis zones.","keywords":["federated learning","survival analysis","hemodialysis","concordance index","random survival forests","privacy-preserving machine learning","India dialysis network"],"falsifier":"A re-analysis of the full 183,063-patient records with a pre-specified, documented exclusion rule that either eliminates the federated model's advantage or reverses the zone-level winners would overturn the central claim; an independent held-out validation cohort on which FedSurf's concordance index no longer beats local random survival forests would also falsify the claimed benefit.","tokens_in":7183,"feed_emoji":"🩺","tokens_out":6440,"duration_ms":53178,"temperature":0.7,"pith_summary":"The paper tries to show that federated learning (training a shared model by exchanging model updates, not patient data) can predict survival for hemodialysis patients as accurately as, and sometimes better than, models trained on each region's data alone. Using records from 24,052 patients across six zones of a nationwide Indian dialysis network, it compares four survival-analysis models in local and federated forms. The federated version of random survival forests wins the concordance-index comparison in four of the six zones, and every zone has at least one federated model that beats its local counterpart. The authors claim this is the first application of federated learning to a large real-life hemodialysis dataset. If the result holds, dialysis networks can improve or maintain predictive accuracy while keeping sensitive patient data within each center.","feed_headline":"Federated model tops local survival scoring in 4 of 6 dialysis zones","feed_subtitle":"Dialysis centers can keep records private and still match—or beat—zone-level survival predictions.","key_machinery":"The paper's central mechanism is the pairing of four survival-analysis models with straightforward federation rules. For Cox proportional hazards, the global model is built by averaging local beta coefficients weighted by each client's sample size. For the deep networks DeepSurv and Cox-nnet, Federated Averaging aggregates model weights across clients. For random survival forests, the global model is assembled by importance-based sampling of trees from each client's local ensemble, forming a federated survival forest. The evaluation metric is Harrell's concordance index, which measures how well predicted risk ranks actual survival times.","core_discovery":"On the paper's own terms, the central discovery is that a federated survival model can be a strict performance improvement for most clients: across the six NephroPlus zones, the federated survival forest (FedSurf) produced the highest concordance index in the North, East, West, and Andhra Pradesh zones, while a locally trained random survival forest won in the South and a locally trained Cox model won in Bihar. Additionally, every one of the six clients had at least one federated model outperform its local counterpart. The paper interprets this as evidence that regional dialysis clients have an incentive to join a federated framework not only for privacy and exposure to more diverse data but also for improved accuracy.","pith_inferences":["A direct consequence the paper does not test: a centralized model trained on the pooled 24,052 records would quantify exactly how much accuracy is lost or gained for the privacy protection FL provides; without that baseline, 'comparable to local' and 'comparable to centralized' are different claims.","If the federated-survival advantage generalizes, the same zone-based setup could be used for live waitlist prioritization, where survival predictions influence transplant wait times; a prospective study would be needed to confirm.","The uneven zone sizes suggest the FedSurf advantage may depend on client heterogeneity; testing with balanced zone partitions or adding differential privacy noise would reveal whether the result survives stronger privacy guarantees.","Because the paper compares FL only against local models, not against a centralized model, the first federated-application claim is best read as 'privacy-compatible accuracy is viable,' not as 'FL exceeds central training.'"],"forward_implications":["Regional dialysis authorities can adopt federated survival forests and expect concordance indices that match or exceed zone-local models for most regions, while patient records never leave the center.","Because four of six zones had a federated model as their single best scorer, accuracy incentives align with privacy incentives for a majority of clients.","Every zone had at least one federated model beat its local counterpart, so even clients whose best model is local can find a federated alternative with comparable or better ranking performance.","The client-level feature-importance differences reported for random survival forests imply that federated models absorb regional heterogeneity rather than forcing one national pattern on every center."],"supporting_citations":[{"why":"Introduces FedAvg, the weighted-averaging rule used to federate DeepSurv and Cox-nnet.","marker":"[8]"},{"why":"Defines the Cox proportional hazards model whose partial likelihood and beta coefficients are used by the local and federated CoxPH baselines.","marker":"[17]"},{"why":"Supplies DeepSurv, the neural-network survival model adapted here as a federated client model.","marker":"[20]"},{"why":"Supplies Cox-nnet, the neural network designed for electronic health records that is federated in this study.","marker":"[21]"},{"why":"Introduces random survival forests, the local model that the federated survival forest (FedSurf) extends and outperforms.","marker":"[22]"},{"why":"Provides the naive global parameters averaging approach used to combine local CoxPH beta coefficients.","marker":"[23]"},{"why":"Provides the importance-based tree sampling method used to aggregate local survival forests into a federated ensemble.","marker":"[24]"},{"why":"Demonstrates federated survival analysis with Cox models, establishing the comparison point the paper extends to real hemodialysis data.","marker":"[10]"}],"fun_headline_variants":["Federated model beats local in 4 of 6 dialysis zones","Privacy-preserving survival model wins most regions","Federated learning improves dialysis survival prediction in most zones","Dialysis survival: federated model tops local in 4 of 6 areas","Federated approach outperforms local survival models for dialysis"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The patient inclusion criteria that shrank the cohort from 183,063 to 24,052 are never described; the study assumes they were fixed in advance and unrelated to outcomes.","fun_headline_variants_meta":{"raw":{"variants":["Federated model beats local in 4 of 6 dialysis zones","Privacy-preserving survival model wins most regions","Federated learning improves dialysis survival prediction in most zones","Dialysis survival: federated model tops local in 4 of 6 areas","Federated approach outperforms local survival models for dialysis"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000147,"raw_usage":{"total_tokens":1127,"prompt_tokens":829,"completion_tokens":298,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":445,"completion_tokens_details":{"reasoning_tokens":212}},"tokens_in":445,"tokens_out":298,"duration_ms":3444,"temperature":1.0,"reasoning_tokens":212,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T15:27:59.797357+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A re-analysis of the full 183,063-patient records with a pre-specified, documented exclusion rule that either eliminates the federated model's advantage or reverses the zone-level winners would overturn the central claim; an independent held-out validation cohort on which FedSurf's concordance index no longer beats local random survival forests would also falsify the claimed benefit.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies Cox-nnet, the neural network designed for electronic health records that is federated in this study."},{"cited_title":"Addressing Data Heterogeneity in Federated Learning of Cox Proportional Hazards Models","cited_arxiv_id":"2407.14960","evidence_quote":"Provides the naive global parameters averaging approach used to combine local CoxPH beta coefficients."},{"cited_title":"Federated Survival Forests","cited_arxiv_id":"2302.02807","evidence_quote":"Provides the importance-based tree sampling method used to aggregate local survival forests into a federated ensemble."}],"review_version":1}