{"id":"d84d48fc-d1b6-4d22-8d07-1879f0c7ea66","arxiv_id":"2412.13779","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"FedSSI combines synaptic intelligence with a locally trained surrogate model anchored to the global model, and reports state-of-the-art accuracy for rehearsal-free continual federated learning under non-IID data.","lead":"A team proposes FedSSI, a way to train federated models on streaming tasks without storing old data, by borrowing a continual-learning trick called synaptic intelligence and adding a personalized model to handle uneven data across devices. The method reports accuracy gains of up to about 12 percentage points over prior continual federated learning approaches on image benchmarks.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Eq. (2) uses α for the SI penalty while experiments use α for data heterogeneity, and the SI-penalty strength is never reported; the FedSSI-vs-FL+SI gains in Table 2 may reflect regularization-strength tuning rather than the PSM mechanism.","rationale":"I read the paper as primarily an empirical contribution: FedSSI is claimed to make SI robust to non-IID CFL by computing importances from a personalized surrogate model that mixes local and global information. The single most load-bearing condition is therefore that the reported gains over FL+SI and other baselines are caused by that mechanism, not by an uncontrolled regularization hyperparameter. The notation collision in Eq. (2) and Section 4.2/5.1 makes this condition insecure: the same symbol α names both the SI penalty and the Dirichlet heterogeneity, and no separate SI-weight value is reported anywhere in the main text or appendices. Combined with the per-dataset selection of λ in Table 3, the experiment cannot currently distinguish 'better PSM importances' from 'better-tuned SI penalty strength.' Proposition 1 does not resolve this because it analyzes Eq. (7), a different objective, under strong convexity, and its λ→0 conclusion is about the PSM's own local-global trade-off, not about the Eq. (2) penalty or the Dirichlet α. I do not claim the method is wrong; I claim the reported evidence underdetermines the mechanism. A matched-penalty sweep is a direct, cheap check. This does not change the reader's CONDITIONAL verdict, since the reader already flagged missing hyperparameters and per-dataset λ tuning; my concern sharpens that flag into a specific confound between the SI penalty coefficient and the heterogeneity parameter. I also agree with the reader that the convergence theorem is imported rather than derived here, but that is secondary to the empirical confound. No credibility or intent concerns are raised; this is an experimental-control issue.","tokens_in":22769,"tokens_out":7301,"duration_ms":71877,"concrete_test":"Re-run the CIFAR10 α=0.1 and Office-Caltech rows of Table 2 for both FedSSI and FL+SI with the SI penalty coefficient swept over {0.01, 0.1, 1, 10, 100}, keeping λ fixed at the paper's chosen value for FedSSI and reporting both best-per-method and matched-penalty accuracies. If FedSSI remains ahead under matched SI penalty strengths, the PSM mechanism is supported; if the advantage collapses or reverses, the headline gain is a tuning artifact rather than evidence for the synergistic importance estimates.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing step is the empirical claim that FedSSI's PSM-derived importance estimates, rather than ordinary SI with a better-tuned penalty, explain the Table 2 gains. Eq. (2) defines the SI penalty coefficient as α, but Section 4.2 and Section 5.1 reuse α for the Dirichlet heterogeneity parameter. The paper never reports the SI penalty strength used for FedSSI or FL+SI, nor whether it is held constant when α is varied. FedSSI also adds a second tuned hyperparameter λ (Eq. 5, Table 3). Under these conditions, the final-accuracy advantages over FL+SI (e.g., +3.3 on CIFAR10, +4.7 on Tiny-ImageNet, +10.8 on Office-Caltech) are not uniquely attributable to the personalized surrogate model; a more favorable SI penalty coefficient alone could produce such gaps. The theory does not close this gap: Proposition 1 only describes how the PSM objective in Eq. (7) moves between local and global solutions for a strongly convex f, and it says nothing about the SI penalty in Eq. (2) or about the Dirichlet α. Theorem 1 is imported from Li et al. 2024a and assumes convergence of the global model, which is precisely what is at issue in non-IID CFL. Thus the paper's central mechanism is not isolated from regularization-strength and λ tuning.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes FedSSI, a rehearsal-free regularization method for continual federated learning (CFL) that adapts Synaptic Intelligence (SI) to non-IID data. The key idea is a personalized surrogate model (PSM) per client, trained on the client's previous local task with a proximal pull toward the global model; the PSM's parameter contributions are used to compute SI importance weights that penalize changes to important weights when learning new tasks. The authors first show that standard regularization methods, especially FL+SI, work well under IID client data but degrade under non-IID partitions, and then argue that FedSSI's PSM restores performance. Experiments on six datasets under Class-IL and Domain-IL scenarios compare FedSSI against FedAvg, FedProx, regularization baselines (FL+LwF, FL+EWC, FL+OGD, FL+SI), and CFL baselines (Re-Fed, FedCIL, GLFC, FOT, FedWeIT), reporting final and average accuracy, sensitivity to the Dirichlet heterogeneity parameter, communication-round efficiency, and resource overhead. The paper claims up to 12.47% improvement in final accuracy over state-of-the-art methods and includes a short analytical section on the PSM.","tokens_in":23113,"tokens_out":5620,"duration_ms":50432,"significance":"If the central claim holds, FedSSI would be a valuable rehearsal-free CFL method that avoids memory and privacy costs of replay while addressing data heterogeneity. The experimental protocol is broad: six datasets, two incremental scenarios, multiple data-heterogeneity levels, and many baselines, with both final and average accuracy reported. The paper also reports communication-efficiency trade-offs and computational overhead, which is useful for practitioners. However, the significance is conditional: the unreported SI penalty strength and the lack of an ablation isolating the PSM mechanism prevent the current results from being uniquely attributed to the proposed method. The analytical section is largely imported from prior work (Hanzely & Richtarik 2020; Li et al. 2024a) and does not establish the paper's main claimed α–λ relationship.","major_comments":[{"comment":"The SI regularization coefficient α in Eq. (2) is never reported for any method, while the same symbol α is used in Section 5.1 as the Dirichlet heterogeneity parameter. This makes the comparison between FedSSI and FL+SI in Table 2 confounded: the gains (e.g., +3.26 on CIFAR10, +4.69 on Tiny-ImageNet) could be due to a more favorable SI penalty strength rather than the PSM mechanism. Please report the exact SI penalty coefficients used for FL+SI and FedSSI for every dataset, and include an ablation in which the SI penalty is held fixed while the PSM is removed or replaced.","section":"Section 5.1, Eq. (2), Table 5"},{"comment":"The claimed positive correlation between the heterogeneity parameter α and the PSM balance parameter λ is not established by the provided analysis. Proposition 1 only shows that as λ→0 the PSM in Eq. (7) approaches the global model for a strongly convex objective, and it does not involve the Dirichlet α or the SI penalty in Eq. (2). Theorem 1 is imported from Li et al. 2024a and assumes the global model converges to the optimum at rate g(t), which is precisely the condition at issue in non-IID CFL, and no proof or independent verification is provided. Table 3 gives empirical evidence for only three λ values. Thus the paper's theoretical support for the central mechanism is insufficient.","section":"Section 4.3, Proposition 1, Theorem 1"},{"comment":"The paper does not provide an ablation that isolates the contribution of the PSM over ordinary SI. The PSM in Eq. (5) is trained only on local previous-task data with a proximal pull to the global model, and the paper asserts that this yields importance estimates reflecting both local and global distributions. Since FedSSI adds both the PSM and a second tuned hyperparameter λ, a comparison against plain SI with the same SI penalty strength but with importance computed from the global model, or against FedSSI with the PSM replaced by a global-mean importance estimate, is needed to show that the PSM, rather than regularization-strength tuning or the extra λ, drives the observed improvements.","section":"Section 4.2, Algorithm 1"}],"minor_comments":[{"comment":"The column header \"CIFAI100\" appears to be a typo for \"CIFAR100\".","section":"Table 2"},{"comment":"The sentence \"In the next version, we will consider using a smaller network model...\" refers to future work and should be removed or the corresponding experiment should be included; published manuscripts should not defer a verification step to a later version.","section":"Section 5.2, Resource Consumption"},{"comment":"The notation α is used for both the SI penalty coefficient in Eq. (2) and the Dirichlet heterogeneity parameter in the experiments; these are different quantities and should be renamed to avoid confusion.","section":"Section 4.2"},{"comment":"The statement \"Denote that α refers to the degree of data heterogeneity and when α has a higher value, indicating a trend towards homogeneity in distribution\" is ambiguous: a larger Dirichlet α means more homogeneous data, so it would be clearer to say that α controls heterogeneity and larger values correspond to more IID settings.","section":"Section 4.2, discussion after Eq. (5)"}],"recommendation":"major_revision","confidential_remarks":"The paper is an ICML 2025 camera-ready submission (PMLR 267). The main concern is that the central empirical claim, FedSSI's advantage over FL+SI, is confounded by the unreported SI penalty strength and the lack of an ablation isolating the PSM. This is fixable through additional experiments and full hyperparameter disclosure. The paper's self-citations in the theory section (Li et al. 2024a) should be checked for independent validation. I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line first: this is a solid empirical paper for the continual-federated-learning subfield, not a conceptual breakthrough. The new piece is modest and real — FedSSI takes the personalized surrogate model from the authors' own Re-Fed line and uses it to compute synaptic-intelligence importances under a global anchor, instead of using it to select replay samples. The ingredients are all known (SI, proximal personalization, Re-Fed's PSM), but the combination is new and the motivation is sensible: plain SI's surrogate loss reflects only the local distribution, which misaligns with the global objective when client data is non-IID. Figures 1–2 document that failure cleanly.\n\nThe experiments are the strength. Six datasets, class-IL and domain-IL, a dozen baselines, plus ablations over heterogeneity level, number of tasks, scalability, and bandwidth. FedSSI wins across the board, and the margins over FL+SI are consistent rather than cherry-picked. Table 3's λ–α trend is a genuine empirical pattern, even if the paper overstates what Proposition 1 proves about it.\n\nNow the soft spots, in proportion. The α notation collision is the most visible sin — α is the SI penalty in Eq. (2) and the Dirichlet heterogeneity parameter. The SI penalty strength is never reported for FedSSI or FL+SI, which is the real gap: the mechanistic claim that the PSM's importance estimates, rather than a better-tuned SI penalty, explain the Table 2 gains is not fully isolated. That said, the stress-test speculation that penalty tuning alone could produce the gaps does not fully land — FedSSI also beats replay baselines, and the gains are consistent across datasets. The cheap fix is to report the penalty strengths and add an ablation where FL+SI gets the same tuning care.\n\nThe theory section is the weakest part. Proposition 1 is imported from Hanzely–Richtárik and only shows λ trades local vs. global in the PSM objective; it does not establish the positive α–λ correlation the text claims. Theorem 1 is a self-citation from Li et al. 2024a with no proof or independent check. The paper would be more credible calling this a plausibility argument.\n\nWho it's for: anyone building rehearsal-free CFL methods or extending regularization-based CL to federated settings will want this as a reference point. My recommendation: send it to review. The revisions needed are reporting and framing, not redoing the experiments.","headline":"A useful empirical CFL paper with a plausible PSM mechanism that is not cleanly isolated from SI-penalty tuning; worth refereeing, but the authors should report missing hyperparameters and trim the theory claims.","tokens_in":23635,"tokens_out":7025,"would_cite":true,"duration_ms":58355,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"FedSSI's rehearsal-free regularizer beats CFL baselines by up to 12.47% in final accuracy.","keywords":["continual federated learning","synaptic intelligence","regularization","non-IID data","catastrophic forgetting","rehearsal-free","personalized surrogate model","class-incremental learning"],"falsifier":"Measure the cosine similarity between the per-parameter importance vector $\\Omega$ computed by FedSSI and the local-only SI importance vector of FL+SI on a strongly non-IID CIFAR10 split ($\\alpha=0.1$); if the two vectors are nearly identical while final accuracy still differs by several points, the global pull through the PSM is not the mechanism producing the reported gains.","tokens_in":22581,"feed_emoji":"🧠","tokens_out":6994,"duration_ms":57797,"temperature":0.7,"pith_summary":"Continual federated learning (CFL) systems that let clients learn from streaming tasks usually fight catastrophic forgetting by replaying cached or synthetic past samples, which costs memory and can violate privacy. This paper makes the case that regularization-based continual learning can work in the federated setting instead, and identifies synaptic intelligence (SI) as the strongest regularization baseline under IID data but one that collapses under non-IID data. The authors propose FedSSI, which adds a personalized surrogate model (PSM) at each client so that the parameter-importance scores SI relies on are computed with both local and global information. The claim is that this simple, rehearsal-free modification lets a regularized CFL method beat state-of-the-art rehearsal and non-rehearsal baselines across six datasets and two incremental scenarios, with gains up to 12.47% in final accuracy.","feed_headline":"No-replay regularizer beats CFL baselines by 12.47%","feed_subtitle":"Synaptic-intelligence importance scores gain global context, lifting non-IID accuracy without sample caching.","key_machinery":"The load-bearing object is the Personalized Surrogate Model (PSM), a per-client auxiliary model that is never used for classification. Before each new task arrives, the client updates the PSM for a few local iterations on its previous-task samples, with an added proximal pull toward the last global model; the scalar $q(\\lambda) = (1-\\lambda)/(2\\lambda)$ controls the trade-off, with $\\lambda\\to 0$ making the PSM converge to the global model and larger $\\lambda$ keeping it local. The gradient of the PSM's loss, rather than the gradient of the target model, is integrated over the training trajectory to compute the synaptic-intelligence contributions $s^k_{l,i}$ in Eq. (6), which then feed the importance scores $\\Omega^k_{l,i}$ that penalize movement of old-task weights when training the new task in Eq. (2). This mechanism is what transfers global knowledge into the regularization penalty without any rehearsal, sample caching, or extra communication.","core_discovery":"FedSSI's central claim is that the failure of regularization-based CFL under data heterogeneity is not a failure of synaptic intelligence itself but of where its importance estimates come from. In vanilla FL+SI, each client computes the surrogate-loss contributions $s^k_{l,i}$ using only its local model and local data, so the importance weights $\\Omega^k_{l,i}$ protect weights that matter locally, which can be misaligned with what the global model needs under non-IID data. FedSSI replaces that local-only computation with a Personalized Surrogate Model (PSM) $v^k_{t-1}$ trained on the client's previous-task data while being pulled toward the received global model $w^{t-1}$ by a proximal term $q(\\lambda)(v - w)$, with $q(\\lambda) = (1-\\lambda)/(2\\lambda)$. The PSM's gradient path is then used in place of the local model's to accumulate SI importance scores, so the surrogate loss in Eq. (2) penalizes changes to weights that matter for both the local and the global data distributions. The paper reports that with this change, FedSSI achieves the best final and average accuracy in all tested cases, including gains of up to 12.47% in final accuracy, and stays ahead of baselines as data heterogeneity $\\alpha$ varies.","pith_inferences":["Extension: the same 'compute importance on a globally-pulled surrogate' recipe should transfer to other regularization-based continual learners that use per-weight importance, such as EWC-style Fisher estimates; if it does, FedSSI is a template rather than a single algorithm.","Extension: the method implicitly assumes the client still has the just-completed task's data when the PSM is updated; if data must be discarded the moment a task ends, PSM training would need to run online during the task itself, a variant the paper does not test.","Extension: the $\\lambda$–$\\alpha$ trend in Table 3 suggests an automatic scheduling rule for $\\lambda$ could be fitted from the Dirichlet concentration parameter, removing the per-dataset tuning burden that the paper leaves manual."],"forward_implications":["Rehearsal-based CFL methods (such as Re-Fed, FedCIL, and GLFC) can be outperformed by a regularization-only method, so memory buffers and generative replay are not necessary for state-of-the-art CFL accuracy.","The PSM training costs about one fortieth of the per-task training budget and requires no extra communication, keeping the rehearsal-free benefit cheap in practice.","The $\\lambda$ knob gives practitioners a way to respond to the degree of data heterogeneity: decreasing $\\lambda$ (more global pull) is the direction that recovers accuracy under stronger non-IID splits, a trend confirmed in Table 3 across three datasets.","FedSSI works for both class-incremental and domain-incremental tasks without needing task boundaries at inference, unlike FOT and FedWeIT, which the paper had to modify to automate task-ID inference."],"supporting_citations":[{"why":"Supplies synaptic intelligence, the base regularization whose importance scores FedSSI re-routes through the PSM.","marker":"(Zenke et al., 2017)"},{"why":"Provides FedAvg, the aggregation protocol and main FL backbone that all baselines and FedSSI build on.","marker":"(McMahan et al., 2017)"},{"why":"Introduces the proximal term to federated optimization that motivates the global pull in the PSM update.","marker":"(Li et al., 2020)"},{"why":"Proves the interpolation result used in Proposition 1 that ties $\\lambda$ to the balance of local and global information.","marker":"(Hanzely & Richtárik, 2020)"},{"why":"Contributes the personalized surrogate model concept and the convergence result reused as Theorem 1, and is also the rehearsal-based baseline Re-Fed that FedSSI is compared against.","marker":"(Li et al., 2024a)"},{"why":"Baseline FedWeIT, a network-extension CFL method that FedSSI must beat.","marker":"(Yoon et al., 2021)"},{"why":"Baseline FedCIL, a generative-replay CFL method whose memory and privacy costs motivate the rehearsal-free design.","marker":"(Qi et al., 2023)"}],"fun_headline_variants":["No-replay CFL: local SI gains global context via surrogate","Synaptic intelligence repaired for non-IID CFL with PSM","FedSSI: personalized surrogate lifts CFL accuracy by 12.47%","Rehearsal-free CFL: importance scores from a global-pulled surrogate","Surrogate-based synaptic intelligence for heterogeneous CFL"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The proximal pull toward the global model in Eq. (5) makes each client's personalized surrogate model carry enough global knowledge that the resulting importance scores protect the right weights under non-IID data; if that pull adds no useful global signal, FedSSI reduces to plain FL+SI and its reported gains lack a mechanism.","fun_headline_variants_meta":{"raw":{"variants":["No-replay CFL: local SI gains global context via surrogate","Synaptic intelligence repaired for non-IID CFL with PSM","FedSSI: personalized surrogate lifts CFL accuracy by 12.47%","Rehearsal-free CFL: importance scores from a global-pulled surrogate","Surrogate-based synaptic intelligence for heterogeneous CFL"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000266,"raw_usage":{"total_tokens":1638,"prompt_tokens":997,"completion_tokens":641,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":613,"completion_tokens_details":{"reasoning_tokens":548}},"tokens_in":613,"tokens_out":641,"duration_ms":5995,"temperature":1.0,"reasoning_tokens":548,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T12:48:50.776185+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the cosine similarity between the per-parameter importance vector $\\Omega$ computed by FedSSI and the local-only SI importance vector of FL+SI on a strongly non-IID CIFAR10 split ($\\alpha=0.1$); if the two vectors are nearly identical while final accuracy still differs by several points, the global pull through the PSM is not the mechanism producing the reported gains.","supporting_citations":[],"review_version":1}