{"id":"478c6ece-6f81-49ed-957a-f51ee9bcc777","arxiv_id":"2607.02681","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":7.5,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"Filtering-based robust multi-task gradient descent matches minimax rates under task contamination and heterogeneity, removing the √d contamination barrier of regularization and score-based methods.","lead":"This paper shows that common robust multi-task methods pay an unnecessary √d cost under task contamination, and gives a filtering-based gradient method that matches minimax rates for both shared and personalized parameters. It matters for federated and multi-task systems that must borrow strength without being wrecked by a few bad clients or studies.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The Reader correctly isolates the local strong-convexity / high-probability Lipschitz pair as the weakest modeling assumption, but that pair is explicitly scoped and does not undermine the matching claim inside the stated regime. The negative results, minimax lower bounds, and filtering upper bounds form a coherent package; the only residual gap is the usual logarithmic factors and the covariance-estimation overhead already quantified in Corollary 1. No stronger load-bearing flaw (e.g., an incorrect reduction from gradient to parameter error, a gap in the stability certificate, or a regime where the claimed optimality fails while the assumptions hold) was identified. Therefore the ACCEPT verdict stands.","tokens_in":91604,"tokens_out":572,"duration_ms":6931,"concrete_test":"Independently re-derive the covariance estimation error of Algorithm 3 (Theorem 8 / 15) under the homogeneous Gaussian mean model with known identity covariance, then plug into the black-box bound of Theorem 7; confirm that the resulting global rate collapses exactly to Õ(√(d/(nK))+ε/√n+√ε h) with no residual (d/n)^{1/4} term when ε=0, matching classical multi-task mean estimation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Corollary 1 matching Theorem 5 up to logs in the regime n≳d or ε^{2}(1∨√(n/d))≲n/K) is internally consistent under the paper's stated assumptions. The negative results (Theorems 1–4) cleanly establish the ε√(d/n) barrier for the listed regularization and score-based families in the Gaussian mean model; the lower bounds (Theorem 5 / 12) correctly capture the √ε h interaction; and the filtering + single-task covariance construction (Algorithms 2–3) removes the extra √d factor while retaining personalization via soft-thresholding. The residual covariance term ε/√n·[(d/n)^{1/4}+(d/n)^{1/2}] is already acknowledged and vanishes in the claimed regime. Local strong convexity / high-probability Lipschitz (Assumptions 2 and 6) are standard for non-asymptotic GD analyses and are not hidden; the paper does not claim rates outside that ball. No internal contradiction or missing step that would overturn the matching rates was found.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper studies multi-task ERM under adversarial contamination of an ε-fraction of K tasks (each of size n) together with heterogeneity among clean tasks. It first proves that several standard paradigms—adaptive/robust center regularization, global matrix penalties, decomposition (dirty) models, and score-based outlier-task detection—incur a worst-case contamination error of order ε√(d/n) in the Gaussian mean model, which is suboptimal relative to the lower bound ε/√n (Theorems 1–4). It then establishes minimax lower bounds for both the global average-risk minimizer θ* and the clean local minimizers θ^(k)* in a general heterogeneous ERM setting, capturing the interaction √(ε)h + h^(k) (Theorem 5 / 12). A filtering-based robust multi-task gradient descent method (Algorithms 1–3), using joint robust gradient estimation and a simple single-task covariance filter, is shown to attain high-probability upper bounds matching these lower bounds up to logs in a broad regime (roughly n ≳ d or ε²(1 ∨ √(n/d)) ≲ n/K), under local strong convexity, smoothness, and sub-Gaussian gradients (Corollary 1). Simulations and a HAR real-data study support robustness and personalization relative to many benchmarks.","tokens_in":91930,"tokens_out":1491,"duration_ms":25006,"significance":"If the matching rates hold as stated, the work cleanly separates a dimension-dependent contamination barrier that affects a wide family of regularization and score-based methods from a filtering approach that removes the extra √d factor while retaining personalization under heterogeneity. The lower bounds improve on prior work by making the √(ε)h interaction explicit and by treating both global and local parameters. The algorithmic construction (JRGE + single-task covariance filtering + soft-thresholded local gradients) is computationally practical and is supported by extensive comparisons. The appendix proofs use standard tools (stability certificates, covering arguments, Taylor expansions for regularizers) and the optimality regime is stated explicitly (Remark 5, Figure 1). These are genuine contributions to robust multi-task / federated learning under simultaneous contamination and heterogeneity.","major_comments":[{"comment":"Algorithms 2 and 3 take the contamination fraction ε as a known input (and λ_Σ, λ are tuned with knowledge of ε-scale quantities). The theoretical rates and the filtering certificate (Lemma 8, Proposition 4) depend on this. The manuscript should either (i) state clearly that ε is assumed known, as is common in strong-contamination analyses, or (ii) add a short discussion/robustness check on misspecification of ε (e.g., over-estimating ε by a constant factor). Without this, the practical claim that the method is ready for use when ε is unknown is slightly stronger than the theory supports.","section":null},{"comment":"Corollary 1 / Remark 5: the upper bound still carries the residual contamination term (ε/√n)[(d/n)^{1/4} + (d/n)^{1/2}]. The abstract and introduction emphasize that the method “removes the extra √d contamination dependence” of regularization methods. That is correct relative to ε√(d/n), but when n is only moderately larger than d the residual is not fully dimension-free. A one-sentence clarification in the abstract or Remark 5 that full minimax optimality (matching ε/√n) holds in the stated regime, and that outside it a milder dimension factor remains, would prevent over-reading of the claim.","section":null}],"minor_comments":[{"comment":"Assumption 6 (high-probability Lipschitz of sample gradients with L' ≲ (nKd)^{C}) is used for uniform control of filtering iterates. A brief remark that this is a high-probability strengthening of population smoothness (Assumption 2), and that it holds for the mean and GLM examples under the stated sub-Gaussian conditions (Lemmas 1–2), would help readers who skip the appendix.","section":null},{"comment":"Section 2.1.1 / Assumption 1: the list of regularizer conditions is long. A short pointer that Lasso, Ridge, Bridge, SCAD, MC+, and hard-thresholding all satisfy it (with verification deferred to Appendix A.6) is already present; ensuring the main-text statement of Theorem 1 explicitly says “for any regularizer satisfying Assumption 1” would make the negative result easier to cite.","section":null},{"comment":"Tables 1–3 and Appendix C: bold/italic marking of best and second/third is helpful. Adding a one-line note that Single-task is omitted from global-error columns because it does not produce a pooled estimator would avoid confusion.","section":null},{"comment":"Notation: the same symbol L is used for the smoothness constant (Assumption 2) and for the proximal radius of the regularizer (Assumption 1). Different letters would reduce cognitive load when both sections are read together.","section":null},{"comment":"Figure 1 caption: “shaded region corresponds to the regime where the upper bound in Corollary 1 is minimax optimal up to logarithmic factors” is clear; a brief axis label or legend for the two boundaries (n ∼ d and n ∼ Kε²) would make the figure self-contained.","section":null},{"comment":"Related work: the discussion of untrusted-batch / batch-contamination models (QV18, CLM20, ABLY26) is useful. A sentence contrasting task-level contamination with within-batch contamination would further situate the contribution.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The paper is technically solid and fills a genuine gap (matching rates for both global and local parameters under simultaneous contamination and heterogeneity). The negative results for a broad class of regularizers are particularly clean and of independent interest. I see no integrity or novelty-disclosure issues. Fit for a strong statistics / ML theory venue is good; the main risk is that some readers may over-interpret “removes √d” without reading Remark 5. Minor revision on the two points above should be sufficient."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is that this paper actually closes a concrete gap. In the Gaussian mean model it proves that a wide family of regularizers and score-based outlier detectors all pay an extra √d in the contamination term; then it gives matching minimax lower and upper bounds for a general ERM multi-task problem and a filtering algorithm that removes that factor while still personalizing.\n\nWhat is new is the systematic negative package (Theorems 1–4), not just one regularizer, plus the joint lower bounds for θ* and the clean local θ(k)* that correctly pick up the √ε h interaction. The algorithm is standard filtering plus a simple single-task covariance estimator and soft-thresholded local gradients; the analysis shows the residual covariance term vanishes in a broad regime (roughly n ≳ d or ε^{2}(1∨√(n/d)) ≲ n/K). That is the right comparison to the lower bound, and the paper is honest about the residual term outside that regime.\n\nThe math looks carefully done. Assumptions are local strong convexity/smoothness and sub-Gaussian gradients, plus a high-probability Lipschitz condition for uniform control of the iterates—standard for non-asymptotic GD and not hidden. Simulations and the HAR experiment are sensible and beat a long list of Byzantine and MTL baselines. No code is a real but ordinary reproducibility soft spot; free parameters (ε, thresholds, stepsizes) are the usual ones for this literature.\n\nThis is for people who work on robust multi-task, federated, or Byzantine estimation and care about rates. It is not a conceptual revolution, but it is a clean, useful advance with matching theory. I would send it to peer review without hesitation and would cite the negative results and the rate statements myself.","headline":"Solid theory paper: clean negative results on the ε√d/n barrier plus a filtering method that matches minimax rates for both global and local parameters under joint contamination and heterogeneity.","tokens_in":92515,"tokens_out":507,"would_cite":true,"duration_ms":10871,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F35","62H12","68T05"],"pacs":[],"model":"grok-4.5","headline":"Filtering multi-task gradients beats regularization: contamination error drops from ε√(d/n) to near the minimax rate ε/√n while still personalizing under heterogeneity.","keywords":["multi-task learning","federated learning","robustness","data contamination","heterogeneity","minimax optimality","filtering","gradient descent"],"falsifier":"In the Gaussian mean model with large d, fix ε and n so that ε√(d/n) is several times larger than ε/√n; if the filtering estimator's worst-case error still tracks ε√(d/n) rather than the claimed near-minimax rate, the barrier-removal claim fails.","tokens_in":92511,"feed_emoji":"🧰","tokens_out":694,"duration_ms":7245,"temperature":0.7,"pith_summary":"When many related learning tasks share information, a few corrupted tasks and ordinary differences among clean tasks can ruin standard multi-task methods. This paper shows that popular regularization schemes and score-based outlier detectors all pay an extra √d factor in the contamination term, producing error of order ε√(d/n) even when the information-theoretic limit is only ε/√n. The authors prove matching minimax lower bounds for a general heterogeneous empirical-risk problem, then give a filtering-based robust multi-task gradient method that attains those rates (up to logs) over a broad sample-size regime. The same procedure returns both a global estimator and clean-task personalized estimators. Simulations and a smartphone activity data set confirm that the method stays accurate under contamination while adapting to task differences.","feed_headline":"Filtering kills the √d contamination barrier in multi-task learning","feed_subtitle":"Near-minimax rates for global and personalized estimators when some tasks are adversarial","key_machinery":"Joint robust gradient estimation (JRGE) via iterative filtering of task-level gradients, fed by a simple robust covariance estimator built from single-task empirical covariances, then plugged into multi-task gradient descent with soft-thresholded local updates.","core_discovery":"In contaminated multi-task ERM with ε-fraction adversarial tasks and heterogeneous clean tasks, regularization families and score-based detectors are fundamentally limited by a dimension-dependent contamination barrier of order ε√(d/n). A filtering-based robust multi-task gradient descent that jointly aggregates gradients and estimates their covariance removes that barrier, matching the minimax rates ε/√n + √(ε)h + √(d/(nK)) for the global parameter and the corresponding personalized rates for clean local parameters, up to logarithmic factors, under local strong convexity, smoothness and sub-Gaussian gradients.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Filtering removes √d contamination barrier in multi-task ERM","Filter-based multi-task GD matches minimax under adversarial tasks","Covariance-aware gradient filtering kills ε√(d/n) barrier","Robust multi-task filtering attains ε/√n rates without dimension hit","Joint gradient filtering enables robust personalization under contamination"],"cache_read_input_tokens":82048,"weakest_assumption_plain":"Each task risk must be strongly convex and smooth inside a fixed-radius ball around its own minimizer, and sample gradients must obey a high-probability Lipschitz condition so that the filtering steps stay controlled throughout the trajectory.","fun_headline_variants_meta":{"raw":{"variants":["Filtering removes √d contamination barrier in multi-task ERM","Filter-based multi-task GD matches minimax under adversarial tasks","Covariance-aware gradient filtering kills ε√(d/n) barrier","Robust multi-task filtering attains ε/√n rates without dimension hit","Joint gradient filtering enables robust personalization under contamination"]},"model":"grok-4.5","effort":"low","cost_usd":0.004342,"raw_usage":{"total_tokens":1349,"prompt_tokens":890,"num_sources_used":0,"completion_tokens":91,"cost_in_usd_ticks":43420000,"prompt_tokens_details":{"text_tokens":890,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":368,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":890,"tokens_out":91,"duration_ms":4035,"temperature":1.0,"reasoning_tokens":368,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T07:44:57.512967+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"In the Gaussian mean model with large d, fix ε and n so that ε√(d/n) is several times larger than ε/√n; if the filtering estimator's worst-case error still tracks ε√(d/n) rather than the claimed near-minimax rate, the barrier-removal claim fails.","supporting_citations":[],"review_version":1}