{"id":"c9502678-aaf6-4195-899f-85b84da090bf","arxiv_id":"2504.12625","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Weighted spectral algorithms with clipped importance weights achieve rates arbitrarily close to the optimal capacity-dependent rates under covariate shift with unbounded density ratios.","lead":"This paper proves convergence rates for spectral regularization algorithms in nonparametric regression when training and test input distributions differ. It shows that normalized importance weights suffer a saturation effect, and that clipped weights recover near-optimal rates under unbounded density ratios.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorems 2 and 4 are stated for arbitrary real-valued responses, but both proofs require a pointwise bound |y - f_rho(x)| <= 2M; without an explicit bounded-output assumption the claimed rates are not established.","rationale":"The reader's weakest assumption identifies exactly this gap, and I agree. It is load-bearing because it sits in the sample-error propositions that directly feed both main theorems; without bounded outputs the Bernstein inequalities do not apply. The concern does not appear to invalidate the mathematical program: adding an explicit bounded-response or sub-Gaussian assumption is a standard, minor amendment, and the rest of the derivation is consistent with that amendment. I therefore keep the reader's conditional verdict rather than escalating to rejection. A secondary presentation issue is that Theorem 5's statement omits the uniform bounded-density-ratio condition used in Lemma 27, but since the surrounding text introduces that theorem under the bounded-ratio case, I do not treat it as the primary concern.","tokens_in":36536,"tokens_out":23041,"duration_ms":224372,"concrete_test":"Independently re-derive the moment bound in Proposition 17 under only finite second moment for y (no uniform bound), and check whether E||eta||^p <= (1/2)p! \\tilde \\sigma^2 L^{p-2} holds for all p >= 2. If the condition fails for any p (e.g., y with finite variance but unbounded fourth moment), then the claimed rates in Theorems 2 and 4 do not follow without adding |y - f_rho| <= 2M to the theorem statements; re-running the proof with that assumption added should close the gap.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In the main setup the response is only declared to be real-valued (Y = R), yet Proposition 17 and Proposition 22 are the sample-error steps for Theorems 2 and 4, and both use the bound |y - f_rho(x)| <= 2M. In Proposition 17, the Bernstein moment calculation for eta(x,y) = (lambda I + L_K)^(-1/2) w(x)(y - f_rho(x)) K_x replaces |y - f_rho(x)|^p by (2M)^p to obtain the moment condition of Lemma 14; no assumption in the paper supplies such a bound. Proposition 22 uses the same 2M factor in bounding E||xi_3||^2 <= 4M^2 D_n \\hat N(lambda). Consequently the rate n^{-r/(2r+beta)+epsilon} in Theorem 4, and the rates in Theorem 2, are not justified for heavy-tailed outputs, and the constants C' and \\tilde C_{r,epsilon} depend on an M that is never defined. The fix is straightforward - add |y| <= M (or a sub-Gaussian/sub-exponential tail condition) to the assumptions or theorem statements - but as written the theorems are broader than what the proof establishes.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies spectral algorithms for nonparametric regression in a reproducing kernel Hilbert space under covariate shift, where the training and test input distributions differ but the conditional distribution of the output given the input is unchanged. The main results are: (i) Theorem 5, an optimal-rate guarantee for the classical unweighted spectral algorithm under a bounded density ratio; (ii) Theorem 2 and Corollary 3, which analyze a normalized-weight spectral algorithm and prove capacity-independent rates when the density ratio is potentially unbounded, subject to a saturation phenomenon at regularity r = 3/2; and (iii) Theorem 4, which introduces a clipped-weight spectral algorithm and proves rates of order n^{-r/(2r+beta)+epsilon}, arbitrarily close to the capacity-dependent optimal rates. The technical core is an operator-valued error decomposition combined with Bernstein-type concentration inequalities in Hilbert spaces and effective-dimension estimates.","tokens_in":36829,"tokens_out":6055,"duration_ms":56880,"significance":"If the stated results hold, they would meaningfully extend the existing theory of importance-weighted kernel ridge regression under covariate shift to the broad family of spectral algorithms, removing restrictive bounded-eigenfunction assumptions used in earlier work (e.g., Ma et al. 2023) and addressing the saturation effect of normalized weights. The paper is methodical: the error decompositions in Section 3 are carefully derived, the constants are made explicit, and the proofs rely on standard, checkable concentration arguments. However, two load-bearing assumptions are missing from theorem statements, so the theorems as written are broader than what the proofs actually establish. These omissions are fixable but currently block acceptance.","major_comments":[{"comment":"The proofs of both Proposition 17 and Proposition 22 require a uniform bound on the response variable. In Proposition 17, the moment calculation for eta(x,y) = (lambda I + L_K)^(-1/2) w(x)(y - f_rho(x)) K_x replaces |y - f_rho(x)|^p by (2M)^p, and the proof of Proposition 22 bounds E||xi_3||^2 by 4M^2 D_n \\hat N(lambda). Neither the general setup in Section 1 (where Y = R is declared) nor the statements of Theorems 2 and 4 include any bounded-output or conditional-sub-Gaussian assumption. Consequently, the claimed rates for heavy-tailed responses are not justified, and the constants C' and \\tilde C_{r,epsilon} depend on an M that is never defined. The fix is straightforward: add an explicit assumption such as |y| <= M almost surely (or a sub-Gaussian/sub-exponential tail condition) to the setup and to the statements of Theorems 2 and 4, and note how M enters the constants.","section":"§4.1, Proposition 17 and §4.2, Proposition 22"},{"comment":"The statement of Theorem 5 does not contain any assumption on the density ratio w, yet its proof invokes Lemma 27, which requires the weight function to be uniformly bounded, |w(x)| <= U. Without this assumption the key estimate ||L_K^{1/2}(lambda I + \\tilde L_K)^(-1/2)||_op <= sqrt(U) is unavailable, so the theorem is not established under the assumptions as written. Please add the bounded-density-ratio assumption to Theorem 5, or modify the proof if the result is intended to hold without it.","section":"Theorem 5 and Lemma 27"}],"minor_comments":[{"comment":"The condition beta + alpha(1 - beta) >= 1, combined with the standing bounds 0 < alpha <= 1 and 0 < beta <= 1, is equivalent to alpha = 1 or beta = 1. The theorem statement should say this explicitly, because as written the phrase 'under Assumption 1 with 0<alpha<=1 and Assumption 3 with 0<beta<=1' suggests coverage of the full range, which is not what the condition allows.","section":"Theorem 2"},{"comment":"The definition of J3 in Proposition 10 is (lambda I + \\hat L_K)^(-1/2)(S_X^\\top \\hat W \\bar y - S_X^\\top \\hat W S_X f_rho), but Proposition 22 states the target as (lambda I + \\hat L_K)^(-1/2)(S_X^\\top \\hat W \\bar y - S_X^\\top \\hat W f_rho), omitting the S_X before f_rho. Please make the notation consistent.","section":"Proposition 22"},{"comment":"The proof of Proposition 17 uses sums over m, writing 1/m and P_{j=1}^m, although the sample size throughout the paper is n. This appears to be a typographical carryover; please replace all occurrences of m with n in that proposition and its proof.","section":"Proposition 17"},{"comment":"There are several typographical errors, including 'yeilds' in Proposition 12, 'conifidence' in Proposition 12, and inconsistent spacing in the Rényi divergence display. A careful proofreading pass is needed.","section":"Throughout"},{"comment":"The statement that rates can 'approach the optimal capacity-dependent convergence rates arbitrarily closely' should be qualified: the exponent becomes n^{-r/(2r+beta)+epsilon}, but the constant \\tilde C_{r,epsilon} depends on epsilon and grows through factors such as the factorial of ceil(1/epsilon). Thus the closeness is in the exponent, with constants that may degrade as epsilon decreases; this is worth stating explicitly for precision.","section":"Theorem 4, discussion after the theorem"}],"recommendation":"major_revision","confidential_remarks":"The two missing assumptions are easy to repair and do not appear to reflect a fundamental flaw in the argument. I recommend returning the manuscript to the authors for revision rather than rejecting it. Please also ensure that the authors check whether the bounded-density-ratio condition in Theorem 5 is intended to be an explicit hypothesis, since the current statement is inconsistent with the proof."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Main plot: the clipped-weight result is a genuine step forward, but the theorems are stated wider than the proofs. The stress-test note is right. Proposition 17 and Proposition 22 both use a pointwise bound |y - f_rho(x)| <= 2M, and no such bound appears in Assumptions 1-3 or in the theorem statements. That makes the rates in Theorems 2 and 4 unsupported for heavy-tailed responses. The fix is one line — add |y| <= M to the model or to the statement — but as written the main results are not proven. Also the constants C' and tilde C_{r,epsilon} depend on an M that is never defined, which is confusing.\n\nWhat is genuinely new: extending the covariate-shift analysis from kernel ridge regression to the general class of spectral algorithms. Gizewski et al. needed bounded density ratios; Gogolashvili et al. handled KRR only. The normalized-weight saturation analysis and the clipping device for unbounded ratios are new, and Theorem 4's rate n^{-r/(2r+beta)+epsilon} is the advertised extension. The proof structure is coherent: standard Bernstein inequalities in Hilbert space, operator interpolation, inverse-operator decompositions. No circularity in what I checked; the use of Guo et al. (2017) for the unweighted case is legitimate.\n\nSoft spots are two. The missing bounded-output assumption is load-bearing — I'd want it fixed before any acceptance. Minor: the rates are oracle rates depending on unknown r, beta, alpha, which is normal for this nonparametric literature, and the epsilon in Theorem 4 comes with constants that blow up as epsilon -> 0, which the paper does not discuss. Also, the paper does not always distinguish normalized weights from the original analysis carefully in notation early on, but that is cosmetic.\n\nWho is this for? A specialist in nonparametric learning theory working on covariate shift or spectral regularization. It deserves a serious referee — the core idea is good and the proofs are mostly careful. The referee should require the bounded-output assumption be stated and the constants cleaned up. I would not cite it yet in its current form; after the fix, it becomes citable.","headline":"Solid, genuinely new extension of covariate-shift rates to spectral algorithms, but the main theorems overstate what the proofs establish because a bounded-output assumption is used and never stated.","tokens_in":37284,"tokens_out":2202,"would_cite":false,"duration_ms":22509,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["68T05","62G08","68Q32"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proves that a clipped-weight spectral algorithm achieves near-optimal regression rates under covariate shift even when density ratios are unbounded.","keywords":["spectral algorithms","covariate shift","kernel ridge regression","reproducing kernel Hilbert space","importance weighting","density ratio","effective dimension","nonparametric regression"],"falsifier":"Take a covariate-shift regression where the density ratio satisfies the paper's moment condition but the output noise is heavy-tailed, e.g. $y=f_\\rho(x)+\\varepsilon$ with $t_2$ noise, and numerically estimate the excess prediction error of the clipped weighted spectral algorithm as $n$ grows. If the observed rate is clearly slower than $n^{-r/(2r+\\beta)+\\epsilon}$ or the confidence statement fails, the missing bounded-output assumption is load-bearing; conversely, if the rate holds, the assumption is removable.","tokens_in":36368,"feed_emoji":"📊","tokens_out":6342,"duration_ms":58222,"temperature":0.7,"pith_summary":"The paper addresses supervised regression when training and test inputs come from different distributions (covariate shift), with the conditional output distribution unchanged. It studies spectral algorithms—regularized kernel methods including kernel ridge regression as a special case—and asks how fast the excess prediction error on the test distribution shrinks. The central claim is that when the density ratio between test and training inputs is unbounded but satisfies a moment condition, the plain weighted algorithm can be suboptimal; a normalized version restores optimal capacity-independent rates but saturates for smooth regression functions; clipping the weights at a data-driven threshold removes the saturation. With clipped weights the convergence rate approaches the optimal capacity-dependent rate $n^{-r/(2r+\\beta)+\\epsilon}$ up to logarithmic factors, where $r$ measures smoothness and $\\beta$ the effective dimension. A companion result shows that when the density ratio is uniformly bounded, the unweighted spectral algorithm already attains minimax optimal rates.","feed_headline":"Clipped weights restore near-optimal covariate-shift rates","feed_subtitle":"Clipped weights approach optimal regression rates even for unbounded density ratios.","key_machinery":"The central object is the weighted spectral estimator in a reproducing kernel Hilbert space, formed by applying a filter function $g_\\lambda$ (with qualification $\\nu_g\\ge 1/2$) to the empirical weighted covariance operator $S_X^\\top W S_X$. The argument runs through three error terms: an approximation term controlled by the filter qualification and smoothness $r$, a sample-error term controlled by Bernstein inequalities in Hilbert space, and a weight-bias term controlled by concentration of the empirical average of the density ratio. The crucial mechanisms are the effective dimension $N(\\lambda)=\\operatorname{Tr}((\\lambda I+L_K)^{-1}L_K)$ for complexity, the normalized weight $\\bar w(x_i)=w(x_i)/\\frac1n\\sum_j w(x_j)$ for bounded average weight, and the clipping threshold $D_n$ that keeps the operator norm of the empirical covariance operator from growing like $O(n)$ while truncating tails of $w$. The clipping is what removes the saturation at $r=3/2$ that afflicts normalized weights.","core_discovery":"For a weighted spectral estimator $f^{\\hat w}_{z,\\lambda}=g_\\lambda(S_X^\\top \\hat W S_X)S_X^\\top \\hat W \\bar y$ with clipped weights $\\hat w(x)=\\min\\{w(x),D_n\\}$, the paper proves an excess prediction error bound $\\|f^{\\hat w}_{z,\\lambda}-f_\\rho\\|_{L^2_{\\rho^{te}_X}}\\le \\tilde C_{r,\\epsilon}\\, n^{-r/(2r+\\beta)+\\epsilon}(\\log 6/\\delta)^2$ under a Rényi-type moment assumption on the density ratio, smoothness $f_\\rho=L_K^r u_\\rho$, and an effective-dimension condition $N(\\lambda)\\le C_0\\lambda^{-\\beta}$. The regularization and clipping thresholds are chosen as $\\lambda=n^{-1/(2r+\\beta)+\\epsilon/r}$ and $D_n=n^{\\alpha\\epsilon}$ for $1/2\\le r\\le 3/2$, or $D_n=n^{\\alpha\\epsilon/(r-1/2)}$ for $r>3/2$. The parameter $\\alpha\\in[0,1]$, which controls how far the test distribution may deviate from the training distribution, affects only the clipping threshold and not the rate. This is what the authors mean by approaching the optimal capacity-dependent rates arbitrarily closely: for any $\\epsilon>0$, the rate exponent is within $\\epsilon$ of $r/(2r+\\beta)$.","pith_inferences":["Editorial inference: Because $\\alpha$ enters only through the clipping threshold $D_n$, a data-driven or Lepski-type choice of $D_n$ might deliver the same near-optimal rates without knowing $\\alpha$ exactly.","Editorial inference: The proofs use the bounded-output condition $|y-f_\\rho(x)|\\le 2M$ in the Bernstein moment bounds, but the main theorems do not state it; if outputs are heavy-tailed, robust losses or output truncation may be needed for the rates to survive.","Editorial inference: The error decomposition depends only on filter qualification, effective dimension, and weight concentration, so the same clipping technique likely transfers to other spectral settings such as functional regression or online kernel learning under covariate shift."],"forward_implications":["If the main theorem is correct, importance-weighted kernel ridge regression and other spectral algorithms under covariate shift achieve rates within $\\epsilon$ of the capacity-dependent optimum even when density ratios are unbounded.","The saturation at $r=3/2$ is not inherent to spectral regularization; clipping resolves it, so smooth target functions no longer waste extra regularity.","For uniformly bounded density ratios, weighting is unnecessary: the classical unweighted spectral algorithm already reaches minimax optimal rates on the test distribution.","The exponent $r/(2r+\\beta)$ matches the standard minimax rate for kernel regression, so covariate shift itself does not degrade the achievable rate once weights are handled properly."],"supporting_citations":[{"why":"Provides the baseline weighted kernel ridge regression analysis under covariate shift with uniformly bounded density ratios, which the paper extends to spectral algorithms and unbounded ratios.","marker":"Ma et al. (2023)"},{"why":"Supplies the moment condition on the density ratio (Assumption 1) and the importance-weighted KRR analysis for unbounded weights.","marker":"Gogolashvili et al. (2023)"},{"why":"Introduces weighted spectral algorithms for covariate shift under a bounded density ratio assumption, the framework this paper pushes to unbounded ratios.","marker":"Gizewski et al. (2022)"},{"why":"Provides the spectral-algorithm learning theory, including the filter-function lemma and operator norm estimates used in the error decomposition.","marker":"Guo et al. (2017)"},{"why":"Supplies the optimal-rate framework for regularized least squares and the Hilbert-space Bernstein inequality used in the sample-error bounds.","marker":"Caponnetto and De Vito (2007)"},{"why":"Gives learning bounds for importance weighting and the normalized-weight construction that the paper adapts to spectral algorithms.","marker":"Cortes et al. (2010)"},{"why":"Provides the vector-valued Bernstein inequality used to control operator-norm fluctuations of the clipped-weight empirical operator.","marker":"Pinelis (1994)"},{"why":"Supplies the scalar Bernstein inequality used to show that normalized and unnormalized weight averages are close with high probability.","marker":"Boucheron et al. (2013)"}],"fun_headline_variants":["Clipped weights fix covariate-shift rate gap","Weight clipping yields near-optimal shift rates","Near-optimal rates with clipped weights under shift","Clipped weighting restores near-optimal regression rates","Unbounded density ratios? Clip the weights"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The proofs assume the response is uniformly bounded around the regression function, $|y-f_\\rho(x)|\\le 2M$, in order to apply Bernstein moment bounds, yet the main theorems do not state this bounded-output assumption; if the outputs are heavy-tailed, the claimed rates are not justified.","fun_headline_variants_meta":{"raw":{"variants":["Clipped weights fix covariate-shift rate gap","Weight clipping yields near-optimal shift rates","Near-optimal rates with clipped weights under shift","Clipped weighting restores near-optimal regression rates","Unbounded density ratios? Clip the weights"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000308,"raw_usage":{"total_tokens":1829,"prompt_tokens":1079,"completion_tokens":750,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":695,"completion_tokens_details":{"reasoning_tokens":678}},"tokens_in":695,"tokens_out":750,"duration_ms":7779,"temperature":1.0,"reasoning_tokens":678,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:27:41.334553+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a covariate-shift regression where the density ratio satisfies the paper's moment condition but the output noise is heavy-tailed, e.g. $y=f_\\rho(x)+\\varepsilon$ with $t_2$ noise, and numerically estimate the excess prediction error of the clipped weighted spectral algorithm as $n$ grows. If the observed rate is clearly slower than $n^{-r/(2r+\\beta)+\\epsilon}$ or the confidence statement fails, the missing bounded-output assumption is load-bearing; conversely, if the rate holds, the assumption is removable.","supporting_citations":[{"cited_title":"On a regularization of unsupervised domain adaptation in RKHS","cited_arxiv_id":null,"evidence_quote":"Introduces weighted spectral algorithms for covariate shift under a bounded density ratio assumption, the framework this paper pushes to unbounded ratios."},{"cited_title":"Learning bounds for importance weighting","cited_arxiv_id":null,"evidence_quote":"Gives learning bounds for importance weighting and the normalized-weight construction that the paper adapts to spectral algorithms."},{"cited_title":"Optimum bounds for the distributions of martingales in B anach spaces","cited_arxiv_id":null,"evidence_quote":"Provides the vector-valued Bernstein inequality used to control operator-norm fluctuations of the clipped-weight empirical operator."}],"review_version":1}