{"id":"e431258f-dde7-4c8d-a30b-638448b2845c","arxiv_id":"2608.07281","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"For ridgeless least squares with a spiked covariance, the asymptotic prediction risk is a closed-form function of spike eigenvalues, the aspect ratio, and target-spike alignment.","lead":"This paper derives formulas for the prediction error of the simplest interpolating linear regression model when the data have a few strong hidden directions. It shows the hidden directions can help or hurt depending on how the true signal lines up with them.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 3.1 is asserted without a proof, and Eq. (3.9) requires a bulk-eigenvector delocalization step that no stated lemma supplies.","rationale":"The reader's verdict is well supported: Theorem 3.1 is the central claim and it is not proved. I agree that the manuscript should not be accepted as is. I go slightly beyond the reader by identifying the specific unproved ingredient in the variance limit: the bulk-eigenvector delocalization needed to replace v_k^T Sigma v_k by tr(Sigma)/p. This is a substantive random matrix fact, not mere bookkeeping, and Lemma 4.4 does not appear to supply it. The reader's additional point about Theorem 4.1 operating with gamma -> infinity while Assumption 2 requires alpha_i > 1 + sqrt(gamma) is also valid; that inconsistency affects the classification claims but is secondary to the missing proof of Theorem 3.1. The formulas may well be correct—they are consistent with known single-spike results and with the heuristics—but as written the paper does not provide the derivation needed to certify them. I recommend keeping the existing REJECT verdict; a revision that supplies the proof of Theorem 3.1, states the needed delocalization condition, and reconciles Theorem 4.1 with the assumptions could change the outcome.","tokens_in":12967,"tokens_out":26525,"duration_ms":275164,"concrete_test":"Derive Eq. (3.9) for M=1 with alpha_1 -> infinity and alpha_1/p -> c > 0, starting from Lemma 4.4 and the Stieltjes transform of S^+ Sigma. Isolate the step where bulk eigenvectors are replaced by their average v^T Sigma v ~ tr(Sigma)/p = 1 + alpha_1/p, and verify that it follows from Assumption 2. If the proof does not go through, run a finite-sample simulation (e.g., n=200, p=800, Gaussian z, M=1, alpha_1=400) and compare n^{-1} tr(S^+ Sigma) with (gamma-1)^{-1}(1 + alpha_1/p); a persistent mismatch outside Monte Carlo error would falsify (3.9) under the stated assumptions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central result, Theorem 3.1, is stated in Section 3 but no proof is given; the Appendix contains only borrowed eigenvalue lemmas. The variance formula (3.9), V_X -> sigma^2/(gamma-1)(sum_j alpha_j/p + (p-M)/p), is not a direct consequence of Lemma 4.4 as stated. To obtain it, one must replace v_k^T Sigma v_k by tr(Sigma)/p = 1 + (sum_j alpha_j)/p for the bulk sample eigenvectors; that is a delocalization statement about bulk eigenvectors relative to each spike direction u_j. Lemma 4.4 controls spike eigenvalues and inverse resolvents, but it does not, by itself, deliver this bulk-eigenvector average. The theorem also claims almost sure convergence, while Lemma 4.4 is stated as convergence in probability, so the almost sure mode is unsupported. Because (3.10) is the paper's central claim, the manuscript does not currently establish it. Theorem 4.1 is additionally suspect: taking gamma -> infinity conflicts with Assumption 2, which requires alpha_i > 1 + sqrt(gamma), and some classification cases explicitly allow fixed alpha_j. Thus the regimes in Theorem 4.1 are not justified by the assumptions.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the out-of-sample prediction risk of the min-norm (ridgeless) least squares estimator under a spiked covariance model in the proportional asymptotics p/n -> γ. The main result, Theorem 3.1, claims that for γ < 1 the risk converges to σ²γ/(1-γ), and for γ > 1 gives explicit bias and variance limits that depend on the spike eigenvalues, the aspect ratio, and the squared alignments ⟨u_i, β⟩² of the regression vector with the spike eigenvectors. Section 4 uses these limits to classify benign, tempered, and catastrophic overfitting as γ → ∞ and includes simulation figures. The central theorem is stated without proof; the appendix only reproduces imported lemmas from other papers.","tokens_in":13232,"tokens_out":6159,"duration_ms":59703,"significance":"If the claimed formula were established, the paper would be a useful extension of the isotropic analysis of Hastie et al. (2022) to multi-spike covariance structures under finite fourth moments, and the alignment-driven phase classification would be a genuinely new contribution. The paper also verifies consistency with known isotropic and single-spike special cases and provides illustrative simulations. However, the central result is not proved, and the variance formula requires a bulk-eigenvector delocalization statement that is neither stated nor supplied. The significance is therefore conditional on a proof that is not present in the current manuscript.","major_comments":[{"comment":"Theorem 3.1, the central claim of the paper, is stated without proof. No derivation of the bias limit (3.8) or the variance limit (3.9) is given in Section 3, and the appendix contains only imported Lemmas 4.2–4.4, which do not by themselves establish the risk limits. The manuscript accordingly does not currently establish its main result.","section":"Section 3, Theorem 3.1"},{"comment":"The variance formula is not a direct consequence of Lemma 4.4 as stated. Lemma 4.4 controls sample spiked eigenvalues and inverse-resolvent quantities, but passing from tr(Σ̂⁺Σ)/n to σ²/(γ-1) (Σ_{j=1}^M α_j/p + (p-M)/p) requires replacing v_kᵀ Σ v_k by tr(Σ)/p for the bulk sample eigenvectors, i.e., a bulk-eigenvector delocalization result relative to the spike directions. No such lemma appears in the paper, so (3.9) lacks a load-bearing derivation.","section":"Eq. (3.9)"},{"comment":"The convergence modes are inconsistent. Theorem 3.1 claims almost sure convergence, while Lemma 4.4 is quoted as convergence in probability, and no argument is provided to upgrade the mode. Additionally, (3.8) and (3.10) contain O_p(1/p) and O_p(⟨u_i,β⟩/√p) remainders inside an almost sure limit statement; these terms need to be shown to be o(1) almost surely (with uniform rates in i), otherwise the displayed convergence is not well-formed.","section":"Theorem 3.1 and Lemma 4.4"},{"comment":"Theorem 4.1 considers γ → ∞ while Assumption 1 fixes p/n → γ > 0 and Assumption 2 requires α_i > 1 + √γ. For any fixed or bounded α_j, the condition α_j > 1 + √γ fails eventually as γ → ∞, so classification cases with constant-order α_j are not covered by the stated assumptions. The theorem needs a separate asymptotic regime with compatible assumptions on the growth of the spikes relative to γ.","section":"Theorem 4.1 and Assumptions 1–2"},{"comment":"The overfitting taxonomy takes a second limit γ → ∞ after defining R_γ as an n,p limit at fixed γ, but the paper does not state a joint limit or a uniformity condition that would justify exchanging these limits. Since the O_p remainders in (3.8) may depend on γ, the double-limit argument is not justified without additional uniformity or rate control.","section":"Section 4.1, Proposition 4.1"}],"minor_comments":[{"comment":"Equations (1.3) and (1.5) are identical and should not both be displayed; one should be removed.","section":"Section 1"},{"comment":"The keyword phrase 'prediction disk' appears to be a typo; it should probably be 'prediction risk'.","section":"Keywords"},{"comment":"The caption of Figure 5(b) says P_{i=1}^5 α_i = 200 while the text in Section 4.2 says 500; also the caption of Figure 5(a) appears to omit the summed quantity in one place.","section":"Figures 5"},{"comment":"The phrasing 'not orthogonal to all vectors in set {u_j}' and 'not orthogonal to any vector in set {u_j}' is ambiguous; the intended quantifiers should be spelled out.","section":"Section 4.1"},{"comment":"The O_p symbols inside the displayed limits should be replaced by explicit remainder terms with stated rates and uniformity in i, since the current notation is non-standard in an almost sure convergence statement.","section":"Section 3, Eqs. (3.8) and (3.10)"},{"comment":"The phrase 'the rate at which γ → ∞' is not formal. If γ_n = p/n is intended to diverge, the paper should define the joint asymptotic regime and state the corresponding versions of Assumptions 1–3.","section":"Theorem 4.1"}],"recommendation":"reject","confidential_remarks":"The paper presents an interesting formula and a plausible classification, but the main theorem is stated without proof, the variance limit relies on an unstated delocalization step, and the γ → ∞ regime in Theorem 4.1 is incompatible with the stated assumptions. These are load-bearing issues rather than presentation problems. If the authors can supply a complete proof and repair the asymptotic regime, a resubmission could be considered, but in the current state the manuscript does not establish its central claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper goes after the right problem and the formulas look right. The multi-spike generalization of ridgeless least-squares risk, with M growing as o(n^{1/4}) and only finite fourth moments, is a genuine extension of the single-spike results of Li and Sonthalia. The formula reduces correctly to the isotropic case, the single-spike case, and the known gamma<1 limit, and the simulations line up with the displayed approximations. That is real value.\n\nThe problem is that Theorem 3.1, the central claim, is stated without proof. The appendix only imports lemmas from other papers. The variance limit in (3.9) is not a direct consequence of the cited Lemma 4.4; you need a bulk-eigenvector delocalization step, essentially replacing v_k^T Sigma v_k by tr(Sigma)/p for bulk eigenvectors, and no lemma or argument supplies that. The theorem claims almost sure convergence while Lemma 4.4 is only convergence in probability. So the paper currently does not establish its main result.\n\nTheorem 4.1 is also stated in a regime the assumptions exclude. Assumption 2 requires alpha_i > 1 + sqrt(gamma). Taking gamma to infinity while allowing bounded or slowly growing spikes means those spikes eventually violate the assumption. The benign/tempered/catastrophic classification may be right, but as a theorem it is not justified.\n\nThere are also minor slips: the keywords include \"prediction disk,\" the Figure 5 caption omits the sum variable, and the text around Corollary 3.2 has some imprecise statements about divergent spikes and o(p). These are small; the two gaps above are what matter.\n\nI think the core formula is believable and the paper would be a useful reference once the proof is actually written. But in its current form, the main theorem and the overfitting classification are unsupported. A serious referee could assess whether the missing delocalization step can be supplied, so I would not desk-reject the topic; but the paper is not close to publishable as is.\n\nRecommendation: send it to peer review only if you expect the authors to fill in the proof; otherwise it will be a waste of referee time. For a reading group, it is a decent case study in what counts as a proof gap, but I would not cite it yet.","headline":"Plausible multi-spike ridgeless risk formula, but the central theorem is not actually proven and the overfitting classification runs outside the assumptions.","tokens_in":13716,"tokens_out":2172,"would_cite":false,"duration_ms":23478,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60B20","60F05"],"pacs":[],"model":"deepseek-v4-flash","headline":"Under a generalized spiked covariance model, the out-of-sample prediction risk of the ridgeless least-squares estimator has a sharp asymptotic limit determined by spike eigenvalues, aspect ratio, and the alignment of the regression vector…","keywords":["prediction risk","linear spectral distribution","random matrix theory","ridgeless least-squares estimator","spiked covariance model","benign overfitting","double descent","target-spike alignment"],"falsifier":"Run the M=1 case with gamma=2, alpha_1=4, $r^{2}$=$sigma^{2}$=1, and $\\beta$ = u_1 so that <u_1,$\\beta$>^2=1; Theorem 3.1 predicts the risk converges to (3)(1/2)^2 + (1/2) + (1)(1/4 + 1) = 0.75 + 0.5 + 1 = 2.25 as n,p -> infinity. If finite-n Monte Carlo estimates of the min-norm estimator's out-of-sample risk do not approach 2.25 (e.g., they saturate elsewhere or fail to converge), the formula is refuted.","tokens_in":12796,"feed_emoji":"🎯","tokens_out":7395,"duration_ms":61593,"temperature":0.7,"pith_summary":"This paper aims to prove that, in proportional-limit high-dimensional regression (p/n to gamma), the prediction risk of the minimum-norm least-squares interpolator is fully characterized by the spiked covariance spectrum, the aspect ratio, and the squared alignment between the regression vector $\\beta$ and each spike eigenvector. In the overparameterized regime gamma > 1, the risk limit is an explicit formula; in the underparameterized regime gamma < 1, the limit is $sigma^{2}$ gamma/(1-gamma) and is insensitive to spikes. The paper further derives a taxonomy of benign, tempered, and catastrophic overfitting driven by the growth of aggregate spike strength relative to gamma, and shows that alignment with spike directions is always detrimental. These results matter because they explain when latent factor structure in features helps or hurts generalization of interpolating models.","feed_headline":"Spike alignment decides when interpolation overfits","feed_subtitle":"New limit formula: risk of min-norm least squares depends on spike eigenvalues, gamma, and beta's alignment with latent factors.","key_machinery":"The machinery is the generalized spiked covariance model Sigma = U diag(alpha_1,...,alpha_M,1,...,1) U^T together with random-matrix-theory limits for the sample spike eigenvalues and eigenvectors. The paper imports three limit results: the almost-sure limit of the sample spiked eigenvalues $\\varphi$(alpha_k) = alpha_k (1 + gamma/(alpha_k - 1)), the almost-sure limit of the squared overlap between the sample and population spike eigenvectors, and a lemma (Lemma 4.4) giving the spike eigenvalue and inverse-resolvent limits when M grows as o($n^{{1/4}}$) under finite fourth moments. These limits are used to decompose the risk functional tr(Sigma_hat^+ Sigma) and $\\beta$^T Pi Sigma Pi $\\beta$ into bulk and spike components, which yields the explicit risk formula.","core_discovery":"Theorem 3.1 asserts that as n,p -> infinity with p/n -> gamma, the conditional prediction risk of the ridgeless estimator satisfies R_X(beta_hat; $\\beta$) - [sum_{i=1}^M (alpha_i - 1)(1 - 1/gamma)^2 <u_i,$\\beta$>^2 + (1 - 1/gamma) $r^{2}$ + $sigma^{2}$/(gamma-1) ( (sum_{j=1}^M alpha_j)/p + (p-M)/p )] -> 0 in probability, for gamma > 1, and R_X -> $sigma^{2}$ gamma/(1-gamma) for gamma < 1. The display combines the bias term, which carries the alignment dependence through <u_i,$\\beta$>^2, and the variance term, which depends on the average spike strength. The paper interprets this formula as showing that spike structure modifies the variance only when the aggregate spike strength is of order p or larger, and modifies the bias whenever $\\beta$ is not orthogonal to the spike eigenvectors.","pith_inferences":["A practical extension the authors do not spell out: if one could estimate the spike subspace from unlabeled features, projecting the design (or orthogonalizing beta) before fitting should lower interpolation risk; this is testable with factor-model datasets.","The imported lemma constrains M to o(n^{1/4}) under finite fourth moments; a self-contained proof or a numerical stress test at M near n^{1/4} would indicate how much of the diverging-M claim is carried by external results.","Theorem 4.1's gamma -> infinity classification formally needs alpha_i > 1 + sqrt(gamma), which cannot hold for bounded or slowly growing spikes; reading it literally, the regime labels apply only to spikes growing at least like sqrt(gamma), an implicit condition that should be stated.","The formula's structure suggests a more general principle: the risk impact of any covariance direction is proportional to (alpha - 1) times the signal squared along that direction; this could be extended to full non-spiked spectra by integration, a natural next step the authors do not pursue."],"forward_implications":["For gamma > 1, the prediction risk limit is an explicit function of the spike eigenvalues, the aspect ratio, and the squared alignment <u_i,beta>^2; the formula is sharp enough to be used as a benchmark for finite-sample behavior.","When the aggregate spike strength is o(p), the variance contribution matches the isotropic covariance case sigma^2/(gamma-1), and only the bias term is changed by the spikes.","Any nonzero alignment between beta and a spike direction increases the asymptotic risk, so restricting interpolation to the orthogonal complement of the spike eigenspace is the paper's suggested way to control risk.","The overfitting taxonomy says benign overfitting occurs when beta is orthogonal to the spike directions and the aggregate spike strength grows slower than gamma; if beta has nonzero alignment, tempered or catastrophic overfitting follow under spike growth conditions.","The underparameterized regime gamma < 1 is spike-insensitive: the risk tends to sigma^2 gamma/(1-gamma), independent of both the spectrum and the alignment."],"supporting_citations":[{"why":"Provides Lemma 2.1, the bias-variance decomposition of the ridgeless estimator, and the isotropic-baseline risk that the new formula extends.","marker":"Hastie et al. (2022)"},{"why":"Supplies Lemma 4.4, the spike eigenvalue and inverse-resolvent limits needed for the diverging-M case.","marker":"Hu et al. (2026)"},{"why":"Supplies Lemma 4.2, the almost-sure limit of the squared overlap between sample and population spike eigenvectors.","marker":"Johnstone and Yang (2018)"},{"why":"Supplies Lemma 4.3, the almost-sure convergence of sample spiked eigenvalues to phi(alpha_k).","marker":"Bai and Yao (2012)"},{"why":"Gives the limiting spectral distribution equation whose Stieltjes transform handles the bulk spectrum.","marker":"Silverstein (1995)"},{"why":"Defines benign overfitting, the phenomenon the spike model can destroy or preserve.","marker":"Bartlett et al. (2020)"},{"why":"Provides the benign/tempered/catastrophic taxonomy used to classify overfitting regimes.","marker":"Mallinar et al. (2022)"},{"why":"Prior single-spike alignment results that the multi-spike classification generalizes and contrasts with.","marker":"Li and Sonthalia (2025)"}],"fun_headline_variants":["Spike alignment, not just spectrum, sets overfitting","Benign overfitting emerges from spike-aligned signal","Ridgeless risk: beta's spike direction is key","Double descent depends on latent spike geometry","Minimal moments, spike alignment: new ridgeless risk law"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The main theorem's diverging-M case rests on an imported lemma (Lemma 4.4) that is not proved in this paper; if that lemma's spike eigenvalue and inverse-resolvent limits fail under the stated finite-fourth-moment condition, the central risk formula's growing-M version does not follow.","fun_headline_variants_meta":{"raw":{"variants":["Spike alignment, not just spectrum, sets overfitting","Benign overfitting emerges from spike-aligned signal","Ridgeless risk: beta's spike direction is key","Double descent depends on latent spike geometry","Minimal moments, spike alignment: new ridgeless risk law"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000308,"raw_usage":{"total_tokens":1771,"prompt_tokens":964,"completion_tokens":807,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":580,"completion_tokens_details":{"reasoning_tokens":728}},"tokens_in":580,"tokens_out":807,"duration_ms":8320,"temperature":1.0,"reasoning_tokens":728,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T11:06:57.218443+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the M=1 case with gamma=2, alpha_1=4, $r^{2}$=$sigma^{2}$=1, and $\\beta$ = u_1 so that <u_1,$\\beta$>^2=1; Theorem 3.1 predicts the risk converges to (3)(1/2)^2 + (1/2) + (1)(1/4 + 1) = 0.75 + 0.5 + 1 = 2.25 as n,p -> infinity. If finite-n Monte Carlo estimates of the min-norm estimator's out-of-sample risk do not approach 2.25 (e.g., they saturate elsewhere or fail to converge), the formula is refuted.","supporting_citations":[{"cited_title":"Surprises in high-dimensional ridgeless least squares interpolation","cited_arxiv_id":null,"evidence_quote":"Provides Lemma 2.1, the bias-variance decomposition of the ridgeless estimator, and the isotropic-baseline risk that the new formula extends."},{"cited_title":"On sample eigenvalues in a generalized spiked population model","cited_arxiv_id":null,"evidence_quote":"Supplies Lemma 4.3, the almost-sure convergence of sample spiked eigenvalues to phi(alpha_k)."},{"cited_title":"Benign overfitting in linear regression","cited_arxiv_id":null,"evidence_quote":"Defines benign overfitting, the phenomenon the spike model can destroy or preserve."},{"cited_title":"Simon and Amirhesam Abedsoltan","cited_arxiv_id":null,"evidence_quote":"Provides the benign/tempered/catastrophic taxonomy used to classify overfitting regimes."}],"review_version":1}