{"id":"86111535-680c-4c0e-b801-615a347e132d","arxiv_id":"2607.08123","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A partially shared subspace spiked model plus RMT-backed estimation recovers shared rank/indices and optimally pools two high-dimensional covariances, including a high-dim contrastive PCA estimator.","lead":"Two related high-dimensional datasets often share only part of their covariance structure. This paper gives a spiked model and estimator that finds that shared subspace, pools the data optimally, and improves the target covariance with random-matrix guarantees.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The unproven leap from oracle consistency (Theorems 5–6) to Algorithm 1 is the load-bearing soft spot for the central estimation claim.","rationale":"The reader correctly isolates the weakest link: oracle consistency plus an unproved iterative algorithm, together with the ad-hoc cutoff margin. That gap is load-bearing because every downstream object (optimal weight α*, Frobenius loss L(α), plug-in covariance, and the high-dimensional CDR estimator) is built on the output of Algorithm 1, not on the oracles. The RMT foundations (Lemma 2, Theorems 3–4) and the closed-form pooling analysis (Theorem 7) appear solid under the stated assumptions; the concern is solely whether the practical estimator inherits those guarantees. Simulations in §4 and Appendix A4 show good finite-sample behavior, but they do not close the theoretical gap the authors themselves flag. Hence the verdict remains CONDITIONAL; no stronger or weaker adjustment is warranted. The concrete test above directly checks whether the leap holds in the regime where the paper claims asymptotic concentration.","tokens_in":31215,"tokens_out":696,"duration_ms":20381,"concrete_test":"Under the exact simulation design of §4.1 (r_S=5, mixed indices, both orthogonal and non-orthogonal distinct blocks, (p,n_X,n_Y)=(1000,250,1000)), record for each of 500 replications whether Algorithm 1’s output (ˆr_S, ˆΨ_X, ˆΨ_Y) equals the oracle pair (ˆr*_S, ˆΨ*_X, ˆΨ*_Y) computed from the true max(Ψ) and true r_S; if the match rate falls below 95 % (or the algorithm cycles), the leap from Theorems 5–6 to the usable estimator is not empirically reliable and the central claim weakens.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim rests on joint estimation of the shared subspace (rank and indices) under the PSS model. Theorems 5–6 establish almost-sure consistency only for the oracle estimators ˆr*_S and ˆΨ*_X, ˆΨ*_Y that already know max(Ψ_X), max(Ψ_Y) or r_S. The paper itself states (end of §3.2) that these results “do not automatically guarantee the consistency of the empirical estimator ˆr_S output by Algorithm 1,” and that termination is not universally guaranteed (only that the finite-state map eventually cycles or fixes). The iterative procedure is therefore an unproved bridge between the RMT limits and the usable estimator that feeds the pooled projection ˆP^(α)_S and the plug-in ˆΣ_X. Finite-sample underestimation of the cutoff is only patched by an ad-hoc margin ε whose three candidates (0, 1/√(n_X+n_Y), bootstrap) have no uniformly best choice (Discussion §6, Appendix A4.2). If the alternating map fails to recover the oracle fixed point with high probability under Assumptions 1–3, the subsequent asymptotic loss L(α) and the claimed efficiency gain over target-only spiked estimation do not attach to the procedure that is actually run.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces a partially shared subspace (PSS) spiked covariance model for two high-dimensional datasets that share an r_S-dimensional subspace of unknown rank and arbitrary spectral position while retaining distinct spikes. Under proportional growth (Assumptions 1–2) and a shared/distinct singular-value separation condition (Assumption 3), it derives almost-sure limits for principal angles between sample eigenspaces (Lemma 2, Theorem 3), constructs oracle estimators for shared rank and index sets that are a.s. consistent (Theorems 5–6), and proposes an iterative Algorithm 1 that alternates between them. A weighted pooled projection estimator of the shared subspace is shown to have asymptotic Frobenius loss L(α) with closed-form optimal weight α* (Theorem 7), yielding a plug-in target covariance estimator and a high-dimensional contrastive dimension-reduction estimator from the distinct subspace. Simulations under the PSS model, misspecification, and degenerate endpoints, plus two real-data illustrations (COVID-era GMV portfolios and LGG gene expression), support the claims.","tokens_in":31631,"tokens_out":1091,"duration_ms":9700,"significance":"If the results hold, the paper supplies a clean geometric alternative to proximity-based transfer learning for high-dimensional covariances and the first contrastive dimension-reduction procedure with asymptotic guarantees in the proportional-growth regime. The RMT foundations (extensions of Johnstone/Paul/BBP) are carefully derived, the optimal pooling weight is closed-form and interpretable, and the framework strictly generalizes existing common-PC and multi-group models (Remark 1). The simulations are extensive (including misspecification and degenerate endpoints) and the real-data examples are coherent. The main practical deliverable—an implementable joint estimator of shared rank, indices, and the pooled covariance—would be useful for portfolio construction, multi-study genomics, and related settings where background data are abundant but only partially aligned.","major_comments":[{"comment":"End of §3.2 and Theorems 5–6: the paper itself states that oracle consistency of ˆr*_S and ˆΨ*_X, ˆΨ*_Y “do not automatically guarantee the consistency of the empirical estimator ˆr_S output by Algorithm 1,” and that termination is not universally guaranteed (only that the finite-state map eventually cycles or fixes). The subsequent asymptotic loss L(α), optimal weight α*, and efficiency claims for ˆΣ_X and the distinct-subspace CDR estimator all attach to the procedure that is actually run. Either a consistency argument for the alternating map under Assumptions 1–3, or a high-probability bound that it recovers the oracle fixed point, is needed before the central estimation claims can be regarded as fully established.","section":null},{"comment":"Assumption 3 and Discussion §6 / Appendix A4.2: the separation condition ϕ_cX(λ_rS,S)ϕ_cY(γ_rS,S) > σ_1(Φ_X,D Q_D^⊤ U_D Φ_Y,D) is load-bearing for Theorems 5–6 and for the cutoff used in Algorithm 1. Finite-sample underestimation of the cutoff is only patched by an ad-hoc margin ε whose three candidates (0, 1/√(n_X+n_Y), parametric bootstrap) have no uniformly best choice. The paper should either supply a data-driven, theoretically justified rule for ε or quantify the probability that the unadjusted cutoff recovers r_S under the stated assumptions, so that practitioners know when the procedure is reliable.","section":null}],"minor_comments":[{"comment":"Notation for debiased spikes (˜λ, ˜γ) and the debiasing map d(ℓ,c) appears in §3.2 without an explicit display of the inverse of the BBP map; a short displayed equation would help readers.","section":null},{"comment":"Figure 1 and Table 1: the mixed-position, small-sample cells show substantial underestimation of r_S by PSS; a brief discussion of when the method is expected to struggle would improve interpretability.","section":null},{"comment":"Appendix A1 Algorithm 1: the initialization ˆΨ_X = {r_X}, ˆΨ_Y = {r_Y} is natural but could be motivated more explicitly (why the largest indices rather than, e.g., a random or top-k start).","section":null},{"comment":"Related-work paragraph on transfer learning: a short comparison of the geometric PSS condition versus the usual Frobenius/sparsity proximity conditions would clarify the novelty claim.","section":null},{"comment":"Typos / polish: “eﬀiciency” (ligature issues), “diﬀiculty”, and a few missing spaces around math operators appear in the main text and appendix.","section":null}],"recommendation":"major_revision","confidential_remarks":"The technical core is solid RMT and the simulations are thorough; the main obstacle to acceptance is the unproved bridge from oracle theorems to the iterative algorithm that is actually used. If the authors can close that gap (even with a high-probability argument under stronger but still realistic conditions) the paper would be a clear contribution. Scope fits a top methods journal in statistics / high-dimensional inference."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a clean methodology paper. The real novelty is the PSS model: two spiked covariances share a subspace of unknown rank that can sit anywhere in the spectrum, while each keeps its own distinct spikes. That strictly generalizes Franks–Hoff and Shi–Al Kontar (Remark 1). They then deliver the two-sample principal-angle asymptotics (Lemma 2, Theorem 3), a closed-form optimal pooling weight for the shared projection, and a high-dimensional contrastive-dimension-reduction estimator with guarantees—the last of which fills an actual gap in the CDR literature.\n\nThe RMT foundations look carefully done. Proposition 1 is standard; the extensions to two samples and the singular-value limits under Assumptions 1–3 are the technical core and appear correct. Simulations under the model, under misspecification, and at the degenerate endpoints (r_S=0 and fully shared) support rank recovery and Frobenius gains over MgCov, PerPCA, and the usual spiked/sample baselines. The two real-data illustrations (COVID GMV portfolio, LGG vs GTEx) are coherent and show the distinct subspace doing something useful that ordinary PCA and existing CDR methods miss.\n\nThe soft spot the stress-test flags is real and the paper itself admits it: Theorems 5–6 give a.s. consistency only for the oracle rank and index estimators that already know max(Ψ) or r_S. Algorithm 1 is an alternating map whose consistency is not proved, and termination is only guaranteed up to cycles. Finite-sample underestimation of the cutoff is patched by an ad-hoc margin ε with no uniformly best choice (Appendix A4.2). That is a genuine gap between the RMT limits and the procedure that feeds the pooled projection and the plug-in covariance. It is not fatal—the simulations converge quickly and the asymptotics of the pooled estimator itself (Theorem 7) are clean—but it keeps the central estimation claim from being fully closed.\n\nWho it is for: people who already work with spiked models, multi-group covariance, or high-dim transfer/CDR. The math is standard RMT, the code and data are promised, and the citation pattern is appropriate. I would send it to peer review; a referee can demand a consistency argument or stronger finite-sample analysis for the iteration without killing the contribution. Worth reading and, for anyone in this niche, worth citing.","headline":"Solid high-dim RMT paper that generalizes partial subspace sharing and gives a usable pooling estimator; the main soft spot is the unproved leap from oracle consistency to the iterative algorithm they actually run.","tokens_in":32211,"tokens_out":592,"would_cite":true,"duration_ms":7175,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62H25","62H12","60B20"],"pacs":[],"model":"grok-4.5","headline":"When two high-dimensional covariances share only part of their spiked subspace, you can recover that shared part, optimally pool it, and improve the target covariance estimate.","keywords":["common principal components","contrastive dimension reduction","high-dimensional covariance estimation","random matrix theory","transfer learning","spiked covariance model","partially shared subspace"],"falsifier":"Generate data from the PSS model with known shared rank and mixed index sets under proportional growth; check whether Algorithm 1 recovers the true shared rank and index sets with probability approaching one, and whether the pooled covariance's Frobenius error falls below the target-only spiked estimator at the rate predicted by the asymptotic loss formula.","tokens_in":32118,"feed_emoji":"📐","tokens_out":967,"duration_ms":9916,"temperature":0.7,"pith_summary":"High-dimensional covariance estimation is usually starved for samples, yet related auxiliary data often exist. This paper argues that the useful overlap is geometric: the two spiked covariances share a low-dimensional subspace of unknown rank that can sit anywhere in the spectrum, while each keeps its own distinct spiked directions. Under proportional growth and random-matrix asymptotics, the sample principal angles between the two spiked subspaces converge to deterministic limits; those limits let you estimate both the shared rank and the indices of the eigenvectors that span it. Once the shared projection is identified, a closed-form weight optimally pools the two sample projections to minimize asymptotic Frobenius loss, and the resulting plug-in estimator for the target covariance improves on any target-only spiked estimator. The same shared-versus-distinct split supplies a high-dimensional contrastive dimension-reduction method with asymptotic guarantees, something existing contrastive PCA-style procedures lack. Applications to pandemic portfolio construction and tumor-versus-normal gene expression illustrate the practical payoff.","feed_headline":"Shared spikes alone improve high-dimensional covariance estimates","feed_subtitle":"A closed-form weight pools only the common subspace; distinct directions stay separate, with asymptotic guarantees.","key_machinery":"The partially shared subspace (PSS) model together with the almost-sure limits of sample principal angles (Lemma 2 / Theorem 3): shared singular values converge to a product of attenuation factors times a rotation, distinct singular values converge separately, and a debiased cutoff separates them, driving the iterative rank-and-index estimator and the optimal pooling weight.","core_discovery":"Under the partially shared subspace (PSS) model, the shared rank and the index sets of the shared spiked eigenvectors can be recovered consistently from the singular values and principal angles of the two sample spiked subspaces; the shared projection can then be estimated by a weighted sum of the two sample projections whose optimal weight has a closed form that depends on relative sample sizes and relative spike strengths, yielding an asymptotic efficiency gain over target-only estimation and a consistent high-dimensional contrastive subspace.","pith_inferences":["The same principal-angle separation idea could extend to more than two groups if a multi-subspace angle or joint-and-individual variation measure replaces the two-matrix singular values.","Combining the optimal shared-projection pooling with existing optimal eigenvalue shrinkage would likely produce a still tighter covariance estimator for downstream tasks such as LDA or GMV portfolios.","The method itself can serve as a diagnostic: a stable zero shared-rank estimate is evidence against a shared-subspace hypothesis, something many common-PCA procedures cannot detect."],"forward_implications":["Target covariance estimates in high dimensions can be improved by any related background dataset that shares even a non-leading subspace, without requiring the whole eigenbasis or parameter proximity.","Negative transfer is self-limiting: if no shared structure exists the estimated shared rank collapses to zero.","Contrastive dimension reduction gains its first high-dimensional asymptotic guarantees via the PSS distinct subspace.","Portfolio risk estimates during regime shifts can exploit pre-shift returns as background while isolating crisis-specific factors.","Tumor-versus-normal gene-expression analysis can separate organ-level shared variation from disease-specific directions with quantified error."],"fun_headline_variants":["Partially shared spikes enable joint high-dim covariance estimation","Closed-form pool of shared subspace beats target-only covariance estimates","Consistent recovery of shared rank and spikes from two sample subspaces","PSS model pools only common spikes while keeping distinct directions separate","Weighted shared projections yield efficiency gains under partial overlap"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The shared directions must produce sample singular values that stay strictly larger than those of the distinct directions, and the alternating algorithm that uses those cutoffs is assumed to inherit the consistency proved only for its oracle versions.","fun_headline_variants_meta":{"raw":{"variants":["Partially shared spikes enable joint high-dim covariance estimation","Closed-form pool of shared subspace beats target-only covariance estimates","Consistent recovery of shared rank and spikes from two sample subspaces","PSS model pools only common spikes while keeping distinct directions separate","Weighted shared projections yield efficiency gains under partial overlap"]},"model":"grok-4.5","effort":"low","cost_usd":0.004352,"raw_usage":{"total_tokens":1282,"prompt_tokens":739,"num_sources_used":0,"completion_tokens":83,"cost_in_usd_ticks":43520000,"prompt_tokens_details":{"text_tokens":739,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":460,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":739,"tokens_out":83,"duration_ms":4442,"temperature":1.0,"reasoning_tokens":460,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-10T12:41:27.784841+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Generate data from the PSS model with known shared rank and mixed index sets under proportional growth; check whether Algorithm 1 recovers the true shared rank and index sets with probability approaching one, and whether the pooled covariance's Frobenius error falls below the target-only spiked estimator at the rate predicted by the asymptotic loss formula.","supporting_citations":[],"review_version":1}