{"id":"2d9a98da-ec17-46d8-a645-4c2bc95820b0","arxiv_id":"2502.05254","paper_version":3,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"For Gaussian i.i.d. data, the singular-value spectrum of the empirical cross-covariance is governed by a cubic Stieltjes equation, with simplified edge formulas in several asymptotic regimes.","lead":"This paper derives the distribution of singular values of a sample cross-covariance matrix computed from two independent Gaussian datasets, covering all relative sizes of the number of samples and the two feature dimensions. The result extends the Marchenko-Pastur law to cross-covariances and gives noise thresholds that could be used to test whether correlations between two high-dimensional datasets are real.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified: the cubic equation is algebraically consistent with the S-transform free-product derivation, and the known product-of-Wishart cases (including the pX=pY=1 edge 27/4) are recovered; remaining limitations are explicitly scoped model assumptions.","rationale":"The reader's ACCEPT verdict is appropriate. My stress-test checked the load-bearing algebraic steps, the scaling conventions for the Wishart S-transform, the free-product construction for independent unitarily invariant Gaussian matrices, the discriminant calculations for the edge formulas, and the zero-mass bookkeeping in Eqs. (12)-(14). No internal inconsistency was found. The known prior work on products of Wishart matrices makes the cubic equation expected, and the numerical agreement confirms the root-selection and density transformation. The weakest point, if any, is the unanalyzed variance-estimation step and the heuristic nature of the signal-detection section, but neither undermines the stated central claim of the null spectral density. Therefore no objection is raised and the reader's verdict should remain unchanged.","tokens_in":11479,"tokens_out":47728,"duration_ms":491550,"concrete_test":"Solve Eq. (15) for p_X=1.5, p_Y=0.5 and compare the resulting singular-value density with the empirical histogram from SVD of independent Gaussian matrices with T=4000, N_X=T/p_X, N_Y=T/p_Y over 20 realizations; agreement within sampling error would close the remaining unplotted-regime gap.","verdict_should_be":"UNCHANGED","load_bearing_attack":"I examined the central derivation rather than the application claims. The S-transform algebra in Appendix A checks out: with W_X^T=(1/(N_X sigma_X^2)) X X^T, the Wishart S-transform is 1/(1+p_X t) when p_X=T/N_X; substituting into S_H(t)=p_X p_Y S_WX(t) S_WY(t) and inverting T(z)=z h-1 reproduces the cubic (15) exactly. The p_X=p_Y=1 case reduces to z^2 h^3 - z h + 1 = 0, whose discriminant places the upper squared-singular-value edge at 27/4, matching the known Fuss-Catalan result for products of two Ginibre matrices. The discriminant computations for the simplifying regimes also reproduce the reported edge formulas in their stated asymptotic limits. The only genuine limitations are those the paper states: the null model assumes independent Gaussian i.i.d. entries (Eq. 1) and known scalar variances, and the signal-detection discussion is explicitly heuristic. The assertion that variance-estimation fluctuations are negligible (Section II) is plausible for the macroscopic density but is not analyzed; it does not weaken the theorem for the model as defined.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper considers two Gaussian i.i.d. matrices X (T×N_X) and Y (T×N_Y) and studies the singular-value spectrum of the normalized empirical cross-covariance matrix C=(1/T)\\tilde Y^\\top \\tilde X. Using the S-transform product rule for asymptotically free Wishart matrices, it derives a cubic equation (Eq. 15) for the Stieltjes transform of H=(1/(σ_X^2 σ_Y^2 T^2)) X X^\\top Y Y^\\top, whose imaginary part yields the eigenvalue density of C^\\top C and hence the singular-value density of C. The paper then derives simplified closed-form band edges in several limits (Eqs. 20–25), confirms the formulas numerically for T=1000 with 500 independent realizations, and discusses implications for detecting shared signals when one or both dimensions exceed the sample size.","tokens_in":11703,"tokens_out":21205,"duration_ms":210626,"significance":"If correct, the paper supplies a parameter-free null distribution for singular values of an unwhitened sample cross-covariance matrix, extending the Marchenko–Pastur result and complementing earlier work on whitened cross-correlations. The S-transform derivation is standard and algebraically consistent, and the numerical simulations in Figs. 1–3 support the density and the edge formulas. The simplified edge formulas are useful for practitioners who want to calibrate significance of cross-correlations in high-dimensional data. The paper honestly acknowledges the equivalence of its auxiliary equation to earlier results (Refs. [11,14]) and the heuristic nature of the signal-detection discussion, which is an explicit strength rather than a hidden assumption.","major_comments":[],"minor_comments":[{"comment":"The notation δ(z) in the Stieltjes-transform relations is incorrect as written: a unit point mass at eigenvalue zero contributes 1/z to the Stieltjes transform, not δ(z). This matters especially for p_X>1 or p_Y>1, where the coefficient 1−p_X is negative and the intended cancellation with the zero-mass part of h_T is otherwise invisible. Please replace δ(z) by 1/z throughout these equations.","section":"Section II, Eqs. (13)–(14)"},{"comment":"The displayed definition of h(z) is missing the matrix inverse: it should read h(z)=lim_{T→∞} (1/T) E[Tr(zI−H)^{-1}], not Tr(zI−H).","section":"Appendix A, Eq. (A3)"},{"comment":"The derivations of the simplified edge formulas are asymptotic expansions of discriminants with statements such as “this can only happen if…” and “the contribution of higher-order terms will be negligible,” but no error bounds or rigorous justification of the root selection are given. Since Eqs. (20)–(25) are presented as results, please state explicitly that these are formal asymptotic expansions, and either provide control of the remainder terms or mark the formulas as asymptotically verified by the numerical experiments.","section":"Appendix B, Eqs. (B10)–(B19)"},{"comment":"The paper does not specify how the physical root of the cubic equation is selected when computing the density from the imaginary part of h. For reproducibility, please state the branch rule (for example, the root whose imaginary part is positive for z in the lower half-plane) or describe the numerical root-selection procedure used for the semi-analytic curves.","section":"Section III, Eq. (15)"},{"comment":"The statement that sampling fluctuations in estimating the scalar variances σ_X^2 and σ_Y^2 are negligible compared with spectral fluctuations is asserted but not analyzed. This is acceptable for the theorem, which is stated for the model with known variances, but because the paper motivates statistical significance of cross-correlations, it would be helpful to state explicitly that this is an additional modeling assumption and to indicate the size of the effect it can have on the band edges.","section":"Section II, paragraph after Eq. (3)"}],"recommendation":"minor_revision","confidential_remarks":"The manuscript is honest about its relationship to Refs. [11] and [14]; the main editorial question is whether the regime analysis and data-science framing provide sufficient incremental contribution for this journal. I believe the central derivation is sound and the paper is acceptable after the notation and asymptotic-presentation issues are cleaned up. I do not see circularity or unsupported claims in the main result."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper does what it says: derives the singular-value density of a sample cross-covariance of two independent Gaussian matrices in the proportional-growth limit, and gives explicit edge formulas in several practically relevant regimes. The central cubic equation is not new—the authors say so themselves in Appendix A—but the paper's contribution is to write down the simplified bounds and frame them for significance testing in under-sampled settings.\n\nWhat is actually new: the explicit formulas for γ± in the severely undersampled, unbalanced, and well-sampled cases (Eqs. 20–25), plus the observation that cross-covariance can detect shared signal when one dataset is poorly sampled and whitening is impossible. The numerical simulations match both the semi-analytic density and the edge formulas across all regimes tested. I checked the S-transform algebra: substituting the Wishart S-transform and tracing through the free-product rule reproduces the cubic exactly, and the pX = pY = 1 case gives the known 27/4 edge for the product of two Ginibre matrices. The derivations are standard but correct.\n\nSoft spots are minor. Novelty is limited to the simplified edge formulas and the significance-testing application; the core spectral density was already obtainable from earlier product-of-Wishart results (Refs. [11,14]), which the authors acknowledge honestly. The Appendix B limiting derivations are heuristic in places—they select dominant terms in the α asymptotics rather than proving uniformity—but they reproduce the reported formulas and the simulations back them. The model assumes independent Gaussian i.i.d. entries with known variances; the paper states the variance-estimation issue and argues it is negligible, but does not analyze it. That is a reasonable assumption for a null model, not a flaw. The signal-detection discussion in Section IV is explicitly heuristic and flagged as such; it does not weaken the mathematical claim.\n\nWho this is for: anyone doing high-dimensional correlation analysis in genomics, neuroscience, or finance, where T ~ N or T << N and you need a null distribution for cross-covariance singular values. It fills a practical gap and the paper is clearly written.\n\nSend it to peer review. It is useful, correct, and honest, even though the main equation is not a new framework. A referee should focus on whether the Appendix B edge-formula derivations are rigorous enough and whether the signal-detectability claims are over-stated relative to the null-model scope.","headline":"Honest, correct extension of product-of-Wishart results to explicit cross-covariance singular-value bounds; worth a serious referee, mainly for the practical edge formulas.","tokens_in":12249,"tokens_out":2596,"would_cite":true,"duration_ms":24654,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60B20","15B52","62H20"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper derives the full singular-value distribution of the sample cross-covariance matrix of two independent Gaussian datasets and shows it is governed by a cubic equation for the Stieltjes transform, with band edges given by simple…","keywords":["cross-covariance matrix","singular value distribution","random matrix theory","Marchenko-Pastur","Wishart product","Stieltjes transform","S-transform","high-dimensional statistics"],"falsifier":"Simulate the model with $T=1000$, $N_X=N_Y=2000$ (so $p_X=p_Y=0.5$) but draw the entries from a heavy-tailed distribution such as Student-$t$ with 5 degrees of freedom instead of Gaussians, and compare the empirical singular-value density and band edges with Eqs. (15) and (20); a mismatch would show the result depends essentially on Gaussianity rather than on the free-product structure. A second check: introduce a shared low-rank signal into $\\mathbf{X}$ and $\\mathbf{Y}$ and see whether the top singular value separates from the predicted $\\gamma_+$ as the sample size grows.","tokens_in":11263,"feed_emoji":"📊","tokens_out":6568,"duration_ms":57843,"temperature":0.7,"pith_summary":"When two high-dimensional datasets are sampled a limited number of times, their sample cross-covariance is full of spurious correlations even when the two datasets are truly independent. This paper derives the distribution of those spurious singular values for Gaussian data: in the large-sample limit with fixed dimension-to-sample ratios, the Stieltjes transform of the squared cross-covariance solves the cubic equation $a h^3 + b h^2 + c h + d = 0$, and the support of the density is a single band whose edges have explicit formulas in several limiting regimes. The result matters because it pins down the 'noise floor' of cross-covariances: a singular value above the band is a candidate for a genuine shared signal. The same formulas cover regimes where the number of samples $T$ is smaller than either or both dimensions, where standard whitening-based methods cannot be applied.","feed_headline":"Cubic equation sets noise band for cross-covariance singular values","feed_subtitle":"Know where random cross-correlations end; everything beyond is a candidate for a genuine shared signal.","key_machinery":"The machinery is the S-transform of free probability: for asymptotically free random matrices $\\mathbf{A}$ and $\\mathbf{B}$, the S-transform of their product is multiplicative, $S_{\\mathbf{AB}}(t) = S_{\\mathbf{A}}(t)S_{\\mathbf{B}}(t)$. The paper writes the squared cross-covariance, up to scaling, as a product of two normalized Wishart matrices, uses the known S-transform of each, and converts the result into a cubic equation for the Stieltjes transform $h(z)$. The eigenvalue density follows from the imaginary part of $h$, and the band edges follow from the zeros of the discriminant of the cubic equation.","core_discovery":"For two independent Gaussian i.i.d. matrices $\\tilde{\\mathbf{X}}$ and $\\tilde{\\mathbf{Y}}$ of dimensions $T\\times N_X$ and $T\\times N_Y$, the normalized empirical cross-covariance matrix $\\mathbf{C} = \\tfrac{1}{T}\\tilde{\\mathbf{Y}}^\\top \\tilde{\\mathbf{X}}$ has nonzero singular values whose density follows from the imaginary part of the solution of the cubic equation (Eq. 15) with coefficients given in Eqs. (16)-(19). The band edges are the zeros of the discriminant of this cubic; in simplifying limits they have closed forms: for equal aspect ratios $p_X=p_Y<1$ the edges are given by Eq. (20), in the severely undersampled limit $p_X=p_Y\\to 0$ by Eq. (21), when one dataset is much higher dimensional by Eq. (22), when both are severely undersampled by Eq. (23), and in the well-sampled case $T>N_X,N_Y$ by Eq. (25) with the lower edge at zero. In all cases the typical scale of the singular values is $\\sqrt{N_XN_Y}/T$, and the result extends the Marchenko-Pastur spectrum of self-covariances to cross-covariances.","pith_inferences":["The same cubic-equation machinery should extend to more than two datasets: because the S-transform of a product of $M$ Wishart matrices is known, the cross-covariance band for a chain of $M$ datasets should follow from a higher-degree polynomial, though the paper does not explore this.","A testable extension is to check robustness to non-Gaussian entries: free-probability universality suggests the cubic equation should hold for a wide class of i.i.d. entries with finite fourth moments, but the paper only proves the Gaussian case, so simulating heavy-tailed entries at the same aspect ratios would test this directly.","The explicit upper edge $\\gamma_+$ could seed a hypothesis test for shared signal: compare the largest observed singular value against $\\gamma_+$ computed from estimated aspect ratios, although the distribution of the largest singular value beyond the bulk edge is not derived here.","The asymmetric band in the $p_X, p_Y \\ll 1$ limit might be exploited to estimate the aspect ratio of one dataset from the observed spectrum of the other, given the scaling of the edges."],"forward_implications":["With the band edges known, singular-value outliers in a cross-covariance of uncorrelated data can be screened: any singular value above $\\gamma_+$ is a candidate for real shared signal.","The noise floor scales as $\\sqrt{N_X N_Y}/T$, so doubling the sample size shrinks spurious singular values by a factor $\\sqrt{2}$ relative to the dimensions.","Signal detection is possible when $T < N_X, N_Y$, a regime where self-covariances cannot be inverted and CCA-style whitening methods fail.","In the well-sampled regime the cross-covariance band edge is smaller than the whitened cross-correlation edge (same scaling, different prefactor), suggesting that whitening with estimated covariances injects extra noise.","A strong shared signal that is invisible in each dataset's own covariance spectrum can be detected in the cross-covariance when one dataset is well sampled enough to compensate for the other's poor sampling."],"supporting_citations":[{"why":"Defines the Marchenko-Pastur spectrum of empirical self-covariance matrices that this work extends to cross-covariances.","marker":"[1]"},{"why":"Supplies the standard toolkit—Stieltjes transform, S-transform, and freeness—used to derive the cubic equation.","marker":"[2]"},{"why":"Gives the whitened cross-correlation spectrum against which the paper's well-sampled limit is compared.","marker":"[3]"},{"why":"Prior product-of-Wishart calculation whose assumption that scalar variance fluctuations are negligible is adopted.","marker":"[10]"},{"why":"Contains the product-Wishart Stieltjes-transform result from which Eq. (A13) can be obtained by reparameterization.","marker":"[11]"},{"why":"Gives the dense product-Wishart spectral density that Eq. (A13) maps onto, omitting the data-science-relevant regimes studied here.","marker":"[14]"},{"why":"Numerical observation that cross-covariance can detect shared signals invisible in individual covariances, motivating the detection discussion.","marker":"[17]"},{"why":"Establishes the freeness theory that licenses the S-transform multiplication for the product of Wishart matrices.","marker":"[18]"}],"fun_headline_variants":["Cubic equation predicts where cross-correlations are noise","Noise floor for cross-correlations from one cubic equation","Extending Marchenko-Pastur to cross-covariance singular spectra","Singular value law for cross-covariance matrices: a cubic equation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything rests on the entries of the two data matrices being independent, identically distributed Gaussian draws with known variances, so that the normalized self-covariances are asymptotically free and the S-transform multiplication rule applies; if the data are non-Gaussian, the two datasets are dependent, or the three sizes are not simultaneously large with fixed ratios, the cubic equation and the edge formulas need not hold.","fun_headline_variants_meta":{"raw":{"variants":["Cubic equation predicts where cross-correlations are noise","Noise floor for cross-correlations from one cubic equation","Extending Marchenko-Pastur to cross-covariance singular spectra","Singular value law for cross-covariance matrices: a cubic equation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001011,"raw_usage":{"total_tokens":4244,"prompt_tokens":889,"completion_tokens":3355,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":505,"completion_tokens_details":{"reasoning_tokens":3283}},"tokens_in":505,"tokens_out":3355,"duration_ms":23147,"temperature":1.0,"reasoning_tokens":3283,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T20:08:03.587182+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate the model with $T=1000$, $N_X=N_Y=2000$ (so $p_X=p_Y=0.5$) but draw the entries from a heavy-tailed distribution such as Student-$t$ with 5 degrees of freedom instead of Gaussians, and compare the empirical singular-value density and band edges with Eqs. (15) and (20); a mismatch would show the result depends essentially on Gaussianity rather than on the free-product structure. A second check: introduce a shared low-rank signal into $\\mathbf{X}$ and $\\mathbf{Y}$ and see whether the top singular value separates from the predicted $\\gamma_+$ as the sample size grows.","supporting_citations":[{"cited_title":"(A13), reduces to: h3z2pX 2 + h2z (pX (1 − pX ) +pX (1 − pX )) + h (1 − pX )(1 − pX ) − zpX 2 + p2 X = 0, (B1) and the discriminant (Eq","cited_arxiv_id":null,"evidence_quote":"Defines the Marchenko-Pastur spectrum of empirical self-covariance matrices that this work extends to cross-covariances."},{"cited_title":"(A13) reduces to: αh3z2p2 X + h2zpX (α(1 − pX ) + (1− αpX )) + h (1 − pX )(1 − αpX ) − zαp2 X + αp2 X = 0","cited_arxiv_id":null,"evidence_quote":"Supplies the standard toolkit—Stieltjes transform, S-transform, and freeness—used to derive the cubic equation."},{"cited_title":"(A13) reduces to: αh3z2p2 X + h2zpX (α + 1) +h 1 − zαp2 X + αp2 X = 0","cited_arxiv_id":null,"evidence_quote":"Gives the whitened cross-correlation spectrum against which the paper's well-sampled limit is compared."},{"cited_title":"Wold, Multivariate analysis , 391 (1966)","cited_arxiv_id":null,"evidence_quote":"Prior product-of-Wishart calculation whose assumption that scalar variance fluctuations are negligible is adopted."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Contains the product-Wishart Stieltjes-transform result from which Eq. (A13) can be obtained by reparameterization."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the dense product-Wishart spectral density that Eq. (A13) maps onto, omitting the data-science-relevant regimes studied here."},{"cited_title":"Benigni and S","cited_arxiv_id":null,"evidence_quote":"Establishes the freeness theory that licenses the S-transform multiplication for the product of Wishart matrices."}],"review_version":1}