Pith. sign in

REVIEW 3 major objections 4 minor 106 references

Bias-Corrected Multiplier Bootstrap Inference for Spectral Edges of Large Covariance Matrices

T0 review · 3 major / 4 minor · reviewed 2026-08-02 · deepseek-v4-flash

Pith's one-line read Bias-corrected multiplier bootstrap yields valid confidence intervals for spectral edges and threshold-free spike counts.

desk verdict Novel bootstrap method for spectral-edge inference with a genuine theoretical core, but the main theorem is stated broader than what the supplement proves: the regular-edge condition is missing from the assumptions. read the letter →

arxiv 2607.08089 v2 pith:TMKZPFHV submitted 2026-07-09 stat.ME math.STstat.TH

classification stat.MEmath.STstat.TH MSC 60B2062E2062F4062H25
keywords spectraledgeinferencemultiplierbootstrapspikedcovariancematrixTracy–Widomfluctuationnumberofspikesscreeplothigh-dimensionalbiascorrection
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

High-dimensional sample eigenvalues at the edge of the bulk spectrum fluctuate on the Tracy–Widom scale, so valid inference normally requires centering and scaling constants that are hard to estimate. This paper proposes instead to perturb the data with random multipliers whose variance is tuned to be larger than the native Tracy–Widom fluctuation, making the leading non-spiked bootstrap eigenvalues approximately Gaussian after a data-driven recentering. The resulting confidence interval for the deterministic spectral edge has asymptotic coverage 1−α under the null hypothesis and vanishing coverage under local alternatives, so the same interval doubles as a spike detector and a threshold-free estimator of the number of spikes. The payoff is a practical procedure that avoids Tracy–Widom constants and supplies a theoretically grounded cutoff for the scree plot.

What carries the argument

The central object is the multiplier-bootstrap sample covariance matrix Q_MB = n^{−1} Σ ξ_i^2 (y_i−ȳ)(y_i−ȳ)^T, with i.i.d. multipliers scaled so Var(ξ^2)∼n^{−1/3+δ}. This perturbation is load-bearing: it is wider than the native Tracy–Widom scale n^{−2/3} but narrow enough to preserve the bulk structure, so the edge eigenvalues become Gaussian after rescaling by v in (3.11). The analysis centers on the self-consistent edge equations F_{n,c} and F_n and a stability/contraction argument showing that the random edge fluctuates around its deterministic limit as a Gaussian average of multiplier transforms. A data-driven recentering Δ_{r0}=μ_{r0}−λ̄_{r0} corrects the edge bias, and the bootstrap

What would settle it

Choose a population covariance spectrum whose limiting density vanishes as (x−E)^β near the edge with β≠1/2, with all stated assumptions otherwise satisfied, and compute F_{n,c} to check whether |∂_y F|·|∂_xx F| vanishes at E_MB. If simulations then show that the coverage of interval (2.5) deviates from 1−α, the missing regular-edge condition is essential and the theorem as stated is incomplete.

Watch

Extended reading notes

Core claim

The paper claims that a deliberate multiplier perturbation regularizes edge fluctuations. With multiplier variances of order n^{−1/3+δ}, the bootstrap edge fluctuates on a scale √(v/n), large enough that the largest few non-spiked bootstrap eigenvalues are conditionally Gaussian after subtracting the deterministic edge plus a bias term (Theorem 3.1). The multiplier perturbation shifts the deterministic edge itself, so the procedure estimates this bias by the difference between the observed r0-th sample eigenvalue and the average of the bootstrap eigenvalues. Theorem 3.2 states that the resulting interval covers the true edge with probability 1−α+o(1) under the null, and with probability o(1)

Load-bearing premise

The load-bearing premise is the regular-edge condition stated in the supplement (Definition D.1): at the perturbed edge, the edge equation must have nonvanishing first derivative in y and nonvanishing second derivative in x; the main text's Assumptions 3.1–3.2 do not explicitly imply it, and if it fails, the stability argument behind the Gaussian approximation collapses.

Editorial extensions

If this is right

  • One can build a confidence interval for the bulk edge without estimating Tracy–Widom centering and scaling constants; the interval's length is only slightly larger than the Tracy–Widom scale.
  • The same interval is a consistent test: under the null the coverage is 1−α asymptotically, while any additional spike that separates locally beyond n^{−1/6} pushes the coverage to zero.
  • Counting eigenvalues above the interval's upper endpoint gives a threshold-free estimator of the number of spikes with asymptotic success probability 1−α/2; the spikes need not be distinct or very large.
  • The upper endpoint can be interpreted as a data-driven scree-plot cutoff with a theoretical guarantee.
  • The procedure applies under general unknown population covariance structures satisfying standard regularity assumptions, with only a data-independent calibration of the multiplier parameter for each dimension pair.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The calibration is performed on a Gaussian reference model and then transferred via universality; the paper does not prove a general transfer theorem, so a natural extension is to test whether the same calibrated multiplier parameter preserves coverage for strongly non-Gaussian or anisotropic bulks beyond the heavy-tailed cases simulated.
  • Because the interval avoids Tracy–Widom constants, it could be inverted to compare two independent samples' bulk edges with minimal assumptions; this extension is not stated in the paper.
  • The regular-edge condition may fail for population spectra with edge singularities other than the standard square-root behavior; checking coverage in such models would delimit the method's domain.
  • The theory assumes a bounded number of spikes, while practice uses r0 of order log n; a testable extension is whether the spike-count guarantee degrades gracefully as r0 grows.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper develops a bias-corrected multiplier bootstrap for inference on the deterministic right edge E of the bulk spectrum of a high-dimensional sample covariance matrix with a general, possibly spiked, population covariance. Multipliers with variance of order n^{-1/3+δ} are used to create a bootstrap eigenvalue fluctuation at a scale larger than the Tracy–Widom scale, so that a conditional Gaussian approximation becomes tractable. The procedure constructs a confidence interval (2.5) by recentering a fresh bootstrap eigenvalue with the empirical bias Δ_{r0} and scaling by the bootstrap standard deviation s. The main results, Theorems 3.1 and 3.2, claim conditional Gaussianity of the leading non-spiked bootstrap eigenvalues after bias correction, asymptotic coverage 1−α under H0, vanishing coverage under H_a, and, in Corollary 3.1, a threshold-free estimator of the number of spikes with asymptotic exactness 1−α/2. Proofs are carried out in a substantial supplement using edge equations, a stability/contraction argument for the perturbed edge, and Berry–Esseen smoothing.

Significance. The idea is original and potentially valuable: inference is built directly from bootstrap eigenvalues, avoiding estimation of Tracy–Widom centering and scaling constants; the same interval yields edge inference, tests, and a spike-counting rule; and the required spike separation n^{-1/6+κ} is weaker than in earlier bootstrap-based factor/PCA methods. The variance formula (3.11) is explicit, the calibration rule in Section 2.3 is data-independent, and the numerical and real-data sections cover several covariance models, Gaussian and t entries, and genomic data. If the missing regularity conditions are supplied, this would be a strong contribution to high-dimensional spectral inference. However, as written, the main theorems are stated more broadly than the proofs in the supplement establish; the omitted edge-regularity and multiplier-boundedness conditions are load-bearing rather than cosmetic.

major comments (3)
  1. [Supplement C.1, Theorem C.1 / Definition D.1; main Theorem 3.1] The proof of Theorem 3.1 is routed through Supplement Theorem C.1 and Lemma D.3, which explicitly assume the regular-edge condition in Definition D.1: |∂_y F_{n,c}(x0,y0)| ≥ c0 and |∂_xx F_{n,c}(x0,y0)| ≥ c0. This condition is not listed in Theorem 3.1 or in Assumptions 3.1–3.2. Assumption B.1 only keeps the denominators 1+σ_i m_{2n,c}(E_MB) and 1+t m_{1n,c}(E_MB) away from zero; it does not control the second partial derivative of F_{n,c}. The contraction argument in Lemma D.3 needs invertibility of the Jacobian, and without it the linear response relation (C.5), and hence the Berry–Esseen step (C.7), do not follow. Since Theorems 3.2 and Corollary 3.1 inherit Theorem 3.1, the main results are broader than the proven statements. The claim in Remark 3.2 that Assumption B.1 ensures regular square-root behavior is asserted, not proved. Please add the regular-edge condition to the main assu
  2. [Example 2.1 and Supplement Remark D.1] The theorem statements are for the chi-squared multipliers of Example 2.1, where ξ^2 = χ^2_N/N has unbounded support. The proof, however, uses bounded multipliers: Lemma D.1 assumes supp(F_ξ²) ⊂ [0,C_ξ], Theorem C.1 states that the multipliers are bounded in the theoretical construction, and Remark D.1 introduces a truncated and recentered version before applying the arguments. This truncation is not part of Definition 2.1 or Example 2.1, and no argument shows that replacing the chi-squared multiplier by its truncated version leaves the bootstrap eigenvalues unchanged to the required order. Either the procedure and theorems should be formulated for the truncated/recentered multiplier, or a proof should be provided that the untruncated Example 2.1 satisfies the bounded-support requirements up to asymptotically negligible error.
  3. [Supplement C.2, Lemma C.3] The power statement (3.14) under H_a relies on Lemma C.3, whose near-critical expansion assumes f''(b) ≥ c0 at the bulk edge. This is a second non-degeneracy condition that is not stated in Assumptions 3.1–3.2. If this condition can fail for population covariances satisfying the current assumptions, the gap between λ_{r0} and E may not be of order n^{-1/3+2κ}, and the proof of vanishing coverage collapses. The condition should be stated explicitly or derived from the same edge-regularity assumption as Definition D.1.
minor comments (4)
  1. [Table 1] The row for GEUV ADIS appears misaligned: '314 4340 24 4' is not parsable as values for the nine methods. Please check the typesetting of this table.
  2. [Assumption 3.2(ii)] Assumption B.1 is used in the main text but defined only in the supplement. For a self-contained main paper, either restate it in Assumption 3.2 or summarize the essential content of the condition.
  3. [Remark 2.3 / Corollary 3.1] The theoretical framework treats r0 as a fixed integer, but the practical recipe suggests r0 = ⌈C log n⌉. Since r0 diverges, Corollary 3.1 as stated does not cover this practical choice. Please clarify whether the theory extends to r0 growing logarithmically or whether the theoretical guarantee should be read with a fixed r0.
  4. [Section 2.3] The calibration rule selects N by empirical coverage on a Gaussian reference model, but no proof is offered that the selected N satisfies the variance scaling (2.2) for every (p,n). A brief statement that the chosen N is then assumed to meet Definition 2.1 would prevent a logical gap between the calibrated implementation and the theoretical conditions.

Circularity Check

0 steps flagged · score 1.0 of 10

No circular derivation; main caveat is an unproved edge-regularity condition, not a circular step.

full rationale

The derivation chain is not circular. Theorem 3.1's conditional Gaussian approximation is obtained by decomposing lambda_{r+i}-E into a Tracy-Widom-scale term (lambda - pEMB), a multiplier-driven term (pEMB - EMB), and a deterministic bias (Delta_edge); the Gaussian law comes from a Berry-Esseen argument on the multiplier average, not from assuming the confidence interval's coverage. The bias-correction term Delta_r0 is a plug-in estimator whose consistency is proved in (C.13), and s^2 is shown in (C.14) to estimate v/n; neither quantity is defined to equal the target E. Theorem 3.2 combines the Gaussian approximation with these plug-in estimates, and Corollary 3.1 follows by thresholding, so no theorem reduces to its own conclusion by construction. The main self-citations (Ding 2021; Ding and Yang 2021) supply edge-rigidity and outlier-localization results that do not assume the bootstrap Gaussian limit, so they are independent support rather than load-bearing circularity. The calibration of N on a Wishart reference model is a tuning rule; although Table A.1 Case (I) with r=0 and equal sigma_i is essentially the same reference family used to calibrate N and is therefore not an independent finite-sample confirmation, the asymptotic claims do not rely on the calibration data. The most serious caveat is that Supplement Theorem C.1 explicitly assumes the regular-edge condition (Definition D.1), which is not stated in Theorem 3.1 and is not shown to follow from Assumption B.1. This is a missing-assumption / omitted-proof gap, not a circular reduction: assuming derivative non-degeneracy does not assume the Gaussian limit or the coverage statement. Hence no step in the claimed derivation is equivalent to its input by construction, and the circularity score is low.

Assumptions & free parameters 4 free parameters · 9 assumptions · 0 invented entities

The method does not introduce a new physical entity, force, or conserved quantity. The main free inputs are the multiplier scale N (and the implied δ), the working upper bound r0, and the number of bootstrap draws B. The theorems also rest on several technical regularity conditions in the supplement that are plausible but not all proven to follow from the main assumptions.

free parameters (4)
  • N (multiplier degrees of freedom) = 4, 9, 15 for (p,n)=(200,500),(500,750),(750,500)
    Controls Var(ξ²)=2/N and hence the Gaussianized fluctuation scale; calibrated on Gaussian reference matrices rather than derived.
  • δ (multiplier variance exponent) = not estimated directly; constrained to (2δ*, 1/3), implied by calibrated N
    Enter through Var(ξ²) ≍ n^{-1/3+δ}; the theoretical results require this order but the exact constant is not fixed by theory.
  • r0 (candidate upper bound on spike count) = ⌈3 log n⌉ in implementation
    Used to select which sample eigenvalue is treated as non-spiked; theory only proves results for fixed r0.
  • B (number of bootstrap replicates) = 2000 in simulations
    Theory requires B ≳ n^c; the specific value is a practical choice.
assumptions (9)
  • domain assumption Assumption 3.1: entries of X are centered, i.i.d., with uniformly bounded moments.
    Used throughout the proof to invoke concentration and local laws for the data matrix.
  • domain assumption Assumption 3.2(i)-(ii): p/n bounded away from 0 and infinity; non-spiked population eigenvalues bounded away from 0 and infinity; regular edge behavior via Assumption B.1.
    Defines the high-dimensional regime and the regularity of the bulk spectrum needed for edge rigidity.
  • domain assumption Assumption 3.2(iii)/(3.10): spike separation t ≳ n^{-1/6+κ} with κ>δ/2.
    Required for the power result and for consistency of the spike-number estimator.
  • ad hoc to paper Definition 2.1 and (2.2): multipliers have mean 1 and variance ≍ n^{-1/3+δ}; the χ² construction is truncated in the proof (Remark D.1).
    The entire method is built on this deliberately calibrated perturbation; it is not a physically or empirically forced scale.
  • domain assumption Assumption B.1: denominator separation |1+σ_i m_{2,n,c}(E_MB)| ≥ τ0 and inf_t |1+t m_{1,n,c}(E_MB)| ≥ τ0.
    Technical but load-bearing: it guarantees the edge of the multiplier-perturbed spectrum is regular and supports the deterministic equivalent equations.
  • ad hoc to paper Definition D.1: regular-edge condition |∂_y F_{n,c}(x0,y0)| ≥ c0 and |∂_{xx}F_{n,c}(x0,y0)| ≥ c0.
    Introduced in the proof of Theorem C.1; without it the stability analysis and Gaussian approximation collapse, yet it is not listed in the main theorem.
  • standard math Standard random matrix theory results: Knowles-Yin isotropic local laws, edge rigidity, Tracy-Widom universality, Theorem 3.7 of Ding (2021).
    The paper explicitly relies on these external results for edge rigidity and outlier localization.
  • standard math Multipliers are independent of the data; product probability space (Remark B.2).
    Needed to separate conditional-on-data Gaussianity from the multiplier CLT.
  • domain assumption Centering: results for uncentered matrices carry over to centered sample covariance matrices via Section 9 of Bloemendal et al. (2016).
    Used to justify analyzing the uncentered model (3.1).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Bias-Corrected Multiplier Bootstrap Inference for Spectral Edges of Large Covariance Matrices." pith.science (2026). https://pith.science/paper/TMKZPFHV

@misc{pith2026260708089,
  author       = {Pith},
  title        = {Pith review of: Bias-Corrected Multiplier Bootstrap Inference for Spectral Edges of Large Covariance Matrices},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/TMKZPFHV}},
  note         = {Machine review of arXiv:2607.08089}
}
abstract

Inference for spectral edges of large covariance matrices is a fundamental problem in high-dimensional statistics. A major difficulty is that the largest non-spiked sample eigenvalues, which serve as natural estimators of the edge, fluctuate on the Tracy--Widom scale. Consequently, valid inference requires accurate centering by the deterministic spectral edge together with a precise scaling constant, both of which are often difficult to estimate in practice under general unknown population covariance structures. In this paper, we propose a bias-corrected multiplier bootstrap procedure for inference on the deterministic edge of the bulk spectrum. The key idea is to introduce a carefully calibrated multiplier perturbation that regularizes the edge fluctuation to a slightly larger scale at which Gaussian approximation becomes tractable. The resulting confidence interval is constructed directly from bootstrap eigenvalues, together with a data-driven recentering step that corrects the bootstrap-induced shift of the deterministic edge. On the theoretical side, we show that, after bias correction and rescaling, the largest few non-spiked bootstrap eigenvalues are asymptotically Gaussian conditionally on the data. Building on this result, we establish the asymptotic validity of the proposed confidence interval, whose length is only slightly larger than the Tracy--Widom scale, and prove vanishing coverage under alternatives in which additional spikes separate from the bulk at a local scale larger than $n^{-1/6}$. As a consequence, the same confidence interval yields a threshold-free estimator for the number of spikes, without requiring the spikes to be distinct or very large. Equivalently, the procedure yields a data-driven and theoretically justified cutoff for the scree plot.

Figures

Figures reproduced from arXiv: 2607.08089 by the authors.

Figure 1
Figure 1. Schematic illustration of the proposed bias-corrected multiplier-bootstrap pro [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Empirical power of the proposed procedure under the alternative hypothesis with [PITH_FULL_IMAGE:figures/full_fig_p024_2.png] view at source ↗
Figure 2
Figure 2. Leading sample eigenvalues for the two real datasets. The left and the right correspond to GEUVADIS RNA-seq data and EUR genotype data, respectively. For both the left and the right, the red dashed line denotes the multiplier bootstrap threshold of the proposed procedure. The inset in the left zooms in on the third through seventh eigenvalues. The threshold in both the left and the right lies between the fourth and … view at source ↗
Figures from the paper (3 more)
Figure 3
Figure 3. Figure 3: Comparison of spike-number estimation for [PITH_FULL_IMAGE:figures/full_fig_p026_3.png]
Figure 4
Figure 4. Figure 4: Comparison of the empirical accuracy of spike-number estimation for [PITH_FULL_IMAGE:figures/full_fig_p027_4.png]
Figure 5
Figure 5. Figure 5: Leading sample eigenvalues for the two real datasets. Panels (a) and (b) cor [PITH_FULL_IMAGE:figures/full_fig_p029_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

106 extracted references · 1 linked inside Pith

  1. [1]

    Johnstone , title =

    Iain M. Johnstone , title =. Proceedings of the International Congress of Mathematicians , volume =. 2007 , publisher =

  2. [2]

    2015 , publisher =

    Yao, Jianfeng and Zheng, Shurong and Bai, Zhidong , title =. 2015 , publisher =

  3. [3]

    , title =

    Anderson, Theodore W. , title =. 2003 , publisher =

  4. [4]

    Econometrica , volume =

    Bai, Jushan and Ng, Serena , title =. Econometrica , volume =

  5. [5]

    Econometrica , volume =

    Onatski, Alexei , title =. Econometrica , volume =

  6. [6]

    The Annals of Statistics , volume =

    Edgar Dobriban , title =. The Annals of Statistics , volume =. 2020 , publisher =

  7. [7]

    IEEE Transactions on Information Theory , volume =

    Ding, Xiucai and Yang, Fan , title =. IEEE Transactions on Information Theory , volume =

  8. [8]

    The Annals of Statistics , volume =

    Nadler, Boaz , title =. The Annals of Statistics , volume =

Show all 106 references
  1. [9]

    and Reich, David , title =

    Patterson, Nick and Price, Alkes L. and Reich, David , title =. PLoS Genetics , volume =

  2. [10]

    and Patterson, Nick J

    Price, Alkes L. and Patterson, Nick J. and Plenge, Robert M. and Weinblatt, Margaret E. and Shadick, Nancy A. and Reich, David , title =. Nature Genetics , volume =

  3. [11]

    Transcriptome and genome sequencing uncovers functional variation in humans , journal =

    Lappalainen, Tuuli and Sammeth, Michael and Friedl. Transcriptome and genome sequencing uncovers functional variation in humans , journal =

  4. [12]

    A global reference for human genetic variation , journal =

  5. [13]

    The Journal of Finance , volume =

    Markowitz, Harry , title =. The Journal of Finance , volume =

  6. [14]

    Journal of Empirical Finance , volume =

    Ledoit, Olivier and Wolf, Michael , title =. Journal of Empirical Finance , volume =

  7. [15]

    Journal of the Royal Statistical Society: Series B (Statistical Methodology) , volume =

    Fan, Jianqing and Liao, Yuan and Mincheva, Martina , title =. Journal of the Royal Statistical Society: Series B (Statistical Methodology) , volume =

  8. [16]

    IEEE Transactions on Information Theory , volume =

    Nadakuditi, Raj Rao , title =. IEEE Transactions on Information Theory , volume =. 2014 , publisher =

  9. [17]

    IEEE Transactions on Information Theory , volume =

    Couillet, Romain and Hachem, Walid , title =. IEEE Transactions on Information Theory , volume =

  10. [18]

    , title =

    Johnstone, Iain M. , title =. The Annals of Statistics , volume =

  11. [19]

    , title =

    Bai, Zhidong and Silverstein, Jack W. , title =. 2010 , publisher =

  12. [20]

    Probability Theory and Related Fields , volume =

    Knowles, Antti and Yin, Jun , title =. Probability Theory and Related Fields , volume =

  13. [21]

    A Dynamical Approach to Random Matrix Theory , series =

    L. A Dynamical Approach to Random Matrix Theory , series =. 2017 , publisher =

  14. [22]

    Probability theory and related fields , volume =

    Bloemendal, Alex and Knowles, Antti and Yau, Horng-Tzer and Yin, Jun , title =. Probability theory and related fields , volume =

  15. [23]

    Communications in Mathematical Physics , volume =

    Tracy, Craig A and Widom, Harold , title =. Communications in Mathematical Physics , volume =. 1994 , publisher =

  16. [24]

    Communications in Mathematical Physics , volume =

    Tracy, Craig A and Widom, Harold , title =. Communications in Mathematical Physics , volume =. 1996 , publisher =

  17. [25]

    The Annals of Probability , volume =

    El Karoui, Noureddine , title =. The Annals of Probability , volume =

  18. [26]

    The Annals of Applied Probability , volume =

    Lee, Ji Oon and Schnelli, Kevin , title =. The Annals of Applied Probability , volume =

  19. [27]

    The Annals of Statistics , volume =

    Bao, Zhigang and Pan, Guangming and Zhou, Wang , title =. The Annals of Statistics , volume =. 2015 , publisher =

  20. [28]

    The Annals of Applied Probability , volume =

    Ding, Xiucai and Yang, Fan , title =. The Annals of Applied Probability , volume =. 2018 , publisher =

  21. [29]

    Electronic Journal of Probability , volume =

    Fan Yang , title =. Electronic Journal of Probability , volume =. 2019 , publisher =

  22. [30]

    The Annals of Applied Probability , volume =

    El Karoui, Noureddine , title =. The Annals of Applied Probability , volume =

  23. [31]

    Journal of Multivariate Analysis , volume =

    Paul, Debashis and Silverstein, Jack W , title =. Journal of Multivariate Analysis , volume =. 2009 , publisher =

  24. [32]

    , title =

    Fan, Zhou and Johnstone, Iain M. , title =. The Annals of Applied Probability , volume =

  25. [33]

    Random Matrices: Theory and Applications , volume =

    Ding, Xiucai , title =. Random Matrices: Theory and Applications , volume =. 2021 , publisher =

  26. [34]

    Phase Transition of the Largest Eigenvalue for Nonnull Complex Sample Covariance Matrices , journal =

    Baik, Jinho and Ben Arous, G. Phase Transition of the Largest Eigenvalue for Nonnull Complex Sample Covariance Matrices , journal =

  27. [35]

    Statistica Sinica , volume =

    Paul, Debashis , title =. Statistica Sinica , volume =

  28. [36]

    Journal of Multivariate Analysis , volume =

    Passemier, Damien and Yao, Jianfeng , title =. Journal of Multivariate Analysis , volume =

  29. [37]

    Braeken, Johan and van Assen, Marcel A. L. M. , title =. Psychological Methods , volume =

  30. [38]

    , title =

    Dobriban, Edgar and Owen, Art B. , title =. Journal of the Royal Statistical Society: Series B (Statistical Methodology) , volume =

  31. [39]

    Journal of the American Statistical Association , volume =

    Fan, Jianqing and Guo, Jianhua and Zheng, Shurong , title =. Journal of the American Statistical Association , volume =

  32. [40]

    Journal of the American Statistical Association , volume =

    Ke, Zheng Tracy and Ma, Yucong and Lin, Xihong , title =. Journal of the American Statistical Association , volume =

  33. [41]

    and Blandino, Andrew and Aue, Alexander , title =

    Lopes, Miles E. and Blandino, Andrew and Aue, Alexander , title =. Biometrika , volume =

  34. [42]

    The 22nd International Conference on Artificial Intelligence and Statistics , pages =

    El Karoui, Noureddine and Purdom, Elizabeth , title =. The 22nd International Conference on Artificial Intelligence and Statistics , pages =

  35. [44]

    Journal of the American Statistical Association , volume =

    Yu, Long and Zhao, Peng and Zhou, Wang , title =. Journal of the American Statistical Association , volume =

  36. [45]

    The Oxford Handbook of Random Matrix Theory , year =

  37. [47]

    The Annals of Statistics , volume =

    Xiucai Ding and Fan Yang , title =. The Annals of Statistics , volume =

  38. [48]

    Ding, Xiucai and Xie, Jiahui and Yu, Long and Zhou, Wang , title =

  39. [49]

    , title =

    van der Vaart, Aad W. , title =. 1998 , publisher =

  40. [50]

    Genetic mapping of cell type specificity for complex traits , journal =

    Watanabe, Kyoko and Umi. Genetic mapping of cell type specificity for complex traits , journal =

  41. [51]

    Efficient toolkit implementing best practices for principal component analysis of population genetic data , journal =

    Priv. Efficient toolkit implementing best practices for principal component analysis of population genetic data , journal =

  42. [52]

    and Chow, Carson C

    Chang, Christopher C. and Chow, Carson C. and Tellier, Laurent C. A. M. and Vattikuti, Shashaank and Purcell, Shaun M. and Lee, James J. , title =. GigaScience , volume =

  43. [53]

    Isotropic local laws for sample covariance and generalized

    Bloemendal, Alex and Erd. Isotropic local laws for sample covariance and generalized. Electronic Journal of Probability , volume =. 2014 , publisher =

  44. [54]

    Communications on Pure and Applied Mathematics , volume =

    Knowles, Antti and Yin, Jun , title =. Communications on Pure and Applied Mathematics , volume =

  45. [55]

    Baik, and P

    Akemann, G., J. Baik, and P. D. Francesco (Eds.) (2011). The Oxford Handbook of Random Matrix Theory . Oxford, UK: Oxford University Press

  46. [56]

    Anderson, T. W. (2003). An Introduction to Multivariate Statistical Analysis\/ (3 ed.). Hoboken, NJ: Wiley

  47. [57]

    Bai, J. and S. Ng (2002). Determining the number of factors in approximate factor models. Econometrica\/ 70\/ (1), 191--221

  48. [58]

    Bai, Z. and J. W. Silverstein (2010). Spectral Analysis of Large Dimensional Random Matrices\/ (2nd ed.). Springer Series in Statistics. New York: Springer

  49. [59]

    Ben Arous, and S

    Baik, J., G. Ben Arous, and S. P \'e ch \'e (2005). Phase transition of the largest eigenvalue for nonnull complex sample covariance matrices. The Annals of Probability\/ 33\/ (5), 1643--1697

  50. [60]

    Pan, and W

    Bao, Z., G. Pan, and W. Zhou (2015). Universality for the largest eigenvalue of sample covariance matrices with general population. The Annals of Statistics\/ 43\/ (1), 382--421

  51. [61]

    Erd o s, A

    Bloemendal, A., L. Erd o s, A. Knowles, H.-T. Yau, and J. Yin (2014). Isotropic local laws for sample covariance and generalized W igner matrices. Electronic Journal of Probability\/ 19 , 1--53

  52. [62]

    Knowles, H.-T

    Bloemendal, A., A. Knowles, H.-T. Yau, and J. Yin (2016). On the principal components of sample covariance matrices. Probability theory and related fields\/ 164\/ (1), 459--552

  53. [63]

    Braeken, J. and M. A. L. M. van Assen (2017). An empirical K aiser criterion. Psychological Methods\/ 22\/ (3), 450--466

  54. [64]

    Chang, C. C., C. C. Chow, L. C. A. M. Tellier, S. Vattikuti, S. M. Purcell, and J. J. Lee (2015). Second-generation plink: rising to the challenge of larger and richer datasets. GigaScience\/ 4 , 7

  55. [65]

    Couillet, R. and W. Hachem (2013). Fluctuations of spiked random matrix models and failure diagnosis in sensor networks. IEEE Transactions on Information Theory\/ 59\/ (8), 5090--5107

  56. [66]

    Dette, H. and A. Rohde (2024). Computationally tractable nonparametric bootstrap of high-dimensional sample covariance matrices. arXiv preprint arXiv:2406.16849\/

  57. [67]

    Ding, X. (2021). Spiked sample covariance matrices with possibly multiple bulk components. Random Matrices: Theory and Applications\/ 10\/ (01), 2150014

  58. [68]

    Li, and F

    Ding, X., Y. Li, and F. Yang (2024). Eigenvector distributions and optimal shrinkage estimators for large covariance and precision matrices. arXiv preprint arXiv:2404.14751\/

  59. [69]

    Ding, X., J. Xie, L. Yu, and W. Zhou (2026). Multiplier bootstrap meets high-dimensional PCA : the good, the bad and the modification

  60. [70]

    Ding, X. and F. Yang (2018). A necessary and sufficient condition for edge universality at the largest singular values of covariance matrices. The Annals of Applied Probability\/ 28\/ (3), 1679--1738

  61. [71]

    Ding, X. and F. Yang (2021). Spiked separable covariance matrices and principal components. The Annals of Statistics\/ 49\/ (2), 1113--1138

  62. [72]

    Ding, X. and F. Yang (2022). Tracy- W idom distribution for heterogeneous G ram matrices with applications in signal detection. IEEE Transactions on Information Theory\/ 68\/ (10), 6682--6715

  63. [73]

    Dobriban, E. (2020). Permutation methods for factor analysis and PCA . The Annals of Statistics\/ 48\/ (5), 2824--2847

  64. [74]

    Dobriban, E. and A. B. Owen (2019). Deterministic parallel analysis: An improved method for selecting factors and principal components. Journal of the Royal Statistical Society: Series B (Statistical Methodology)\/ 81\/ (1), 163--183

  65. [75]

    El Karoui, N. (2007). T racy-- W idom limit for the largest eigenvalue of a large class of complex sample covariance matrices. The Annals of Probability\/ 35\/ (2), 663--714

  66. [76]

    El Karoui, N. (2009). Concentration of measure and spectra of random matrices: Applications to correlation matrices, elliptical distributions and beyond. The Annals of Applied Probability\/ 19\/ (6), 2362--2405

  67. [77]

    El Karoui, N. and E. Purdom (2019). The non-parametric bootstrap and spectral analysis in moderate and high-dimension. In The 22nd International Conference on Artificial Intelligence and Statistics , pp.\ 2115--2124

  68. [78]

    and H.-T

    Erd o s, L. and H.-T. Yau (2017). A Dynamical Approach to Random Matrix Theory , Volume 28 of Courant Lecture Notes in Mathematics . Providence, RI: American Mathematical Society and Courant Institute of Mathematical Sciences

  69. [79]

    Guo, and S

    Fan, J., J. Guo, and S. Zheng (2022). Estimating number of factors by adjusted eigenvalues thresholding. Journal of the American Statistical Association\/ 117\/ (538), 852--861

  70. [80]

    Liao, and M

    Fan, J., Y. Liao, and M. Mincheva (2013). Large covariance estimation by thresholding principal orthogonal complements. Journal of the Royal Statistical Society: Series B (Statistical Methodology)\/ 75\/ (4), 603--680

  71. [81]

    Fan, Z. and I. M. Johnstone (2022). T racy-- W idom at each edge of real covariance and MANOVA estimators. The Annals of Applied Probability\/ 32\/ (4), 2967--3003

  72. [82]

    Johnstone, I. M. (2001). On the distribution of the largest eigenvalue in principal components analysis. The Annals of Statistics\/ 29\/ (2), 295--327

  73. [83]

    Johnstone, I. M. (2007). High dimensional statistical inference and random matrices. In Proceedings of the International Congress of Mathematicians , Volume I, pp.\ 307--333. European Mathematical Society

  74. [84]

    Ke, Z. T., Y. Ma, and X. Lin (2023). Estimation of the number of spiked eigenvalues in a covariance matrix by bulk eigenvalue matching analysis. Journal of the American Statistical Association\/ 118\/ (541), 374--392

  75. [85]

    Knowles, A. and J. Yin (2013). The isotropic semicircle law and deformation of W igner matrices. Communications on Pure and Applied Mathematics\/ 66\/ (11), 1663--1749

  76. [86]

    Knowles, A. and J. Yin (2017). Anisotropic local laws for random matrices. Probability Theory and Related Fields\/ 169\/ (1), 257--352

  77. [87]

    Sammeth, M

    Lappalainen, T., M. Sammeth, M. R. Friedl \"a nder, et al. (2013). Transcriptome and genome sequencing uncovers functional variation in humans. Nature\/ 501 , 506--511

  78. [88]

    Ledoit, O. and M. Wolf (2003). Improved estimation of the covariance matrix of stock returns with an application to portfolio selection. Journal of Empirical Finance\/ 10\/ (5), 603--621

  79. [89]

    Lee, J. O. and K. Schnelli (2016). T racy-- W idom distribution for the largest eigenvalue of real sample covariance matrices with general population. The Annals of Applied Probability\/ 26\/ (6), 3786--3839

  80. [90]

    Lopes, M. E., A. Blandino, and A. Aue (2019). Bootstrapping spectral statistics in high dimensions. Biometrika\/ 106\/ (4), 781--801

  81. [91]

    Markowitz, H. (1952). Portfolio selection. The Journal of Finance\/ 7\/ (1), 77--91

  82. [92]

    Nadakuditi, R. R. (2014). Optshrink: An algorithm for improved low-rank signal matrix denoising by optimal, data-driven singular value shrinkage. IEEE Transactions on Information Theory\/ 60\/ (5), 3002--3018

  83. [93]

    Nadler, B. (2008). Finite sample approximation results for principal component analysis: A matrix perturbation approach. The Annals of Statistics\/ 36\/ (6), 2791--2817

  84. [94]

    Onatski, A. (2009). A formal statistical test for the number of factors in the approximate factor models. Econometrica\/ 77\/ (5), 1447--1480

  85. [95]

    Passemier, D. and J. Yao (2014). Estimation of the number of spikes, possibly equal, in the high-dimensional case. Journal of Multivariate Analysis\/ 127 , 173--183

  86. [96]

    Patterson, N., A. L. Price, and D. Reich (2006). Population structure and eigenanalysis. PLoS Genetics\/ 2\/ (12), e190

  87. [97]

    Paul, D. (2007). Asymptotics of sample eigenstructure for a large dimensional spiked covariance model. Statistica Sinica\/ 17\/ (4), 1617--1642

  88. [98]

    Paul, D. and J. W. Silverstein (2009). No eigenvalues outside the support of the limiting empirical spectral distribution of a separable covariance matrix. Journal of Multivariate Analysis\/ 100\/ (1), 37--57

  89. [99]

    Price, A. L., N. J. Patterson, R. M. Plenge, M. E. Weinblatt, N. A. Shadick, and D. Reich (2006). Principal components analysis corrects for stratification in genome-wide association studies. Nature Genetics\/ 38\/ (8), 904--909

  90. [100]

    Priv \'e , F., K. Luu, M. G. B. Blum, J. J. McGrath, and B. J. Vilhj \'a lmsson (2020). Efficient toolkit implementing best practices for principal component analysis of population genetic data. Bioinformatics\/ 36\/ (16), 4449--4457

  91. [101]

    A global reference for human genetic variation

    The 1000 Genomes Project Consortium (2015). A global reference for human genetic variation. Nature\/ 526 , 68--74

  92. [102]

    Tracy, C. A. and H. Widom (1994). Level-spacing distributions and the A iry kernel. Communications in Mathematical Physics\/ 159 , 151--174

  93. [103]

    Tracy, C. A. and H. Widom (1996). On orthogonal and symplectic matrix ensembles. Communications in Mathematical Physics\/ 177 , 727--754

  94. [104]

    van der Vaart, A. W. (1998). Asymptotic Statistics . Cambridge Series in Statistical and Probabilistic Mathematics. Cambridge: Cambridge University Press

  95. [105]

    Umi \'c evi \'c Mirkov, C

    Watanabe, K., M. Umi \'c evi \'c Mirkov, C. A. de Leeuw, M. P. van den Heuvel, and D. Posthuma (2019). Genetic mapping of cell type specificity for complex traits. Nature Communications\/ 10\/ (1), 3222

  96. [106]

    Yang, F. (2019). Edge universality of separable covariance matrices . Electronic Journal of Probability\/ 24\/ (none), 1 -- 57

  97. [107]

    Zheng, and Z

    Yao, J., S. Zheng, and Z. Bai (2015). Large Sample Covariance Matrices and High-Dimensional Data Analysis . New York: Cambridge University Press

  98. [108]

    Zhao, and W

    Yu, L., P. Zhao, and W. Zhou (2025). Testing the number of common factors by bootstrapped sample covariance matrix in high-dimensional factor models. Journal of the American Statistical Association\/ 120\/ (549), 448--459

Pith tools

Reviewed August 2, 2026 · model on record in the stance chip above.