Pith. sign in

REVIEW 3 major objections 8 minor 44 references

Edgeworth corrections for the spiked eigenvalues of non-Gaussian sample covariance matrices with applications

T0 review · 3 major / 8 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper proves, under four moment and smoothness assumptions, a first-order Edgeworth expansion for the spiked eigenvalues of sample covariance matrices with non-Gaussian entries, reducing the approximation error to $o(n^{-1/2})$ and…

desk verdict A serious attempt at a real open problem, but the key universality theorem contradicts the paper's own variance computation. read the letter →

arxiv 2507.09584 v2 pith:ARY37DKG submitted 2025-07-13 math.ST math.PRstat.MEstat.TH

classification math.STmath.PRstat.MEstat.TH MSC 62E2060B2062H25
keywords Edgeworthexpansionspikedcovariancemodellargesteigenvaluenon-Gaussiandatapartialgeneralizedfourmomenttheoremconfidenceintervalsnumberofspikeshigh-dimensionalstatistics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper establishes a first-order Edgeworth expansion for the distribution of the leading sample eigenvalues in a spiked covariance model when the data are not Gaussian, an extension left open by the Gaussian treatment of Yang and Johnstone (2018). The expansion adds an explicit $n^{-1/2}$ correction term built from the third cumulant of the squared entries and from the other spikes, shrinking the error of the normal approximation from $O(n^{-1/2})$ to $o(n^{-1/2})$. If the result is right, confidence intervals for population spikes can be shortened to a level of accuracy previously available only under Gaussianity, and estimation of the number of spikes becomes more reliable in low-dimensional settings. The paper also provides consistent estimators for the entry moments needed to compute the correction, and simulations show the corrected densities track the empirical ones far better than the Gaussian approximation.

What carries the argument

The argument rests on three approximations chained together. First, Theorem 5 shows the eigenvalue statistic can be replaced, up to $o(n^{-1/2})$, by a linear spectral statistic $\Omega(\rho_n, Z)$ that is a sum of independent contributions after conditioning on the noise spectrum. Second, Theorem 6 uses the partial generalized four-moment theorem (PG4MT) of Jiang and Bai (2021b) to assert that the distribution of this statistic is the same for genuinely non-Gaussian entries $Z$ and for Gaussian entries $Y$, again up to $o(n^{-1/2})$ — this is the step that converts the Gaussian-only analysis into a universal one. Third, Theorem 7 feeds the Gaussian version, written as $n^{-1/2}\sum_i c_{ni}(W_i^2 - 1)$ conditionally on the eigenvalues of $n^{-1}YY'$, into the classical Edgeworth expansion for sums of independent variables (Petrov 1975; the lemma is inherited from Yang and Johnstone 2018), producing the explicit $\Phi + n^{-1/2}p_1\phi$ form. The cumulants $\kappa_{2,k}$, $\kappa_{3,k}$ and the mean shift $\mu(g_{nk})$ enter exactly at this last step.

What would settle it

Compute, at increasing $n$ and for an entry distribution that matches a Gaussian only in its first four moments (for example a three-point lattice with zero mean, unit variance and zero third moment), the sup-norm distance between the empirical distribution functions of the two statistics $\Omega_s(Z_1,Z)$ and $\Omega_s(Z_1,Y)$ defined in Theorem 6; if this distance does not decay faster than $n^{-1/2}$, the central expansion of Theorem 1 is false.

Watch

Extended reading notes

Core claim

The central discovery is a universality statement with a rate: the distribution of the normalized spiked eigenvalue statistic $$R_k = $n^{{1/2}}$(\hat l_k - \rho_{nk})/\tilde\sigma_{nk}$$ is, up to an error of $o(n^{-1/2})$, independent of the entry distribution except through two low-order cumulants of $Z_{11}^2$ and the cross-spike interaction term $A(g_{nk})$. Concretely, Theorem 1 gives $$\sup_x \left|P(R_k \le x) - \Phi(x) - $n^{{-1/2}}$\left[\tfrac16 \kappa_{2,k}^{-3/2}\kappa_{3,k}(1-$x^{2}$) - \kappa_{2,k}^{-1/2}(\mu(g_{nk}) + A(g_{nk}))\right]\$\varphi$(x)\right| = o($n^{{-1/2}}$),$$ where $\kappa_{2,k},\kappa_{3,k}$ are constructed cumulants of $\tilde Z_{1k}^2 -1$ and $A(g_{nk})$ sums contributions from the other population spikes. The same structure resolves the open problem of Yang and Johnstone (2018) for the single-spike case and, for multiple spikes, shows that each spiked eigenvalue's Edgeworth correction depends not only on its own spike but on all spikes through $A(g_{nk})$.

Load-bearing premise

The load-bearing premise is that the partial generalized four moment theorem delivers a distributional error of order $o(n^{-1/2})$ for the specific linear statistics $\Omega_s(Z_1,Z)$ versus $\Omega_s(Z_1,Y)$; this rate is asserted in Remark 9 with the detailed cancellations deferred, and if it is invalid the non-Gaussian Edgeworth expansion is not established.

Editorial extensions

If this is right

  • Confidence intervals for a population spike built from the Edgeworth-corrected E-type pivot have coverage error $o(n^{-1/2})$, one order better than the $O(n^{-1/2})$ error of the Gaussian Z-type pivot (Theorem 4).
  • In the multi-spike case, the correction contains the explicit interaction term $A(g_{nk}) = \frac{l_k-1}{(l_k-1)^2-\gamma_n}\sum_{j\ne k}\frac{l_j-1}{l_k-l_j}$, so the distribution of the $k$-th spiked eigenvalue depends on all other spikes, not just on $l_k$.
  • A new cardinality estimator $\hat r = \sum_{k=1}^p I(\hat l_k \in C_k)$ built on Edgeworth-corrected intervals outperforms existing spike-counting methods in low-dimensional (small $n,p$) regimes across Gamma, Uniform, and Gaussian data.
  • The moments $\beta_z$, $\Gamma$ and $\Delta$ needed to compute the correction are consistently estimated from leave-one-out inverses of $S$, making the expansion implementable without prior knowledge of the entry distribution (Theorem 3).
  • The authors state that a second-order Edgeworth expansion remains open because it would require a first-order approximation for the associated linear spectral statistic under non-Gaussianity.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The $o(n^{-1/2})$ universality suggests the same three-step strategy (linear-statistic reduction, four-moment comparison, conditional Edgeworth expansion) could be exported to other spiked ensembles — correlation matrices, F-matrices, or beta ensembles — where only Gaussian or first-order CLT results exist; the paper does not claim this.
  • Because the Edgeworth correction is dominated by the third cumulant of the squared entries, the method is most valuable for skewed or heavy-tailed data; for nearly Gaussian data the estimated correction adds variance without much bias reduction, which may explain the simulation cases where the Edgeworth-based method trails its Gaussian rival.
  • The explicit dependence of $A(g_{nk})$ on the gap $l_k - l_j$ implies that when two spikes are close, the correction is large and single-spike formulas mislead; a testable prediction is that the E-type confidence interval for the smaller spike degrades as $l_k \to l_j$, a quality the paper does not examine.
  • The moment estimators rely on leave-one-out inverses and become unstable when $p$ is close to $n$; a practical extension would be to replace them by ridge-regularized inverses, mirroring the pseudo-inverse treatment already used for $p>n$ in the paper.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 8 minor

Summary. The paper claims to close the open problem, posed by Yang and Johnstone (2018), of first-order Edgeworth corrections for spiked eigenvalues of sample covariance matrices with non-Gaussian entries. Theorem 1 (multi-spike) and Theorem 2 (single-spike) assert sup_x |P(R_k <= x) - Phi(x) - n^{-1/2}[(1/6)kappa_{2,k}^{-3/2}kappa_{3,k}(1-x^2) - kappa_{2,k}^{-1/2}(mu(g_{nk}) + A(g_{nk}))]phi(x)| = o(n^{-1/2}) with explicit formulas for kappa_{2,k}, kappa_{3,k}, mu(g_{nk}), and A(g_{nk}). Theorem 3 provides consistent estimators of the fourth- and sixth-moment parameters beta_z, Gamma, and Delta; Theorem 4 claims O(n^{-1/2}) and o(n^{-1/2}) coverage errors for the Z-type and E-type pivots; Section 3.2 constructs a spike-number estimator; Section 4 reports extensive simulations. The proof is organized as three steps: Theorem 5 approximates R_n by a linear statistic of the non-Gaussian matrix, Theorem 6 asserts an o(n^{-1/2}) non-Gaussian/Gaussian universality via a partial generalized four moment theorem, and Theorem 7 gives the conditional Edgeworth expansion for the Gaussian statistic. The critical load-bearing step is Theorem 6, whose statement conflicts with the paper's own variance computation in Section A.5.

Significance. If the main results held, this would be a useful contribution: explicit, parameter-free Edgeworth corrections for spiked eigenvalues in non-Gaussian models, consistent cumulant estimators, and a credible route to confidence intervals and spike counting. The paper deserves credit for writing out the correction terms (including the cross-spike interaction A(g_{nk})), for honestly reporting in Section 4.2 that estimated Edgeworth coefficients can degrade performance, and for structuring a serious proof attempt rather than a heuristic derivation. The decisive universality claim of Theorem 6 is, however, contradicted by the paper's own variance formula in Section A.5: the non-Gaussian and Gaussian statistics compared there have asymptotic variances differing by beta_z l^{-2}, an O(1) gap, so the claimed o(n^{-1/2}) equivalence cannot hold. Theorem 6 is the bridge in the proofs of Theorems 1 and 2 (equations (1.1)-(1.3) and (1.4)-(1.6)), and its proof also depends on an unproved rate assertion (Remark 9). The main theorems, and the applications built on them (Theorem 4 and Section 3), are therefore not established by the proof given.

major comments (3)
  1. [Theorem 6; Section 5.2; Section A.5; Section A.6] The stress-test concern about an O(1) variance gap is confirmed on reading the paper. Theorem 6 compares the unstandardized statistics tilde S_n(g_n) + n^{-1/2} tilde S_n(g_n h_n) rho_n^{-1} tilde sigma_n x and S_n(g_n) + n^{-1/2} S_n(g_n h_n) rho_n^{-1} tilde sigma_n x, and claims their distribution functions differ by o(n^{-1/2}). But from the identity Omega(rho_n,Z) = -rho_n tilde S_n(g_n) in Section 5.2 and the paper's own variance computation at the end of Section A.5, Var(Omega(rho_n,Z)) = 2 rho^2 F(g^2) + rho^2 beta_z F^2(g), hence Var(tilde S_n(g_n)) tends to 2F(g^2) + beta_z F^2(g), whereas the Gaussian term S_n(g_n) = n^{-1/2} sum g_n(lambda_i)(omega_i^2 - 1) has Var tending to 2F(g^2) because beta_z = 0 for Y. The gap beta_z F^2(g) = beta_z l^{-2} is O(1). Consequently the two CDFs in Theorem 6 have sup-norm distance bounded away from zero; for instance, the Gaussian event in Section 5.2 has limiting probability Phi(a x) with a = sqrt(1 + beta_z l^{-2} sigma_n^2/4) different from 1, while the non-Gaussian event has limit Phi(x). The proof in Section A.6 says it 'matches the first four moments of z_k and y_k,' but E z_k^4 = 3 + beta_z differs from E y_k^4 = 3, so the fourth moment does not match, and the telescoping characteristic-function argument cannot remove an O(1) variance difference. Since Theorem 6 is the step converting the non-Gaussian statistic to its Gaussian counterpart in equations (1.4)-(1.6) of Section A.2 and (1.1)-(1.3) of Section A.1, the proofs of Theorem 2 and Theorem 1 do not go through.
  2. [Remark 9; Section A.6; Remark 10] Even setting aside the variance problem, the proof of Theorem 6 does not establish the claimed rate. The final telescoping step in Section A.6 asserts o(n^{-1/2}) 'established through Lemmas 1, 2 and 3, along with Remark 10,' but Lemma 1 gives only E_k(alpha_{ki0}) = O(n^{-1/2}) for the individual conditional expectations, and summing such bounds over k would leave a contribution of order n^{1/2} without cancellation. The decisive cancellation that would close the argument is placed entirely on the unproved assertion in Remark 10 that alpha_{ki0} - alpha_{ki0y} is of order O(n^{-1}) 'demonstrated' by Jiang and Bai (2021b), with no derivation given. Remark 9, which is the actual load-bearing rate assumption of Theorem 6, states without proof that the conclusion of the partial generalized four moment theorem 'remains valid, and the asymptotic error bound can be shown to be of order o(n^{-1/2}).' The appendices repeatedly defer bounds with phrases such as 'the proofs are similar' (Lemmas 1-3 in Section A.6) and 'using the same method' (Appendix B.4), so the missing rate is not documented elsewhere in the manuscript. This is an omitted proof of a load-bearing step.
  3. [Theorem 4; Section A.4] Theorem 4(2) claims a coverage error of o(n^{-1/2}) for the E-type pivot, but the proof in Section A.4 obtains the bound |u^E_n(hat rho_k, l_k) - bar F_{kn}(hat rho_k, l)| <= C_n n^{-1/2} with only the sentence 'building upon our theoretical framework established in previous sections.' This bound is essentially the uniform Edgeworth approximation of Theorem 1 itself, so Theorem 4(2) inherits the failure of Theorem 1 identified above. The argument also does not show why the post-selection conditioning on hat l_k > theta_n preserves the o(n^{-1/2}) rate rather than only the O(n^{-1/2}) rate claimed for the Z-type pivot, and the displayed probability calculation does not by itself establish the conditional claim. The confidence-interval construction in Section 3.1 and the spike-number estimator in Section 3.2 therefore rest on results that are not established.
minor comments (8)
  1. [Throughout] There are numerous typos: 'eatimator' (Section 3.1), 'converagence' (Remark 6), 'varibales' (Theorem 7), 'Corolllary' (Lemma 6), and 'compansion' for 'companion' (repeatedly in Section A.1); the paper would benefit from a careful proofreading pass.
  2. [Section 2.2; Remark 5] Remark 5 cites 'Theorem 2.7 of Zheng et al. (2019)', but the reference list contains Zheng, Bai and Yao (2015) and Zhang, Hu and Bai (2019), and no Zheng et al. (2019); the citation should be corrected.
  3. [Theorem 6 statement] Theorem 6 compares statistics that contain a fixed x while asserting convergence 'for any t in R'; the theorem should state explicitly that x is fixed and whether the rate is uniform in x, since the application in Theorem 5 needs uniformity in x over the real line.
  4. [Section A.5, final display] In the final display of the proof of Theorem 5, the term 'n^{-1/2} tilde S_n(g_n)' appears twice with different implied coefficients, and the intermediate algebra leading to the displayed equation is difficult to verify; the authors should check for a typesetting slip and expand the derivation.
  5. [Section A.7, condition R3] The verification of condition R3 of Lemma 6 reads 'n^{1/2} integral_{|t|>epsilon} |t|^{-1} |E exp(it bar V_n^{-1/2} sum X_{ni})| dt <= n^{1/2} integral_{|t|>epsilon} |t|^{-1} delta_n^i dt = o(1)' with delta_n^i undefined; since R3 is genuinely needed for the o(n^{-1/2}) rate, this bound needs a real argument rather than an undefined symbol.
  6. [Section 4.2; Tables 3-6] In Tables 3-6 the Y&J-E method reports 0% estimation accuracy in several high-dimensional settings (for example Table 3, rows (80,400), (60,200), and (120,400)), and accuracy is sometimes non-monotone in n; the discussion in Section 4.2 attributes this to coefficient estimation difficulty but does not explain the mechanism or the non-monotonicity, which is important for assessing the practical claims.
  7. [Remark 3] Remark 3 states that kappa_{2,k} and kappa_{3,k} are 'not the exact conditional cumulants of Z11 but rather carefully constructed approximations,' while Theorem 1 asserts an o(n^{-1/2}) expansion with these exact formulas; the paper should clarify in what sense the approximate cumulants are sufficient for the claimed exact rate, since a reader cannot tell from the text whether a remainder has been absorbed.
  8. [Section A.3, near equation (1.9)] Equation (1.9) contains the term -(Delta + 12 hat beta_z + 6/(1 - gamma_n) - 15 hat beta_z - 21 + 8(1 + gamma_n)/(1 - gamma_n)^2)^2, whose signs on the beta_z terms are opposite to those in the definition (2.5); as written the display also appears to square a random variable (it contains hat beta_z) rather than a constant, so the formula for E(hat Delta - Delta)^2 should be checked.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the Edgeworth derivation is assembled from external, independently checkable results rather than from the paper's own target claim.

full rationale

The claimed derivation is not circular. Theorem 1 and Theorem 2 are obtained through three distinct steps: a characteristic-equation reduction to linear spectral statistics (Theorem 5), a universality comparison between non-Gaussian and Gaussian statistics (Theorem 6), and a conditional Edgeworth expansion for sums of independent variables in the Gaussian case (Theorem 7), the last of which invokes Petrov (1975) and Yang and Johnstone (2018). Each of these steps uses external results whose stated assumptions do not include the target Edgeworth expansion. The paper does rely on Jiang and Bai (2021a, 2021b) for the o(n^{-1/2}) universality rate asserted in Theorem 6, Remark 9, and Remark 10; since Bai is a co-author, this is a self-citation, and it is load-bearing for the non-Gaussian extension. However, those cited results are peer-reviewed, parameter-free statements about four-moment universality, not the paper's own fitted values or its own conclusion, so under the review rules they count as independent support rather than circularity. The proof of Theorem 6 also contains a gap: the key refinement in Remark 10 is attributed to Jiang and Bai (2021b) with the detailed cancellation deferred, and the paper's own variance computation Var(Omega(rho_n,Z)) = 2rho^2 F(g^2) + rho^2 beta_z F^2(g) suggests a possible incompatibility with the asserted o(n^{-1/2}) closeness of unstandardized statistics. That is a correctness concern, not a circular reduction of the expansion to its own inputs. The moment estimators, confidence intervals, and spike-number estimator use the derived Edgeworth formula as a plug-in, which is standard inference rather than fitting a parameter and then renaming the fit a prediction.

Assumptions & free parameters 2 free parameters · 6 assumptions · 0 invented entities

No free parameters are fit in the core theorems; the expansion coefficients are functions of the population moments. The applied spike-number estimation introduces hand-tuned r0 and heuristic bootstrap estimates. The main unproved input is the o(n^{-1/2}) four-moment universality rate.

free parameters (2)
  • Initial number of spikes r0 = 5
    Section 4.2 selects r0 = 5 by hand for the iterative spike-number estimation; results are not shown to be insensitive to this choice.
  • Bootstrap-estimated Edgeworth coefficients
    Section 4.2 estimates the correction coefficients via a bootstrap procedure 'similar to Xie and Zhang (2024)', with initial spike values obtained by expanding the sample to size 1000 through resampling; the procedure is heuristic and its variability is not quantified.
assumptions (6)
  • domain assumption EZ11=0, EZ11^2=1, EZ11^4<infinity, EZ11^6<infinity (Assumption a)
    Section 2.1; finite sixth moment is required for the third cumulant of squared observations in the first-order Edgeworth term.
  • domain assumption gamma_n = p/n -> gamma in (0, infinity) (Assumption b)
    Section 2.1; standard high-dimensional asymptotic regime.
  • domain assumption l_r > 1 + sqrt(gamma) (Assumption c)
    Section 2.1; phase transition condition ensuring the spike separates from the bulk.
  • domain assumption Cramer's condition for Z11 (Assumption d)
    Section 2.1; standard smoothness condition for Edgeworth expansions.
  • ad hoc to paper Partial generalized four moment theorem extends with rate o(n^{-1/2}) (Remark 9)
    Section 5.2; the paper relies on an unproved rate refinement of Jiang and Bai (2021b) to swap non-Gaussian and Gaussian linear statistics; this is the load-bearing step for Theorem 6.
  • standard math Edgeworth expansion for sums of independent random variables (Petrov 1975; Lemma 6 from Yang and Johnstone 2018)
    Section 5.3 and Appendix A.7; standard result applied to the conditional sum.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Edgeworth corrections for the spiked eigenvalues of non-Gaussian sample covariance matrices with applications." pith.science (2026). https://pith.science/paper/ARY37DKG

@misc{pith2026250709584,
  author       = {Pith},
  title        = {Pith review of: Edgeworth corrections for the spiked eigenvalues of non-Gaussian sample covariance matrices with applications},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ARY37DKG}},
  note         = {Machine review of arXiv:2507.09584}
}
read the original abstract

Yang and Johnstone (2018) established an Edgeworth correction for the largest sample eigenvalue in a spiked covariance model under the assumption of Gaussian observations, leaving the extension to non-Gaussian settings as an open problem. In this paper, we address this issue by establishing first-order Edgeworth expansions for spiked eigenvalues in both single-spike and multi-spike scenarios with non-Gaussian data. Leveraging these expansions, we construct more accurate confidence intervals for the population spiked eigenvalues and propose a novel estimator for the number of spikes. Simulation studies demonstrate that our proposed methodology outperforms existing approaches in both robustness and accuracy across a wide range of settings, particularly in low-dimensional cases.

Figures

Figures reproduced from arXiv: 2507.09584 by the authors.

Figure 1
Figure 1. Edgeworth expansion for different samples under Setting 1 [PITH_FULL_IMAGE:figures/full_fig_p011_1.png] view at source ↗
Figure 2
Figure 2. Edgeworth expansion for different samples under Setting 2 [PITH_FULL_IMAGE:figures/full_fig_p011_2.png] view at source ↗
Figure 3
Figure 3. Edgeworth expansion for different samples under Setting 3 [PITH_FULL_IMAGE:figures/full_fig_p055_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Edgeworth expansion for different samples under Setting 4 [PITH_FULL_IMAGE:figures/full_fig_p055_4.png]
Figure 5
Figure 5. Figure 5: Edgeworth expansion for different samples under Setting 5 [PITH_FULL_IMAGE:figures/full_fig_p056_5.png]
Figure 6
Figure 6. Figure 6: Edgeworth expansion for different samples under Setting 6 [PITH_FULL_IMAGE:figures/full_fig_p056_6.png]
Figure 7
Figure 7. Figure 7: Edgeworth expansion for different samples under Setting 7 [PITH_FULL_IMAGE:figures/full_fig_p057_7.png]
Figure 8
Figure 8. Figure 8: Edgeworth expansion for different samples under Setting 8 [PITH_FULL_IMAGE:figures/full_fig_p057_8.png]
Figure 9
Figure 9. Figure 9: Edgeworth expansion for different samples under Setting 9 [PITH_FULL_IMAGE:figures/full_fig_p058_9.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

44 extracted references · 39 canonical work pages

  1. [1]

    P., and Fujikoshi, Y

    Bai, Z., Choi, K. P., and Fujikoshi, Y. (2018). Consistency of aic and bic in estimating the number of significant components in high-dimensional principal component analysis. The Annals of Statistics , 46(3):1050--1076

  2. [2]

    and Ding, X

    Bai, Z. and Ding, X. (2012). Estimation of spiked eigenvalues in spiked models. Random Matrices: Theory and Applications , 1(02):1150011

  3. [3]

    and Hu, J

    Bai, Z. and Hu, J. and Pan, G. and Zhou, W. (2015). Convergence of the empirical spectral distribution function of Beta matrices. Bernoulli , 21(3):1538--1574

  4. [4]

    and Miao, B

    Bai, Z. and Miao, B. and Pan, G. (2007). On asymptotics of eigenvectors of large sample covariance matrix. The Annals of Probability , 35(4):1532--1572

  5. [5]

    and Silverstein, J

    Bai, Z. and Silverstein, J. W. (2004). CLT for linear spectral statistics of large-dimensional sample covariance matrices . The Annals of Statistics , 32(1A):553--605

  6. [6]

    and Silverstein, J

    Bai, Z. and Silverstein, J. W. (2010). Spectral analysis of large dimensional random matrices . Springer, New York

  7. [7]

    and Yao, J

    Bai, Z. and Yao, J. (2008). Central limit theorems for eigenvalues in a spiked population model. In Annales de l'IHP Probabilit \'e s et Statistiques , volume 44, pages 447--474

  8. [8]

    and Yao, J

    Bai, Z. and Yao, J. (2012). On sample eigenvalues in a generalized spiked population model. Journal of Multivariate Analysis , 106:167--177

Show all 44 references
  1. [9]

    Baik, J., Ben Arous, G., and P \'e ch \'e , S. (2005). Phase transition of the largest eigenvalue for nonnull complex sample covariance matrices. The Annals of Probability , 33(5):1643--1697

  2. [10]

    Bao, Z., Hu, J., Pan, G., and Zhou, W. (2019). Canonical correlation coefficients of high-dimensional Gaussian vectors: Finite rank case . The Annals of Statistics , 47(1):612--640

  3. [11]

    Berk, R., Brown, L., Buja, A., Zhang, K., and Zhao, L. (2013). Valid post-selection inference. The Annals of Statistics , 41(2):802--837

  4. [12]

    T., Han, X., and Pan, G

    Cai, T. T., Han, X., and Pan, G. (2020). Limiting laws for divergent spiked eigenvalues and largest nonspiked eigenvalue of sample covariance matrices. The Annals of Statistics , 48(3):1255--1280

  5. [13]

    and Ji, H

    Ding, X. and Ji, H. C. (2023). Spiked multiplicative random matrices and principal components. Stochastic Processes and their Applications , 163:25--60

  6. [14]

    and Paul, D

    D \"o rnemann, N. and Paul, D. (2024). Detecting spectral breaks in spiked covariance models. arXiv preprint arXiv:2404.19176

  7. [15]

    Hall, P. (2013). The bootstrap and Edgeworth expansion . Springer Science & Business Media

  8. [16]

    Huang, H. (2017). Asymptotic behavior of support vector machine for spiked population model. Journal of Machine Learning Research , 18(45):1--21

  9. [17]

    and Bai, Z

    Jiang, D. and Bai, Z. (2021a). Generalized four moment theorem and an application to CLT for spiked eigenvalues of high-dimensional covariance matrices . Bernoulli , 27(1):274--294

  10. [18]

    and Bai, Z

    Jiang, D. and Bai, Z. (2021b). Partial generalized four moment theorem revisited. Bernoulli , 27(4):2337--2352

  11. [19]

    Jiang, D., Hou, Z., and Hu, J. (2021). The limits of the sample spiked eigenvalues for a high-dimensional generalized fisher matrix and its applications. Journal of Statistical Planning and Inference , 215:208--217

  12. [20]

    Johnstone, I. M. (2001). On the distribution of the largest eigenvalue in principal components analysis. The Annals of Statistics , 29(2):295--327

  13. [21]

    Johnstone, I. M. and Nadler, B. (2017). Roy’s largest root test under rank-one alternatives. Biometrika , 104(1):181--193

  14. [22]

    Johnstone, I. M. and Paul, D. (2018). Pca in high dimensions: An orientation. Proceedings of the IEEE , 106(8):1277--1292

  15. [23]

    Lamarre, P. Y. G. and Shkolnikov, M. (2019). Edge of spiked beta ensembles, stochastic airy semigroups and reflected brownian motions. Annales de l’Institut Henri Poincaré , 55:1402--1438

  16. [24]

    D., Sun, D

    Lee, J. D., Sun, D. L., Sun, Y., and Taylor, J. E. (2016). Exact post-selection inference, with application to the lasso. The Annals of Statistics , 44(3):907--927

  17. [25]

    Li, Z., Han, F., and Yao, J. (2020). Asymptotic joint distribution of extreme eigenvalues and trace of large sample covariance matrix in a generalized spiked population model. The Annals of Statistics , 48(6):3138--3160

  18. [26]

    Passemier, D., Li, Z., and Yao, J. (2017). On estimation of the noise variance in high dimensional probabilistic principal component analysis. Journal of the Royal Statistical Society Series B: Statistical Methodology , 79(1):51--67

  19. [27]

    and Yao, J

    Passemier, D. and Yao, J. (2012). On determining the number of spikes in a high-dimensional spiked population model. Random Matrices: Theory and Applications , 1(01):1150002

  20. [28]

    Paul, D. (2007). Asymptotics of sample eigenstructure for a large dimensional spiked covariance model. Statistica Sinica , 17:1617--1642

  21. [29]

    P \'e ch \'e , S. (2003). Universality of local eigenvalue statistics for random sample covariance matrices . PhD thesis, Ecole Polytechnique F\' e d\' e rale de Lausanne

  22. [30]

    Petrov, V. V. (1975). Sums of independent random variables . Springer, New York

  23. [31]

    Shen, D., Shen, H., Zhu, H., and Marron, J. (2016). The statistics and mathematics of high dimension low sample size asymptotics. Statistica Sinica , 26(4):1747

  24. [32]

    Tang, R., Yuan, M., and Zhang, A. R. (2025). Mode-wise principal subspace pursuit and matrix spiked covariance model. Journal of the Royal Statistical Society Series B: Statistical Methodology , 87(1):232--255

  25. [33]

    W., and Yao, J.-f

    Wang, Q., Silverstein, J. W., and Yao, J.-f. (2014). A note on the clt of the lss for sample covariance matrix from a spiked population model. Journal of Multivariate Analysis , 130:194--207

  26. [34]

    and Yao, J

    Wang, Q. and Yao, J. (2013). On the sphericity test with large-dimensional observations. Electronic Journal of Statistics , 7:2164--2192

  27. [35]

    and Yao, J

    Wang, Q. and Yao, J. (2017). Extreme eigenvalues of large-dimensional spiked Fisher matrices with application . The Annals of Statistics , 45(1):415--460

  28. [36]

    and Fan, J

    Wang, W. and Fan, J. (2017). Asymptotics of empirical eigenstructure for high dimensional spiked covariance. Annals of Statistics , 45(3):1342--1374

  29. [37]

    and Zhang, Y

    Xie, F. and Zhang, Y. (2024). Higher-order entrywise eigenvectors analysis of low-rank random matrices: Bias correction, edgeworth expansion, and bootstrap. arXiv preprint arXiv:2401.15033

  30. [38]

    Xie, J., Zeng, Y., and Zhu, L. (2021). Limiting laws for extreme eigenvalues of large-dimensional spiked Fisher matrices with a divergent number of spikes . Journal of Multivariate Analysis , 184:104742

  31. [39]

    Yang, J. (2019). Edgeworth Approximations for Spiked PCA Models and Applications . PhD thesis, Stanford University

  32. [40]

    and Johnstone, I

    Yang, J. and Johnstone, I. M. (2018). Edgeworth correction for the largest eigenvalue in a spiked PCA model . Statistica Sinica , 28(4):2541--2564

  33. [41]

    R., Cai, T

    Zhang, A. R., Cai, T. T., and Wu, Y. (2022a). Heteroskedastic pca: Algorithm, optimality, and applications. The Annals of Statistics , 50(1):53--80

  34. [42]

    and Hu, J

    Zhang, Q. and Hu, J. and Bai, Z. (2019). Invariant test based on the modified correction to LRT for the equality of two high-dimensional covariance matrices . Electronic Journal of Statistics , 13:850--881

  35. [43]

    Zheng, S., Bai, Z., and Yao, J. (2015). Substitution principle for CLT of linear spectral statistics of high-dimensional sample covariance matrices with applications to hypothesis testing . The Annals of Statistics , 43(2):546--591

  36. [44]

    Zhang, Z., Zheng, S., Pan, G., and Zhong, P.-S. (2022b). Asymptotic independence of spiked eigenvalues and linear spectral statistics for large sample covariance matrices. The Annals of Statistics , 50(4):2205--2230

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.