REVIEW 3 major objections 5 minor 17 references
Limiting eigen-structure of spiked sample covariance matrices under missing observations
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Under entrywise missing observations, distant spiked sample eigenvalues and eigenvectors still have Gaussian limits, with shift and covariance formulas that differ from the complete-data case.
desk verdict First CLTs for spiked eigenvalues/eigenvectors under Bernoulli missingness; core theory plausible, but the application section has an unproved plug-in theorem and a narrower-than-claimed power analysis. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the effective covariance matrix \tilde{\Sigma} = P\Sigma P + (P - $P^{2}$) \circ \Sigma, which converts the Bernoulli-masked problem into a no-missing problem with a modified population, together with the spike-to-eigenvalue map \psi(\tilde{\$\alpha$}) = \tilde{\$\alpha$}(1 + c \int t/(\tilde{\$\alpha$} - t) H(dt)) whose derivative sign separates distant spikes from closed ones. The Gaussian limit is carried by an N × N symmetric random matrix C(\psi(\tilde{\$\alpha$}_k)) whose covariance, specified in Theorem 2.5, is expressed through \tau_0, \nu_0 and mixed fourth moments such as E(y_{j1} y_{s1} y_{\ell 1} y_{t1}). The proofs condition on the resolvent of the bulk block, extract the deterministic part [1 + c f_1(\$\lambda$)] \tilde{\Sigma}, apply a central limit theorem for random quadratic forms, and use an eigenvector perturbation resolvent bound; the application then turns the joint normal vector from Theorem 2.7 into a chi-square statistic.
What would settle it
Simulate a one-spike model with n = p = 500 under the paper's assumptions, record the empirical variance of \sqrt{n}(\lambda_{n,1} - \psi(\tilde{\$\alpha$}_1)) over many replications, and compare it with the variance from Theorem 2.5; then repeat the same simulation but make the Bernoulli mask depend on |x_{ij}| so that missingness is not completely at random. A systematic mismatch in the second setting, while the first matches, would confirm that the missing-completely-at-random, zero-imputation premise is the load-bearing condition.
Extended reading notes
Core claim
Under a block-diagonal spiked population, the sample covariance matrix built from y = b ◦ x behaves spectrally as if the population covariance were \tilde{\Sigma} = P \Sigma P + (P - $P^{2}$) \circ \Sigma, where P = diag(\theta_i) and \circ is the Hadamard product. For a distant spike \tilde{\$\alpha$}_k with \psi'(\tilde{\$\alpha$}_k) > 0, where \psi(\tilde{\$\alpha$}) = \tilde{\$\alpha$}(1 + c \int t/(\tilde{\$\alpha$} - t) H(dt)), Proposition 2.1 gives \lambda_{n,j} \to \psi(\tilde{\$\alpha$}_k) almost surely. Theorem 2.5 states that \sqrt{n}(\lambda_{n,j} - \psi(\tilde{\$\alpha$}_k)) converges in distribution to the eigenvalues of an explicit Gaussian random matrix whose covariance involves fourth-order moments of the observed entries, the effective covariance \tilde{\Sigma}, and the spectral functionals \tau_0 and \nu_0; Theorem 2.9 gives \sqrt{n}(a_k - \tilde{e}v_k) \Rightarrow N(0, \Gamma_k) for the normalized eigenvectors. The paper's key message is that normality survives missingness, but the shift and the spread are governed by the missing rates and the effective spectrum, not by the original spectrum.
Load-bearing premise
The whole analysis rests on the assumption that each entry is missing completely at random and independently of the data, with missing entries recorded as zero; if missingness is correlated with the value that would have been observed, or with other coordinates, or if a different imputation is used, the stated limits no longer apply.
Editorial extensions
If this is right
- Confidence intervals and hypothesis tests for spiked eigenvalues can be built from the new centering \psi(\tilde{\alpha}_k) and the covariance in Theorem 2.5 instead of the complete-data formulas.
- The independence test statistic T_N is asymptotically \chi^2(N) under the null and has power tending to 1 against the modeled dependence alternative, giving a practical tool for block independence in incomplete data.
- Eigenvector-based inference, such as subspace estimation, must use the effective eigenvectors \tilde{e}v_k and the covariance \Gamma_k from Theorem 2.9; the plain sample eigenvector is biased by a factor governed by 1/(1 + c f_2(\psi(\tilde{\alpha}_k)) \tilde{\alpha}_k).
- Uniform missingness does not simply rescale the spectrum: it can shift a spike below the phase-transition boundary and turn a spiked sample eigenvalue into a bulk eigenvalue.
- The delta-method estimator \hat{T}_N with plug-in estimates of \psi and \tilde{\alpha} retains the chi-square limit, allowing the test to be applied without knowing the population parameters.
Reading between the lines
- Because the proof reduces the masked problem to an effective covariance, the same Gaussian-limit structure should survive under more general independent-coordinate masks, such as deterministic missing proportions or heterogeneous Bernoulli rates, with only the effective covariance and fourth-moment terms changed; the paper does not pursue this generalization.
- The variance formulas suggest a testable extension: comparing the empirical variance of \sqrt{n}(\lambda_{n,1} - \psi(\tilde{\alpha}_1)) across different missing rates could be used to estimate the missingness probabilities, although inversion is not developed here.
- For practitioners, the paper implies that applying classical PCA after zero-filling and using complete-data CLT formulas will misstate uncertainty whenever missingness is non-negligible; the corrected formulas provide the benchmark that simpler approaches would need to match.
- The framework deliberately excludes missing-not-at-random mechanisms and non-zero imputation; adapting the limiting formulas to those settings is a natural next step, and the covariance formulas here give a concrete reference point for such extensions.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies the asymptotic eigen-structure of sample covariance matrices under a spiked population model with entrywise missing observations. Missingness is modeled as y = b ◦ x, where b is an independent Bernoulli mask and x has an independent-component structure, so the effective population covariance is tilde Σ = PΣP + (P − P^2) ◦ Σ. Under assumptions on the effective spiked eigenvalues, the authors state an almost sure convergence result for distant spiked eigenvalues (Proposition 2.1), a central limit theorem for the spiked eigenvalues (Theorems 2.5 and 2.7), and CLTs for the associated eigenvectors (Theorems 2.8 and 2.9), with explicit limiting covariance matrices. The paper then applies these results to propose a chi-square statistic for testing independence between two groups of variables (Theorem 4.1) and a plug-in version (Theorem 4.3), supported by simulations in Section 3 and Section 4.2.
Significance. If the main CLT results are correct, the paper makes a useful contribution by extending spiked PCA inference to data with missing-at-random entries and by showing that the missingness mechanism changes the limiting variances and the eigenvector basis in a nontrivial way. The explicit covariance formulas in Theorems 2.5 and 2.9 are concrete and falsifiable, and the simulations in Section 3 provide a prima facie check of the Gaussian approximations. The reliance on the published machinery of Li et al. (2024) is appropriate in substance, though the paper should clearly delineate which statements are new. However, the application section has a serious internal inconsistency: the power analysis in Theorem 4.2 is conducted under a block-diagonal alternative that corresponds to independence in the Section 2 model, not the stated dependence alternative H1. In addition, the proofs of Theorem 2.7 and Theorem 4.3 are not supplied, and the core CLT lemma is only available in a supplement with a placeholder DOI. These gaps prevent the paper from being accepted in its current form; they are fixable within the manuscript's scope, so I recommend major revision.
major comments (3)
- [Section 4.1, Theorem 4.2] The power analysis does not address the stated alternative H1: x_N is not independent of x_S. In the Section 2 model, x = Σ^{1/2} \tilde x with \tilde x having independent entries, so if Σ = diag(Σ_{N+L}, Σ_{S−L}) as assumed in the proof of Theorem 4.2, then x_N and x_S are independent, i.e., the proof is analyzing a variant of H0, not H1. The argument that some eα_j^{[N+L]} > eα_j shows only that the test can detect an increase in the number of large eigenvalues in the first block, which is a change in the marginal covariance of x_N, not a cross-covariance between x_N and x_S. The power simulation in Section 4.2 sets pairwise covariance 0.5 between variables in x_N and x_S, which is an off-diagonal perturbation not covered by Theorem 4.2. Therefore the claimed consistency of T_N and \hat T_N against dependence is unsupported; the theorem should either be proved for a genuine cross-covariance alternative or the application should be reframed.
- [Section 5.1 and Theorem 4.3] The proof of Theorem 2.7 is omitted: Section 5.1 states that the proof is similar to that of Theorem 2.5 and therefore omitted. This is load-bearing because Theorem 2.7 provides the joint CLT for the N spiked eigenvalues that underlies the chi-square statistic in Theorem 4.1. Similarly, Theorem 4.3, the plug-in version of the test statistic used in the simulations, is stated without proof. The claimed χ²(N) limit for \hat T_N involves estimation of eα_i, ev_i, ψ, K_4, and the covariance matrix; the delta-method assertion needs a careful derivation, especially because the estimation errors enter at the √n scale. Please provide complete proofs or detailed, verifiable derivations for both results.
- [Section 6, Lemma 6.1] Lemma 6.1 is the main ingredient for the CLT in Theorem 2.5, but its proof is not in the manuscript; it is relegated to the supplementary material [9], whose DOI is given as 'to be provided by the typesetter'. For a submitted manuscript, this is not a minor presentation issue: the core lemma's proof must be available in a permanent supplement or in the paper itself. Please include a complete proof or ensure that the supplement is accessible.
minor comments (5)
- [Remark 2.3] Remark 2.3 says that examples are given below Theorem 2.9, but the examples actually appear in Remark 2.4, before Theorem 2.5; the cross-reference should be corrected.
- [Proof of Theorem 2.5] The covariance formula for C(λ) is written using E(y_{j1} y_{s1} y_{ℓ1} y_{t1}), while the proof displays E(d_{j1}d_{s1}d_{ℓ1}d_{t1} x_{j1}x_{s1}x_{ℓ1}x_{t1}); the equivalence relies on independence of the mask D and the underlying X and should be stated explicitly.
- [Table 1] The simulated values of T_2 deviate from the theoretical values by up to about 0.03 in the Gamma and t scenarios (e.g., 0.8529 vs. 0.8825 in column (2)); the paper should comment on whether this reflects finite-sample bias or a missing term in the approximation.
- [Section 3] The simulation generates the missing probabilities as random draws (1−θ_i ∼ Uniform(0,0.5)), whereas Section 2 treats θ_i as fixed constants; please clarify whether θ_i are fixed per replication or resampled, and how this affects the reported empirical quantities.
- [Proposition 2.1] The second part of Proposition 2.1 uses the symbol α̃_k, but this quantity is never defined; it should be eα_k or a clearly defined alternative.
Circularity Check
No circularity: the spiked CLTs are derived from an independent bulk-spectrum input, not from the target result.
full rationale
The central derivation chain is not circular. Proposition 2.1 and Theorem 2.5 expand A_n(lambda) around the deterministic equivalent [1+c f_1(lambda)]tilde-Sigma_N, with the Gaussian fluctuation C_n(lambda) obtained from Lemmas 6.1 and 6.2; the target normality of the spiked eigenvalues appears as the conclusion, not as an input. The bulk ingredients imported from [13] (ESD convergence, eigenvalue support, and the o_p(n^{-1/2}) trace CLT) are published, peer-reviewed theorems about the non-spiked Gram matrix under missing observations. Although [13] shares two co-authors with the present paper, those results do not state or presuppose the spiked CLT, so the citation is independent support rather than self-referential loading. The plug-in construction of hat-T_N uses standard delta-method and nuisance-estimation steps from [16] and [17] after Theorem 4.1, again without feeding the target distribution back into the assumptions. The only omitted proof, Lemma 6.1, is deferred to the supplement [9] and is part of the paper's own derivation rather than an imported assumption. The main weakness is not circular: Theorem 4.2's power analysis is stated for a block-diagonal N+L / S-L alternative, and Theorem 4.5 inherits this, while the declared H1 is non-independence; this is a correctness or coverage concern, not a by-construction reduction of the result to its inputs.
Assumptions & free parameters
assumptions (5)
- domain assumption The missing data mechanism is independent Bernoulli scaling: y = b ◦ x with b_i ~ Ber(theta_i), independent of x.
- domain assumption The complete data has block-diagonal covariance Sigma = diag(Sigma_N, Sigma_S) with N fixed and an independent component structure x = Sigma^{1/2} \bar x with iid entries.
- domain assumption Assumptions (a)-(d): p/n -> c in (0, infinity), finite fourth moment, ESD of tilde Sigma_S converges to H, and distant spikes satisfy psi'(tilde alpha_k) > 0.
- standard math The CLT results for linear spectral statistics and quadratic forms of Gram matrices with missing data from Li et al. (2024), Theorems 2.1-2.3 and Section 2.5, are taken as input.
- standard math The perturbation lemma for eigenvectors (Lemma 6.4) from Johnstone and Yang (2018) is assumed.
Cite this review
Pith. "Pith review of Limiting eigen-structure of spiked sample covariance matrices under missing observations." pith.science (2026). https://pith.science/paper/5SSFUW6J
@misc{pith2026260809135,
author = {Pith},
title = {Pith review of: Limiting eigen-structure of spiked sample covariance matrices under missing observations},
year = {2026},
howpublished = {\url{https://pith.science/paper/5SSFUW6J}},
note = {Machine review of arXiv:2608.09135}
}
read the original abstract
High-dimensional Principal Component Analysis (PCA) has become an essential tool in modern data analysis, offering dimensionality reduction and feature extraction. However, the presence of missing data introduces significant challenges, distorting the performance of PCA and complicating statistical inference. In this paper, we study the asymptotic behavior of PCA under a spiked population model with missing observations, leveraging recent advances in random matrix theory. We demonstrate that while the spiked sample eigenvalues exhibit asymptotic normality, the limiting parameters differ substantially from those in the complete data case, reflecting the non?trivial influence of the missing data mechanism. As an application of our results, we propose a test to evaluate the independent structure of a spiked population.
Figures
Reference graph
Works this paper leans on
-
[9]
Limiting Eigen-Structure of Spiked Sample Covariance Matrices under Missing Observations
CHENG, H., LI, H., YIN, Y. and ZHANG, Z. (2026). Supplement to “Limiting Eigen-Structure of Spiked Sample Covariance Matrices under Missing Observations”. DOI to be provided by the typesetter
work page 2026
-
[1]
ANDERSON, T. W. (1958).An Introduction to Multivariate Statistical Analysis. Wiley, New York. 24Cheng, Li, Yin and Zhang
work page 1958
-
[2]
BAI, Z. D. and SILVERSTEIN, J. W. (2004). CLT for linear spectral statistics of large-dimensional sample covariance matrices.Ann. Probab.32553–605
work page 2004
-
[3]
BAI, Z. D. and YAO, J. F. (2008). Central limit theorems for eigenvalues in a spiked population model.Ann. Inst. Henri Poincaré Probab. Stat.44447–474
work page 2008
-
[4]
BAI, Z. D. and YAO, J. F. (2012). On sample eigenvalues in a generalized spiked population.J. Multivariate Anal.106167–177
work page 2012
-
[5]
BAIK, J., BENAROUS, G. and PÉCHÉ, S. (2005). Phase transition of the largest eigenvalue for nonnull complex sample covariance matrices.Ann. Probab.331643–1697
work page 2005
-
[6]
BAIK, J. and SILVERSTEIN, J. W. (2006). Eigenvalues of large sample covariance matrices of spiked popu- lation models.J. Multivariate Anal.971382–1408
work page 2006
-
[7]
BAO, Z., DING, X., WANG, J. and WANG, K. (2022). Statistical inference for principal components of spiked covariance matrices.Ann. Statist.501144–1169
work page 2022
Show all 17 references
-
[8]
and PAROLYA, N
BODNAR, T., DETTE, H. and PAROLYA, N. (2019). Testing for independence of large dimensional vectors. Ann. Statist.472977–3008
2019
-
[10]
and BAI, Z
JIANG, D. and BAI, Z. (2021). Generalized four moment theorem and an application to CLT for spiked eigenvalues of large-dimensional covariance matrices.Bernoulli27274–294
2021
-
[11]
JOHNSTONE, I. M. (2001). On the distribution of the largest eigenvalue in principal components analysis. Ann. Statist.29295–327
2001
-
[12]
JOHNSTONE, I. M. and YANG, J. (2018). Notes on asymptotics of sample eigenstructure for spiked covari- ance models with non-Gaussian data.arXiv preprint arXiv:1810.10427
2018 arXiv
-
[13]
and ZHOU, W
LI, H., PAN, G., YIN, Y. and ZHOU, W. (2024). Spectral analysis of Gram matrices with missing at random observations: Convergence, central limit theorems, and applications in statistical inference.Ann. Statist.52 1254–1275
2024
-
[14]
M., MCKAY, M
MORALES-JIMENEZ, D., JOHNSTONE, I. M., MCKAY, M. R. and YANG, J. (2021). Asymptotics of eigen- structure of sample correlation matrices for high-dimensional spiked models.Statist. Sinica31571–601
2021
-
[15]
PAUL, D. (2007). Asymptotics of sample eigenstructure for a large dimensional spiked covariance model. Statist. Sinica171617–1642
2007
-
[16]
WATERNAUX, C. M. (1976). Asymptotic distribution of the sample roots for a nonnormal population. Biometrika63639–645
1976
-
[17]
and ZHONG, P
ZHANG, Z., ZHENG, S., PAN, G. and ZHONG, P. S. (2022). Asymptotic independence of spiked eigenvalues and linear spectral statistics for large sample covariance matrices.Ann. Statist.502205–2230
2022
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.