Pith. sign in

REVIEW 3 major objections 5 minor 17 references

Limiting eigen-structure of spiked sample covariance matrices under missing observations

T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Under entrywise missing observations, distant spiked sample eigenvalues and eigenvectors still have Gaussian limits, with shift and covariance formulas that differ from the complete-data case.

desk verdict First CLTs for spiked eigenvalues/eigenvectors under Bernoulli missingness; core theory plausible, but the application section has an unproved plug-in theorem and a narrower-than-claimed power analysis. read the letter →

arxiv 2608.09135 v1 pith:5SSFUW6J submitted 2026-08-10 math.ST stat.TH

classification math.STstat.TH MSC 62H1562B2062D10
keywords centrallimittheoremhigh-dimensionalstatisticsmissingdatasamplecovariancematrixspikedeigenvalueseigenvectorsindependencetest
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks what happens to principal-component analysis when a high-dimensional dataset has entries missing in a random coordinate-wise fashion, and it models missingness by multiplying each observation by an independent Bernoulli mask with zeros in unobserved entries. Its central claim is that in the spiked population model the distant sample spiked eigenvalues still converge almost surely, and after centering by a modified spike map they obey a Gaussian central limit theorem, while the normalized spiked eigenvectors are asymptotically Gaussian around the eigenvectors of a modified population covariance. The limiting parameters are not the complete-data ones: missingness changes the effective covariance, the centering map, and the variance. If correct, standard eigenvalue-based inference, including confidence intervals and tests of independent block structure, can be extended to incomplete high-dimensional data using the formulas supplied. The paper also constructs a chi-squared independence test for two blocks of variables and shows it has asymptotic power against the modeled dependence alternative.

What carries the argument

The machinery is the effective covariance matrix \tilde{\Sigma} = P\Sigma P + (P - $P^{2}$) \circ \Sigma, which converts the Bernoulli-masked problem into a no-missing problem with a modified population, together with the spike-to-eigenvalue map \psi(\tilde{\$\alpha$}) = \tilde{\$\alpha$}(1 + c \int t/(\tilde{\$\alpha$} - t) H(dt)) whose derivative sign separates distant spikes from closed ones. The Gaussian limit is carried by an N × N symmetric random matrix C(\psi(\tilde{\$\alpha$}_k)) whose covariance, specified in Theorem 2.5, is expressed through \tau_0, \nu_0 and mixed fourth moments such as E(y_{j1} y_{s1} y_{\ell 1} y_{t1}). The proofs condition on the resolvent of the bulk block, extract the deterministic part [1 + c f_1(\$\lambda$)] \tilde{\Sigma}, apply a central limit theorem for random quadratic forms, and use an eigenvector perturbation resolvent bound; the application then turns the joint normal vector from Theorem 2.7 into a chi-square statistic.

What would settle it

Simulate a one-spike model with n = p = 500 under the paper's assumptions, record the empirical variance of \sqrt{n}(\lambda_{n,1} - \psi(\tilde{\$\alpha$}_1)) over many replications, and compare it with the variance from Theorem 2.5; then repeat the same simulation but make the Bernoulli mask depend on |x_{ij}| so that missingness is not completely at random. A systematic mismatch in the second setting, while the first matches, would confirm that the missing-completely-at-random, zero-imputation premise is the load-bearing condition.

Watch

Extended reading notes

Core claim

Under a block-diagonal spiked population, the sample covariance matrix built from y = b ◦ x behaves spectrally as if the population covariance were \tilde{\Sigma} = P \Sigma P + (P - $P^{2}$) \circ \Sigma, where P = diag(\theta_i) and \circ is the Hadamard product. For a distant spike \tilde{\$\alpha$}_k with \psi'(\tilde{\$\alpha$}_k) > 0, where \psi(\tilde{\$\alpha$}) = \tilde{\$\alpha$}(1 + c \int t/(\tilde{\$\alpha$} - t) H(dt)), Proposition 2.1 gives \lambda_{n,j} \to \psi(\tilde{\$\alpha$}_k) almost surely. Theorem 2.5 states that \sqrt{n}(\lambda_{n,j} - \psi(\tilde{\$\alpha$}_k)) converges in distribution to the eigenvalues of an explicit Gaussian random matrix whose covariance involves fourth-order moments of the observed entries, the effective covariance \tilde{\Sigma}, and the spectral functionals \tau_0 and \nu_0; Theorem 2.9 gives \sqrt{n}(a_k - \tilde{e}v_k) \Rightarrow N(0, \Gamma_k) for the normalized eigenvectors. The paper's key message is that normality survives missingness, but the shift and the spread are governed by the missing rates and the effective spectrum, not by the original spectrum.

Load-bearing premise

The whole analysis rests on the assumption that each entry is missing completely at random and independently of the data, with missing entries recorded as zero; if missingness is correlated with the value that would have been observed, or with other coordinates, or if a different imputation is used, the stated limits no longer apply.

Editorial extensions

If this is right

  • Confidence intervals and hypothesis tests for spiked eigenvalues can be built from the new centering \psi(\tilde{\alpha}_k) and the covariance in Theorem 2.5 instead of the complete-data formulas.
  • The independence test statistic T_N is asymptotically \chi^2(N) under the null and has power tending to 1 against the modeled dependence alternative, giving a practical tool for block independence in incomplete data.
  • Eigenvector-based inference, such as subspace estimation, must use the effective eigenvectors \tilde{e}v_k and the covariance \Gamma_k from Theorem 2.9; the plain sample eigenvector is biased by a factor governed by 1/(1 + c f_2(\psi(\tilde{\alpha}_k)) \tilde{\alpha}_k).
  • Uniform missingness does not simply rescale the spectrum: it can shift a spike below the phase-transition boundary and turn a spiked sample eigenvalue into a bulk eigenvalue.
  • The delta-method estimator \hat{T}_N with plug-in estimates of \psi and \tilde{\alpha} retains the chi-square limit, allowing the test to be applied without knowing the population parameters.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the proof reduces the masked problem to an effective covariance, the same Gaussian-limit structure should survive under more general independent-coordinate masks, such as deterministic missing proportions or heterogeneous Bernoulli rates, with only the effective covariance and fourth-moment terms changed; the paper does not pursue this generalization.
  • The variance formulas suggest a testable extension: comparing the empirical variance of \sqrt{n}(\lambda_{n,1} - \psi(\tilde{\alpha}_1)) across different missing rates could be used to estimate the missingness probabilities, although inversion is not developed here.
  • For practitioners, the paper implies that applying classical PCA after zero-filling and using complete-data CLT formulas will misstate uncertainty whenever missingness is non-negligible; the corrected formulas provide the benchmark that simpler approaches would need to match.
  • The framework deliberately excludes missing-not-at-random mechanisms and non-zero imputation; adapting the limiting formulas to those settings is a natural next step, and the covariance formulas here give a concrete reference point for such extensions.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper studies the asymptotic eigen-structure of sample covariance matrices under a spiked population model with entrywise missing observations. Missingness is modeled as y = b ◦ x, where b is an independent Bernoulli mask and x has an independent-component structure, so the effective population covariance is tilde Σ = PΣP + (P − P^2) ◦ Σ. Under assumptions on the effective spiked eigenvalues, the authors state an almost sure convergence result for distant spiked eigenvalues (Proposition 2.1), a central limit theorem for the spiked eigenvalues (Theorems 2.5 and 2.7), and CLTs for the associated eigenvectors (Theorems 2.8 and 2.9), with explicit limiting covariance matrices. The paper then applies these results to propose a chi-square statistic for testing independence between two groups of variables (Theorem 4.1) and a plug-in version (Theorem 4.3), supported by simulations in Section 3 and Section 4.2.

Significance. If the main CLT results are correct, the paper makes a useful contribution by extending spiked PCA inference to data with missing-at-random entries and by showing that the missingness mechanism changes the limiting variances and the eigenvector basis in a nontrivial way. The explicit covariance formulas in Theorems 2.5 and 2.9 are concrete and falsifiable, and the simulations in Section 3 provide a prima facie check of the Gaussian approximations. The reliance on the published machinery of Li et al. (2024) is appropriate in substance, though the paper should clearly delineate which statements are new. However, the application section has a serious internal inconsistency: the power analysis in Theorem 4.2 is conducted under a block-diagonal alternative that corresponds to independence in the Section 2 model, not the stated dependence alternative H1. In addition, the proofs of Theorem 2.7 and Theorem 4.3 are not supplied, and the core CLT lemma is only available in a supplement with a placeholder DOI. These gaps prevent the paper from being accepted in its current form; they are fixable within the manuscript's scope, so I recommend major revision.

major comments (3)
  1. [Section 4.1, Theorem 4.2] The power analysis does not address the stated alternative H1: x_N is not independent of x_S. In the Section 2 model, x = Σ^{1/2} \tilde x with \tilde x having independent entries, so if Σ = diag(Σ_{N+L}, Σ_{S−L}) as assumed in the proof of Theorem 4.2, then x_N and x_S are independent, i.e., the proof is analyzing a variant of H0, not H1. The argument that some eα_j^{[N+L]} > eα_j shows only that the test can detect an increase in the number of large eigenvalues in the first block, which is a change in the marginal covariance of x_N, not a cross-covariance between x_N and x_S. The power simulation in Section 4.2 sets pairwise covariance 0.5 between variables in x_N and x_S, which is an off-diagonal perturbation not covered by Theorem 4.2. Therefore the claimed consistency of T_N and \hat T_N against dependence is unsupported; the theorem should either be proved for a genuine cross-covariance alternative or the application should be reframed.
  2. [Section 5.1 and Theorem 4.3] The proof of Theorem 2.7 is omitted: Section 5.1 states that the proof is similar to that of Theorem 2.5 and therefore omitted. This is load-bearing because Theorem 2.7 provides the joint CLT for the N spiked eigenvalues that underlies the chi-square statistic in Theorem 4.1. Similarly, Theorem 4.3, the plug-in version of the test statistic used in the simulations, is stated without proof. The claimed χ²(N) limit for \hat T_N involves estimation of eα_i, ev_i, ψ, K_4, and the covariance matrix; the delta-method assertion needs a careful derivation, especially because the estimation errors enter at the √n scale. Please provide complete proofs or detailed, verifiable derivations for both results.
  3. [Section 6, Lemma 6.1] Lemma 6.1 is the main ingredient for the CLT in Theorem 2.5, but its proof is not in the manuscript; it is relegated to the supplementary material [9], whose DOI is given as 'to be provided by the typesetter'. For a submitted manuscript, this is not a minor presentation issue: the core lemma's proof must be available in a permanent supplement or in the paper itself. Please include a complete proof or ensure that the supplement is accessible.
minor comments (5)
  1. [Remark 2.3] Remark 2.3 says that examples are given below Theorem 2.9, but the examples actually appear in Remark 2.4, before Theorem 2.5; the cross-reference should be corrected.
  2. [Proof of Theorem 2.5] The covariance formula for C(λ) is written using E(y_{j1} y_{s1} y_{ℓ1} y_{t1}), while the proof displays E(d_{j1}d_{s1}d_{ℓ1}d_{t1} x_{j1}x_{s1}x_{ℓ1}x_{t1}); the equivalence relies on independence of the mask D and the underlying X and should be stated explicitly.
  3. [Table 1] The simulated values of T_2 deviate from the theoretical values by up to about 0.03 in the Gamma and t scenarios (e.g., 0.8529 vs. 0.8825 in column (2)); the paper should comment on whether this reflects finite-sample bias or a missing term in the approximation.
  4. [Section 3] The simulation generates the missing probabilities as random draws (1−θ_i ∼ Uniform(0,0.5)), whereas Section 2 treats θ_i as fixed constants; please clarify whether θ_i are fixed per replication or resampled, and how this affects the reported empirical quantities.
  5. [Proposition 2.1] The second part of Proposition 2.1 uses the symbol α̃_k, but this quantity is never defined; it should be eα_k or a clearly defined alternative.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the spiked CLTs are derived from an independent bulk-spectrum input, not from the target result.

full rationale

The central derivation chain is not circular. Proposition 2.1 and Theorem 2.5 expand A_n(lambda) around the deterministic equivalent [1+c f_1(lambda)]tilde-Sigma_N, with the Gaussian fluctuation C_n(lambda) obtained from Lemmas 6.1 and 6.2; the target normality of the spiked eigenvalues appears as the conclusion, not as an input. The bulk ingredients imported from [13] (ESD convergence, eigenvalue support, and the o_p(n^{-1/2}) trace CLT) are published, peer-reviewed theorems about the non-spiked Gram matrix under missing observations. Although [13] shares two co-authors with the present paper, those results do not state or presuppose the spiked CLT, so the citation is independent support rather than self-referential loading. The plug-in construction of hat-T_N uses standard delta-method and nuisance-estimation steps from [16] and [17] after Theorem 4.1, again without feeding the target distribution back into the assumptions. The only omitted proof, Lemma 6.1, is deferred to the supplement [9] and is part of the paper's own derivation rather than an imported assumption. The main weakness is not circular: Theorem 4.2's power analysis is stated for a block-diagonal N+L / S-L alternative, and Theorem 4.5 inherits this, while the declared H1 is non-independence; this is a correctness or coverage concern, not a by-construction reduction of the result to its inputs.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The paper introduces no new free parameters or invented entities. Its central results rest on the Bernoulli missingness model, the block-diagonal spiked covariance structure, and standard asymptotic conditions, plus the published random-matrix machinery of Li et al. (2024). The main unstated cost is the reliance on self-cited technical results and the unproved plug-in test theorem.

assumptions (5)
  • domain assumption The missing data mechanism is independent Bernoulli scaling: y = b ◦ x with b_i ~ Ber(theta_i), independent of x.
    Section 2, model definition. This gives the effective covariance tilde Sigma = P Sigma P + (P-P^2) ◦ Sigma. If missingness depends on data values or coordinates, the analysis does not apply.
  • domain assumption The complete data has block-diagonal covariance Sigma = diag(Sigma_N, Sigma_S) with N fixed and an independent component structure x = Sigma^{1/2} \bar x with iid entries.
    Section 2. This is the classical spiked model framework; it makes the spiked and bulk parts independent.
  • domain assumption Assumptions (a)-(d): p/n -> c in (0, infinity), finite fourth moment, ESD of tilde Sigma_S converges to H, and distant spikes satisfy psi'(tilde alpha_k) > 0.
    Section 2. These are standard asymptotic conditions for the central limit theorems to hold.
  • standard math The CLT results for linear spectral statistics and quadratic forms of Gram matrices with missing data from Li et al. (2024), Theorems 2.1-2.3 and Section 2.5, are taken as input.
    Used throughout the proofs in Section 5. These are published results from Ann. Statist. 52 (2024), by overlapping authors.
  • standard math The perturbation lemma for eigenvectors (Lemma 6.4) from Johnstone and Yang (2018) is assumed.
    Used in the proof of Theorem 2.9 for the first-order eigenvector perturbation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Limiting eigen-structure of spiked sample covariance matrices under missing observations." pith.science (2026). https://pith.science/paper/5SSFUW6J

@misc{pith2026260809135,
  author       = {Pith},
  title        = {Pith review of: Limiting eigen-structure of spiked sample covariance matrices under missing observations},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5SSFUW6J}},
  note         = {Machine review of arXiv:2608.09135}
}
read the original abstract

High-dimensional Principal Component Analysis (PCA) has become an essential tool in modern data analysis, offering dimensionality reduction and feature extraction. However, the presence of missing data introduces significant challenges, distorting the performance of PCA and complicating statistical inference. In this paper, we study the asymptotic behavior of PCA under a spiked population model with missing observations, leveraging recent advances in random matrix theory. We demonstrate that while the spiked sample eigenvalues exhibit asymptotic normality, the limiting parameters differ substantially from those in the complete data case, reflecting the non?trivial influence of the missing data mechanism. As an application of our results, we propose a test to evaluate the independent structure of a spiked population.

Figures

Figures reproduced from arXiv: 2608.09135 by the authors.

Figure 1
Figure 1. Q-Q plot of the empirical distribution of [PITH_FULL_IMAGE:figures/full_fig_p008_1.png] view at source ↗
Figure 2
Figure 2. Q-Q plot of the empirical distribution of [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

17 extracted references · 16 canonical work pages

  1. [9]

    Limiting Eigen-Structure of Spiked Sample Covariance Matrices under Missing Observations

    CHENG, H., LI, H., YIN, Y. and ZHANG, Z. (2026). Supplement to “Limiting Eigen-Structure of Spiked Sample Covariance Matrices under Missing Observations”. DOI to be provided by the typesetter

  2. [1]

    ANDERSON, T. W. (1958).An Introduction to Multivariate Statistical Analysis. Wiley, New York. 24Cheng, Li, Yin and Zhang

  3. [2]

    BAI, Z. D. and SILVERSTEIN, J. W. (2004). CLT for linear spectral statistics of large-dimensional sample covariance matrices.Ann. Probab.32553–605

  4. [3]

    BAI, Z. D. and YAO, J. F. (2008). Central limit theorems for eigenvalues in a spiked population model.Ann. Inst. Henri Poincaré Probab. Stat.44447–474

  5. [4]

    BAI, Z. D. and YAO, J. F. (2012). On sample eigenvalues in a generalized spiked population.J. Multivariate Anal.106167–177

  6. [5]

    and PÉCHÉ, S

    BAIK, J., BENAROUS, G. and PÉCHÉ, S. (2005). Phase transition of the largest eigenvalue for nonnull complex sample covariance matrices.Ann. Probab.331643–1697

  7. [6]

    and SILVERSTEIN, J

    BAIK, J. and SILVERSTEIN, J. W. (2006). Eigenvalues of large sample covariance matrices of spiked popu- lation models.J. Multivariate Anal.971382–1408

  8. [7]

    and WANG, K

    BAO, Z., DING, X., WANG, J. and WANG, K. (2022). Statistical inference for principal components of spiked covariance matrices.Ann. Statist.501144–1169

Show all 17 references
  1. [8]

    and PAROLYA, N

    BODNAR, T., DETTE, H. and PAROLYA, N. (2019). Testing for independence of large dimensional vectors. Ann. Statist.472977–3008

  2. [10]

    and BAI, Z

    JIANG, D. and BAI, Z. (2021). Generalized four moment theorem and an application to CLT for spiked eigenvalues of large-dimensional covariance matrices.Bernoulli27274–294

  3. [11]

    JOHNSTONE, I. M. (2001). On the distribution of the largest eigenvalue in principal components analysis. Ann. Statist.29295–327

  4. [12]

    JOHNSTONE, I. M. and YANG, J. (2018). Notes on asymptotics of sample eigenstructure for spiked covari- ance models with non-Gaussian data.arXiv preprint arXiv:1810.10427

  5. [13]

    and ZHOU, W

    LI, H., PAN, G., YIN, Y. and ZHOU, W. (2024). Spectral analysis of Gram matrices with missing at random observations: Convergence, central limit theorems, and applications in statistical inference.Ann. Statist.52 1254–1275

  6. [14]

    M., MCKAY, M

    MORALES-JIMENEZ, D., JOHNSTONE, I. M., MCKAY, M. R. and YANG, J. (2021). Asymptotics of eigen- structure of sample correlation matrices for high-dimensional spiked models.Statist. Sinica31571–601

  7. [15]

    PAUL, D. (2007). Asymptotics of sample eigenstructure for a large dimensional spiked covariance model. Statist. Sinica171617–1642

  8. [16]

    WATERNAUX, C. M. (1976). Asymptotic distribution of the sample roots for a nonnormal population. Biometrika63639–645

  9. [17]

    and ZHONG, P

    ZHANG, Z., ZHENG, S., PAN, G. and ZHONG, P. S. (2022). Asymptotic independence of spiked eigenvalues and linear spectral statistics for large sample covariance matrices.Ann. Statist.502205–2230

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.