primePCA iteratively imputes missing entries via projection onto current principal component estimates and updates the estimate with the leading right singular space, achieving geometric error convergence in the noiseless case under an incoherence condition when signal strength is sufficient.
Heteroskedastic PCA: Algorithm, Optimality, and Applications
2 Pith papers cite this work. Polarity classification is still indexing.
abstract
A general framework for principal component analysis (PCA) in the presence of heteroskedastic noise is introduced. We propose an algorithm called HeteroPCA, which involves iteratively imputing the diagonal entries of the sample covariance matrix to remove estimation bias due to heteroskedasticity. This procedure is computationally efficient and provably optimal under the generalized spiked covariance model. A key technical step is a deterministic robust perturbation analysis on singular subspaces, which can be of independent interest. The effectiveness of the proposed algorithm is demonstrated in a suite of problems in high-dimensional statistics, including singular value decomposition (SVD) under heteroskedastic noise, Poisson PCA, and SVD for heteroskedastic and incomplete data.
representative citing papers
A robust, heteroskedastic matrix factorization method generalizes PCA to handle per-feature uncertainties, missing data, and outlier detection via Student-t likelihood iterative reweighting.
citing papers explorer
-
High-dimensional principal component analysis with heterogeneous missingness
primePCA iteratively imputes missing entries via projection onto current principal component estimates and updates the estimate with the leading right singular space, achieving geometric error convergence in the noiseless case under an incoherence condition when signal strength is sufficient.
-
Robust Heteroskedastic Matrix Factorization: A Generalization of PCA that Flags Outliers and Handles Missing Data
A robust, heteroskedastic matrix factorization method generalizes PCA to handle per-feature uncertainties, missing data, and outlier detection via Student-t likelihood iterative reweighting.