REVIEW 3 major objections 4 minor 5 cited by
Estimating shared subspace with AJIVE: the power and limitation of multiple data matrices
T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read This paper establishes a two-regime theory for AJIVE shared-subspace estimation: high signal-to-noise gives a minimax-optimal $1/\sqrt K$ error decay, while low signal-to-noise leaves a non-diminishing error that also plagues an…
desk verdict Solid high-SNR minimax theory for AJIVE; the low-SNR 'fundamental barrier' is honestly labeled in the body but overclaimed in the abstract. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the averaged projection matrix $M := \frac{1}{K}\sum_{k=1}^K \widehat U_k \widehat U_k^\top$, formed from per-matrix estimates of the top $(r+r_k)$ left singular subspaces; AJIVE outputs the top-$r$ eigenspace of $M$. The analysis separates the noiseless leading term $U^\star U^{\star\top} + \frac{1}{K}\sum_k U_k^\star U_k^{\star\top}$ from a perturbation $\Delta$ built from first-order SVD expansions, and the misalignment parameter $\theta$ controls the eigengap of that leading term. For the oracle lower bound the machinery is a fourth-order polynomial approximation $Q$ of the oracle matrix $M$, whose expectation contains a cross term $\alpha_4 \frac{1}{K}\sum_k (U^\star U_k^{\star\top} + U_k^\star U^{\star\top})$ with $\alpha_4 \asymp \sigma^4 nd$; because $V^{\star\top}W^\star = 0.6I$ in the hard configuration, this cross term tilts the leading eigenspace away from $U^\star$ and the tilt does not shrink as $K$ grows.
What would settle it
Run a low-SNR sequence with $K$ from 25 to 10,000 on the paper's hard configuration ($V_k^{\star\top}W_k^\star = 0.6I$, $\sigma\sqrt n/\sigma_{\min}$ small, $r\le n/3$) and test a full joint nonconvex estimator, not the oracle-initialized one-step estimator: if its error converges to 0, the claimed fundamental barrier is false. Since the paper explicitly lacks an information-theoretic lower bound in this regime, a minimax calculation that forces positive error for all estimators would be the definitive confirmation.
Extended reading notes
Core claim
On the paper's own terms, the discovery is a complete first-order characterization of AJIVE for the JIVE model $A_k = U^\star V_k^{\star\top} + U_k^\star W_k^{\star\top} + E_k$ with i.i.d. Gaussian noise of variance $\sigma^2$. Theorem 1 (simplified at equation (8)) says that under the high-SNR condition $\sigma\sqrt n/\sigma_{\min} \ll \min\{\sqrt\theta,\sqrt{K\theta}\}$ the estimation error $\|\widehat U\widehat U^\top - U^\star U^{\star\top}\|$ is bounded by $(\sigma/\sigma_{\min})(\sqrt{n/K} + \sqrt{r/(K\theta)})$ plus a second-order term; Theorem 2 shows the first-order rate is minimax optimal, with the rate separating a large-$\theta$ regime where unique subspaces barely interfere from a small-$\theta$ regime where their alignment degrades the rate. In the complementary low-SNR regime the second-order term dominates and converges to $\frac{1}{\theta}\frac{\sigma^2 n}{\sigma_{\min}^2}$ as $K\to\infty$. Theorem 3 constructs a worst-case loading configuration $V_k^{\star\top}W_k^\star = 0.6I$ on which even the oracle estimator, namely one alternating-minimization step initialized at the true $U^\star$, has error bounded below by $C\sigma^4 n^2/\sigma_{\min}^4$, so the non-diminishing error is presented as a structural barrier rather than an artifact of AJIVE.
Load-bearing premise
The load-bearing premise is that the oracle-aided spectral estimator, namely one alternating-minimization step initialized at the true shared subspace, represents the fundamental low-SNR limit; the paper states it cannot prove an information-theoretic lower bound, so any estimator outside this spectral class that drives the error to zero would overturn the barrier conclusion.
Editorial extensions
If this is right
- In the high-SNR regime, averaging more matrices improves shared-subspace estimation at the minimax-optimal rate $1/\sqrt K$, so additional views give real accuracy gains.
- The optimal rate worsens as the unique subspaces become more aligned: the $r/(K\theta)$ term shows that near-identifiability, not noise alone, controls difficulty.
- In the low-SNR regime, AJIVE's error has a component $\frac{1}{\theta}\frac{\sigma^2 n}{\sigma_{\min}^2}$ that survives $K\to\infty$; beyond a threshold the number of matrices is irrelevant to the error floor.
- An oracle-aided spectral estimator, given the true unique components in the first stage, still fails on a worst-case alignment, indicating the floor is shared by spectral methods rather than a quirk of AJIVE.
- The persistent error fits the classical incidental-parameter phenomenon: the shared subspace is the structural parameter and the per-matrix unique subspaces are incidental parameters that contaminate likelihood-type estimation.
Reading between the lines
- A likely testable extension is that full alternating minimization, which re-estimates $U$ jointly with the unique components rather than stopping after one step from the truth, may behave differently from the oracle estimator in the low-SNR regime; the paper only analyzes the one-step version.
- Because the theory assumes equal noise variance across matrices and known ranks, an immediate extension would allow per-matrix $\sigma_k^2$; the second-stage average would then need inverse-variance weighting, which may shift or remove the non-diminishing term.
- The hard configuration uses adversarial loadings with $V^{\star\top}W^\star = 0.6I$; the numerical experiments show the low-SNR floor under shared loading while random loadings show the $1/\sqrt K$ decay, suggesting that alignment between $V_k^\star$ and $W_k^\star$, not just subspace misalignment $\theta$, controls whether the floor appears.
- A practical diagnostic suggested by the theory: estimate $\theta$ and the SNR before deciding whether to pool more matrices; when $\theta$ and SNR are both small, the shared subspace should be treated as only partially identified.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies the AJIVE algorithm for estimating the shared column subspace in the JIVE model with K noisy data matrices. The main theoretical contribution is a two-regime characterization: in a high-SNR regime, the estimation error of AJIVE is shown to decay as σ/σmin times roughly sqrt(n/K) plus a misalignment-dependent term, and a matching minimax lower bound is established; in a low-SNR regime, a non-diminishing error term is identified. Since the paper does not prove an information-theoretic lower bound for this low-SNR barrier, it instead proves a lower bound for an oracle-aided spectral estimator (Theorem 3) and presents numerical evidence. The paper also contains a detailed proof appendix based on SVD perturbation expansions, concentration inequalities, and Fano/packing arguments.
Significance. The high-SNR results are a solid contribution: Theorem 1 provides a nearly complete rate characterization of AJIVE, and Theorem 2 shows the first-order rate is minimax-optimal. The dependence on the misalignment parameter θ and the explicit benefit of multiple matrices (1/sqrt(K) rate) are new and quantitatively useful. The paper also ships detailed, self-contained proofs, including a nontrivial Fano construction for the θ-dependent term that may be of independent interest. The low-SNR part is suggestive but not at the same level of rigor: Theorem 3 is an algorithm-specific lower bound for an oracle-initialized spectral estimator, not a minimax lower bound, and the paper's own caveats in Sections 5 and 7 are more accurate than the abstract. If the claims are appropriately qualified, the paper is a valuable contribution to the theory of multi-matrix subspace estimation.
major comments (3)
- [Section 3, Eqs. (6)-(8)] The abstract and Section 1 state that in low-SNR settings AJIVE exhibits a non-diminishing error as K grows, but Theorem 1 does not prove this. Theorem 1 is an upper bound that holds only under the high-SNR conditions (6a)-(6b); when σ√n/σmin is not ≪ √θ, those conditions fail and the bound, including the simplified expression (8), is not a valid consequence. The discussion immediately after (8) interprets the second-order term E2 as a low-SNR non-diminishing error, but this is an extrapolation outside the theorem's assumptions. Please either prove a low-SNR lower bound for AJIVE itself or explicitly label this statement as a conjecture supported by the oracle-estimator analysis and the simulations.
- [Section 5, Theorem 3] The phrase 'fundamental barrier' in the abstract and in the contribution list is stronger than what Theorem 3 establishes. Theorem 3 is a lower bound for the oracle-aided spectral estimator (Algorithm 2) on a specially constructed configuration with V⋆ᵀW⋆ = 0.6I; it is not a minimax lower bound over all estimators, and it does not rule out non-spectral methods that could have vanishing error as K grows. The paper itself concedes this in Section 5 ('we are unable to deliver an information-theoretic limit') and Section 7 ('might be a fundamental limit'). The abstract, Section 1, and the concluding remarks should be aligned with these qualifications; the current wording overstates the strength of Theorem 3.
- [Section 5.3] The connection to Neyman and Scott's problem is presented as if the oracle estimator being 'one step of alternating minimization starting from the ground truth' makes it representative of the MLE. This is an analogy, not a proof: Algorithm 2 is neither the nonconvex least-squares estimator nor a full alternating-minimization run, and the paper does not show inconsistency of the MLE itself. The text should clearly state that the Neyman-Scott connection is heuristic, or provide an analysis of the actual nonconvex estimator.
minor comments (4)
- [Abstract] The abstract's phrase 'a fundamental barrier' is inconsistent with Section 7's accurate wording 'might be a fundamental limit'; please use the qualified version throughout.
- [Section D.3.1] The statement 'Without loss of generality, assume n is divisible by 3 and K is even' is not a true WLOG; rounding or a padding argument should be supplied, or the assumption should be stated as a technical condition.
- [Section C.1.1] In the proof of (22) and (23), the sentence 'Moreover, this truncation is tight...' is followed by a probability statement for a deterministic expectation; the phrase 'with probability at least 1 - O(N^-11)' around (23b) is misleading since (23b) is an exact identity, not a high-probability event.
- [Section 6, Figure 3] The figure caption states the y-axis is in log scale but does not specify whether the x-axis is √K or log √K; adding the axis scaling in the caption would improve readability.
Circularity Check
No significant circularity: the derivation chains are independent; the low-SNR 'fundamental barrier' is an overclaim, but not a circular step.
full rationale
This paper is a proof-based analysis and does not derive its conclusions from its own inputs by construction. The central upper-bound result (Theorem 1) is established through a first-order SVD perturbation expansion attributed to the external result [Xia21], followed by matrix concentration inequalities and a spectral perturbation argument. The minimax lower bound (Theorem 2) uses Fano's method with explicit packing constructions, and the constants and rates are obtained from the constructions, not from the target upper bound. Theorem 3 is explicitly an algorithm-specific lower bound for the oracle-aided spectral estimator; the paper honestly states that it cannot prove an information-theoretic limit ('we are unable to deliver an information-theoretic limit') and only conjectures that the non-diminishing error 'might be a fundamental limit.' That the abstract uses the stronger phrase 'fundamental barrier' is a scope/correctness concern about overclaiming, not a circularity. The self-citation to [CCFM21] (a monograph coauthored by Cong Ma) is used for standard truncated matrix Bernstein inequalities and a geometric lemma; these are established published results with stated assumptions, not special premises whose validity depends on the present paper's claims. The simulations generate θ-misaligned subspaces via the randomized construction in Section 2.2, but that construction is a direct implementation of Definition 1, not a fitted parameter renamed as a prediction. No fitted quantity is later called a prediction, and no uniqueness theorem from the authors' prior work is used to force the choice of estimator. Therefore the derivation chain is self-contained against its declared inputs, and no circular step can be exhibited with the specificity required by the review rules.
Assumptions & free parameters
free parameters (2)
- cross-loading constant 0.6 in the Theorem 3 configuration =
0.6
- unique-component signal scale gamma =
0.5
assumptions (6)
- domain assumption Noise entries are i.i.d. Gaussian with common variance σ² across all K matrices
- domain assumption Ranks r and rk are known exactly
- domain assumption The unique subspaces are θ-misaligned with θ > 0, so ∩k col(U⋆k) = ∅
- standard math The SVD perturbation expansion of Xia (2021), Theorem 1, is valid and its remainder is controlled by conditions (6a)-(6b)
- standard math Standard concentration and information-theoretic tools: matrix Bernstein, Vershynin bounds, Gilbert-Varshamov, Fano's method
- ad hoc to paper The oracle spectral estimator (one step of alternating minimization from the truth) is representative of achievable low-SNR performance
Cite this review
Pith. "Pith review of Estimating shared subspace with AJIVE: the power and limitation of multiple data matrices." pith.science (2026). https://pith.science/paper/Z3I3XEZX
@misc{pith2026250109336,
author = {Pith},
title = {Pith review of: Estimating shared subspace with AJIVE: the power and limitation of multiple data matrices},
year = {2026},
howpublished = {\url{https://pith.science/paper/Z3I3XEZX}},
note = {Machine review of arXiv:2501.09336}
}
read the original abstract
Integrative data analysis often requires disentangling joint and individual variations across multiple datasets, a challenge commonly addressed by the Joint and Individual Variation Explained (JIVE) model. While numerous methods have been developed to estimate the shared subspace under JIVE, the theoretical understanding of their performance remains limited, particularly in the context of multiple matrices and varying degrees of subspace misalignment. This paper bridges this gap by providing a systematic analysis of shared subspace estimation in multi-matrix settings. We focus on the Angle-based Joint and Individual Variation Explained (AJIVE) method, a two-stage spectral approach, and establish new performance guarantees that uncover its strengths and limitations. Specifically, we show that in high signal-to-noise ratio (SNR) regimes, AJIVE's estimation error decreases with the number of matrices, demonstrating the power of multi-matrix integration. Conversely, in low-SNR settings, AJIVE exhibits a non-diminishing error, highlighting fundamental limitations. To complement these results, we derive minimax lower bounds, showing that AJIVE achieves optimal rates in high-SNR regimes. Furthermore, we analyze an oracle-aided spectral estimator to demonstrate that the non-diminishing error in low-SNR scenarios is a fundamental barrier. Extensive numerical experiments corroborate our theoretical findings, providing insights into the interplay between SNR, the number of matrices, and subspace misalignment.
Figures
Figures from the paper (1 more)
Forward citations
Cited by 5 Pith papers
-
Stacked SVD or SVD stacked? A Random Matrix Theory perspective on data integration
In the proportional high-dimensional limit, the paper derives exact squared-overlap formulas and phase transitions for Stack-SVD and SVD-Stack, and proves optimally weighted Stack-SVD always beats optimally weighted S...
-
Transfer Learning in High-Dimensional Clustering: Minimax Thresholds and Applications in Single-Cell Data
In high-d two-community GMMs, consistent target clustering via transfer is possible iff either the target SNR clears the usual (d/n)^{1/4} barrier or the source is strong and aligned enough that µ∆_T, ∆_S, and µ∆_S∆_T...
-
Unveiling Invariant and Transferable Latent Factors Across Heterogeneous Environments via ATLAS
ATLAS is a multi-environment factor model estimator that separates invariant (shared-loading) factors from environment-specific factors and uses auxiliary labels to align and transfer prediction-relevant signals.
-
Spectral Joint Subspace Estimation for Heterogeneous Multi-View Data: Geometry and Reweighting
Geometry of the individual components, not the number of views, determines whether AJIVE's error vanishes at the K^{-1/2} rate; a weighted variant handles heterogeneous views.
-
A functional tensor model for dynamic multilayer networks with common invariant subspaces and the RKHS estimation
A functional tensor model with common invariant subspaces and RKHS estimation for analyzing dynamic multilayer networks, applied to bike-share and food-trade data.
Reference graph
Works this paper leans on
-
[1]
Spectral methods for data science: A statistical perspective
Yuxin Chen, Yuejie Chi, Jianqing Fan, and Cong Ma. Spectral methods for data science: A statistical perspective. Foundations and Trends in Machine Learning , 14(5):566--806, 2021
work page 2021
-
[2]
Changxiao Cai, Gen Li, Yuejie Chi, H. Vincent Poor, and Yuxin Chen. Subspace estimation from unbalanced and incomplete data matrices: _ 2, statistical guarantees. Ann. Statist. , 49(2):944--967, 2021
work page 2021
-
[3]
Tony Cai, Zongming Ma, and Yihong Wu
T. Tony Cai, Zongming Ma, and Yihong Wu. Sparse PCA : optimal rates and adaptive estimation. Ann. Statist. , 41(6):3074--3110, 2013
work page 2013
-
[4]
T. Tony Cai and Anru Zhang. Rate-optimal perturbation bounds for singular subspaces with applications to high-dimensional statistics. Ann. Statist. , 46(1):60--89, 2018
work page 2018
-
[5]
Detection limits in the high-dimensional spiked rectangular model
Ahmed El Alaoui and Michael I Jordan. Detection limits in the high-dimensional spiked rectangular model. In Conference On Learning Theory , pages 410--438. PMLR, 2018
work page 2018
-
[6]
Angle-based joint and individual variation explained
Qing Feng, Meilei Jiang, Jan Hannig, and JS Marron. Angle-based joint and individual variation explained. Journal of multivariate analysis , 166:241--265, 2018
work page 2018
-
[7]
Distributed estimation of principal eigenspaces
Jianqing Fan, Dong Wang, Kaizheng Wang, and Ziwei Zhu. Distributed estimation of principal eigenspaces. Ann. Statist. , 47(6):3009--3031, 2019
work page 2019
-
[8]
Structural learning and integrative decomposition of multi-view data
Irina Gaynanova and Gen Li. Structural learning and integrative decomposition of multi-view data. Biometrics , 75(4):1121--1132, 2019
work page 2019
Show all 29 references
-
[9]
Covariate-driven factorization by thresholding for multiblock data
Xing Gao, Sungwon Lee, Gen Li, and Sungkyu Jung. Covariate-driven factorization by thresholding for multiblock data. Biometrics , 77(3):1011--1023, 2021
2021
-
[10]
On a formula for the product-moment coefficient of any order of a normal frequency distribution in any number of variables
Leon Isserlis. On a formula for the product-moment coefficient of any order of a normal frequency distribution in any number of variables. Biometrika , 12(1/2):134--139, 1918
1918
-
[11]
Information and coding theory
Gareth A Jones and J Mary Jones. Information and coding theory . Springer Science & Business Media, 2012
2012
-
[12]
The incidental parameter problem since 1948
Tony Lancaster. The incidental parameter problem since 1948. Journal of econometrics , 95(2):391--413, 2000
1948
-
[13]
Joint and individual variation explained ( JIVE ) for integrated analysis of multiple data types
Eric F Lock, Katherine A Hoadley, James Stephen Marron, and Andrew B Nobel. Joint and individual variation explained ( JIVE ) for integrated analysis of multiple data types. The annals of applied statistics , 7(1):523, 2013
2013
-
[14]
P. W. MacDonald, E. Levina, and J. Zhu. Latent space models for multiplex networks with shared structure. Biometrika , 109(3):683--706, 2022
2022
-
[15]
Optimal estimation of shared singular subspaces across multiple noisy matrices
Zhengchi Ma and Rong Ma. Optimal estimation of shared singular subspaces across multiple noisy matrices. arXiv preprint arXiv:2411.17054 , 2024
2024 arXiv
-
[16]
Consistent estimates based on partially consistent observations
Jerzy Neyman and Elizabeth L Scott. Consistent estimates based on partially consistent observations. Econometrica: journal of the Econometric Society , pages 1--32, 1948
1948
-
[17]
Fourth-order properties of normally distributed random matrices
Heinz Neudecker and Tom Wansbeek. Fourth-order properties of normally distributed random matrices. Linear Algebra and its Applications , 97:13--21, 1987
1987
-
[18]
Data integration via analysis of subspaces ( DIVAS )
Jack Prothero, Meilei Jiang, Jan Hannig, Quoc Tran-Dinh, Andrew Ackerman, and JS Marron. Data integration via analysis of subspaces ( DIVAS ). TEST , pages 1--42, 2024
2024
-
[19]
RaJIVE : Robust angle based JIVE for integrating noisy multi-source data
Erica Ponzi, Magne Thoresen, and Abhik Ghosh. RaJIVE : Robust angle based JIVE for integrating noisy multi-source data. arXiv preprint arXiv:2101.09110 , 2021
2021 arXiv
-
[20]
Triple component matrix factorization: Untangling global, local, and noisy components
Naichen Shi, Salar Fattahi, and Raed Al Kontar. Triple component matrix factorization: Untangling global, local, and noisy components. Journal of Machine Learning Research , 25(332):1--76, 2024
2024
-
[21]
Personalized PCA : Decoupling shared and unique features
Naichen Shi and Raed Al Kontar. Personalized PCA : Decoupling shared and unique features. Journal of machine learning research , 25:1--82, 2024
2024
-
[22]
Heterogeneous matrix factorization: When features differ by datasets
Naichen Shi, Raed Al Kontar, and Salar Fattahi. Heterogeneous matrix factorization: When features differ by datasets. arXiv preprint arXiv:2305.17744 , 2023
2023 arXiv
-
[23]
A spectral method for multi-view subspace learning using the product of projections
Renat Sergazinov, Armeen Taeb, and Irina Gaynanova. A spectral method for multi-view subspace learning using the product of projections. arXiv preprint arXiv:2410.19125 , 2024
2024 arXiv
-
[24]
High-dimensional probability: An introduction with applications in data science , volume 47
Roman Vershynin. High-dimensional probability: An introduction with applications in data science , volume 47. Cambridge university press, 2018
2018
-
[25]
Perturbation theory for pseudo-inverses
Per- ke Wedin. Perturbation theory for pseudo-inverses. BIT Numerical Mathematics , 13:217--232, 1973
1973
-
[26]
Normal approximation and confidence region of singular subspaces
Dong Xia. Normal approximation and confidence region of singular subspaces. Electronic Journal of Statistics , 15(2):3798--3851, 2021
2021
-
[27]
Assouad, F ano, and L e C am
Bin Yu. Assouad, F ano, and L e C am. In Festschrift for L ucien L e C am , pages 423--435. Springer, New York, 1997
1997
-
[28]
Group component analysis for multiblock data: Common and individual feature extraction
Guoxu Zhou, Andrzej Cichocki, Yu Zhang, and Danilo P Mandic. Group component analysis for multiblock data: Common and individual feature extraction. IEEE transactions on neural networks and learning systems , 27(11):2426--2439, 2015
2015
-
[29]
Limit results for distributed estimation of invariant subspaces in multiple networks inference and PCA
Runbing Zheng and Minh Tang. Limit results for distributed estimation of invariant subspaces in multiple networks inference and PCA . arXiv preprint arXiv:2206.04306 , 2022
2022
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.