{"id":"1d757a23-75da-4490-8804-13f71bb01bd2","arxiv_id":"2412.07987","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"A new U-statistic based test, combined with a sparse singular value decomposition estimator, is proposed for testing whether a high-dimensional matrix mean has rank at most K.","lead":"This statistics paper proposes a new hypothesis test for the rank of the mean matrix of high-dimensional matrix-valued data, plus a sparse SVD estimator to make the test work in practice. A generalist might care because such rank tests could be used to detect objects in surveillance video by treating each frame as a noisy matrix.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 3 never restricts K; for q=p=n, Sigma=I, K=n/2, the oracle statistic V_n has asymptotic variance 0.75 times the claimed denominator, so the N(0,1) calibration fails under the paper's own assumptions.","rationale":"The paper's central claim is Theorem 3, which states that the oracle statistic V_n, scaled by sqrt(2 tr(Sigma^2)/n^2), is asymptotically N(0,1) under H0 whenever Assumptions A1 and A4 hold, with Theorem 5 extending this to the plug-in statistic used in practice. This scaling is only correct if the subtracted U-statistics sum_{k=1}^K T_k are asymptotically negligible in variance relative to U_n. That is not guaranteed by A4, and it is false when the projection P_V otimes P_U captures a non-negligible fraction of the spectral mass of Sigma. The reader's weakest_assumption identifies exactly this gap: the variance of V_n is not the variance of U_n once K > 0. My analysis makes the mechanism precise: V_n is a degenerate U-statistic with kernel x^T Q y, Q = I - P_V otimes P_U, so its variance is governed by tr(Q Sigma Q Sigma). For Sigma = I and K of order p, the correction is of constant order. A concrete sequence with q = p = n, K = n/2 satisfies A1 and A4 and yields asymptotic variance 3/4, so Theorem 3 is not merely missing a proof detail; it is false as stated. The empirical sections cannot rescue the claim because all simulations use K = 1, where K^2 ||Sigma||_2^2 is indeed negligible relative to tr(Sigma^2). The real-data applications also use sequential tests starting at small K, so they do not probe the broken regime. I therefore agree with the reader's verdict: the central asymptotic calibration is unsupported and the rejection stands. I would only note that the reader's phrase '2 tr(Sigma_2)/n^2' is likely an extraction artifact for '2 tr(Sigma^2)/n^2'; the substantive concern is independent of that typographic issue and becomes stronger when the correct denominator is used.","tokens_in":13567,"tokens_out":19097,"duration_ms":188768,"concrete_test":"Run the oracle-statistic counterexample: set q = p = n = 200, Sigma = I, K = 100, Pi_0 = sum_{k=1}^{100} e_k e_k^T, U_0 = V_0 = [e_1,...,e_{100}], and X_i = Pi_0 + Z_i with i.i.d. N(0,1) entries. Compute V_n and G_n = V_n / sqrt(2 tr(Sigma^2)/n^2) using the known U_0,V_0 over 10,000 Monte Carlo replications, and estimate Var(G_n). The predicted value is (p^2 - K^2)/p^2 = 3/4, not 1; the empirical quantiles will also be compressed by the factor sqrt(3)/2. If instead one derives tr(Q Sigma Q Sigma) analytically, the same ratio follows immediately, confirming that a corrected theorem must either restrict K (e.g., K^2 ||Sigma||_2^2 = o(tr(Sigma^2))) or replace the denominator by 2 tr(Q Sigma Q Sigma)/n^2.","verdict_should_be":"REJECT","load_bearing_attack":"Under H0 the oracle statistic is V_n = (1/(n(n-1))) sum_{i!=j} x_i^T (I - P_V otimes P_U) x_j, where P_U = U_0 U_0^T, P_V = V_0 V_0^T, and x_i = vec(X_i). Writing x_i = mu + u_i, the mean term satisfies Q mu = 0 with Q = I - P_V otimes P_U, so V_n is a degenerate U-statistic whose variance is 2 tr(Q Sigma Q Sigma)/n^2 + o(n^{-2}), not the 2 tr(Sigma^2)/n^2 used in Theorem 3. No condition on K appears in Theorem 3 or Theorem 5. Take q = p = n, Sigma = I_{n^2}, K = floor(n/2), and Pi_0 of rank K with U_0,V_0 the first K standard basis vectors. Then P_V otimes P_U has rank K^2 ~ n^2/4, so tr(Q^2) = 3n^2/4 + o(n^2), and the scaled statistic G_n converges in law to N(0,3/4), not N(0,1). Assumptions A1 and A4 hold in this sequence, so Theorem 3 is false as stated. The plug-in Theorem 5 inherits the same missing restriction, and the simulation section only exercises K = 1, where the projection term is negligible; it does not test the regime in which the theorem fails.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a test for H0: rank(Π0) ≤ K for the mean of high-dimensional matrix-valued data with separable covariance. The authors first argue that the classical Cragg-Donald minimum-discrepancy test fails when matrix dimensions grow with the sample size, then introduce an oracle U-statistic that subtracts the first K estimated singular-value contributions from the total squared Frobenius norm, claim asymptotic normality under the null and local alternatives, and provide a sparse SVD estimator to make the statistic plug-in. Simulations and two surveillance-video case studies are used to support the method.","tokens_in":14071,"tokens_out":20653,"duration_ms":203289,"significance":"The testing problem is timely, and the oracle statistic is a natural U-statistic extension of high-dimensional mean tests to matrix rank testing. The sparse SVD estimator with concave penalties is a useful contribution, and the video surveillance applications are interesting. If the central distributional claims were valid, the paper would be a meaningful step for matrix-valued inference. However, the main theorem is not valid in the stated generality, and the plug-in theorem inherits the same gap; the paper's own limitations and the missing proofs must be addressed before the claims can be accepted.","major_comments":[{"comment":"The variance of the oracle statistic is miscalibrated when K is not negligible. Writing Q = I - V0V0ᵀ ⊗ U0U0ᵀ and x_i = vec(X_i), the statistic V_n equals (1/(n(n-1))) Σ_{i≠j} x_iᵀ Q x_j. Under H0, Q vec(Π0) = 0, so this is a degenerate U-statistic with variance 2 tr(QΣQΣ)/n² + o(n⁻²), not 2 tr(Σ²)/n² as used in the denominator of (3.5). The difference tr(Σ²) − tr(QΣQΣ) = 2 tr((V0V0ᵀ⊗U0U0ᵀ)Σ²) − tr((V0V0ᵀ⊗U0U0ᵀ)Σ(V0V0ᵀ⊗U0U0ᵀ)Σ) can be of the same order as tr(Σ²) when K grows. For example, take q = p = n, Σ = I_{n²}, and Π0 of rank K = floor(n/2) with U0 = V0 equal to the first K standard basis vectors. Assumptions A1 and A4 hold, but tr(QΣQΣ) = (3/4) tr(Σ²) + o(tr(Σ²)), so G_n converges to N(0, 3/4), not N(0, 1). No condition on K appears in Theorem 3 or in A4; a condition such as K² ||Σ||₂² = o(tr(Σ²)), or an explicit variance calculation, is required.","section":"3.2, Theorem 3"},{"comment":"The plug-in statistic inherits the missing variance correction from Theorem 3. Assumptions A9 and A10 impose no restriction linking K to tr(Σ²); A10 only requires the projection error to be o_P(√tr(Σ²)/n), which does not make the Q-factor in the variance vanish. Consequently the local-power statement (3.9) is also unproved in the stated generality. The simulation section (Section 4, Tables 1–3) only exercises K = 1: the null models in Model (a) with c = 0 and Model (b) with c = 0 have rank-1 mean matrices, so the subtracted projection has rank K² = 1 and the variance correction is negligible. The simulations therefore do not test the regime in which Theorem 3 fails.","section":"3.4, Theorem 5"},{"comment":"The null hypothesis H0: rank(Π0) ≤ K includes cases where the true rank R is strictly less than K. In those cases the singular vectors μ_k, ν_k for k > R are not identified, so the oracle statistic (3.4) is not a well-defined functional of the data-generating process; for general Σ, its null distribution can depend on the arbitrary completion of U0 and V0 through terms such as tr((V0V0ᵀ⊗U0U0ᵀ)Σ²). Theorem 4 explicitly assumes 0 < K ≤ R, so the estimation theory does not cover the R < K part of the null. The paper should either restrict the null to R = K or prove that the statistical behavior is invariant to the completion and that A10 is attainable in that case.","section":"3.2 and 3.4, null with R < K"},{"comment":"The manuscript states in Section 1 and elsewhere that proofs are presented in the appendix, but no appendix is included in the posted text. Since the core results are asymptotic theorems for degenerate U-statistics and a new sparse SVD estimator, the absence of the proofs makes the central claims unverifiable as submitted. This is a load-bearing omission for a theory-focused paper.","section":"Appendix / proofs"}],"minor_comments":[{"comment":"The display for Theorem 2 appears to be missing a square root in the denominator; it should be (T_R − (q−R)(p−R)) / √(2(q−R)(p−R)) → N(0, 1), not the expression with 2(q−R)(p−R) in the denominator as printed.","section":"3.1, Theorem 2"},{"comment":"The condition for λ_v uses max_{v∈Nu0}, which should be max_{v∈Nv0}; the two displayed conditions for λ_u and λ_v otherwise appear identical, suggesting a copy-paste error.","section":"3.3, Assumption A8"},{"comment":"There are typos including 'subgaussion' for 'subgaussian', 'Euclidan' for 'Euclidean', and 'can ’t' in Remark 1; the label 'REMARK 1' in Table 3 should be replaced by a proper model description.","section":"2.2 and 4"},{"comment":"The displayed estimator \\S\\S^2_{n1} is garbled: the first term as printed, tr(2XiX T j )2, should presumably be tr(X_i X_jᵀ)² with the correct combinatorial factors; please check the formula against Li and Chen (2012).","section":"3.4, estimator of tr(Σ²)"}],"recommendation":"major_revision","confidential_remarks":"The counterexample in the stress-test note is correct, and the variance computation above confirms that Theorem 3 is false as stated. The error appears fixable by imposing a small-K condition (e.g., K² ||Σ||₂² = o(tr(Σ²))) and by restricting the null derivation to cases where the singular vectors are identified or showing invariance; the paper should also supply the missing appendix. I do not see a basis for questioning the authors' integrity; the self-citations are to directly relevant prior work on power-enhanced mean tests and sparse SVD."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper takes a real problem — rank testing for the mean of matrix-valued data when dimensions grow with n — and proposes a natural solution: a U-statistic designed to separate the top-K singular value contribution from the rest, plugged with a sparse SVD estimator. The motivation is solid, and the paper correctly shows that the classic Cragg–Donald minimum discrepancy test breaks down as dimensions increase. The oracle statistic is a straightforward but sensible adaptation of Chen et al. (2010), and the sparse SVD theory is a credible extension of Lee et al. (2010). The simulations and the two video case studies are also a plus, demonstrating that the method works in practice for K=1 and that the minimum discrepancy test does not control size.\n\nThe soft spot is load-bearing. Theorem 3, as stated, claims the oracle statistic is asymptotically N(0,1) under H0 for all K, with variance 2 tr(Sigma^2)/n^2. That is false. The variance calculation ignores the projection term Q = I - P_V ⊗ P_U. The stress-test example is correct: with q=p=n, Sigma=I, and K around n/2, the asymptotic variance is 3/4 of the claimed denominator, so the calibration fails under the paper's own assumptions. The theorem needs an explicit condition on K — something like K^2 = o(tr(Sigma^2)) or K fixed — and a rigorous variance computation. This is not a minor typo; it propagates to Theorem 5, which inherits the same gap. The simulations only exercise K=1, where the projection term is negligible, so they never expose the flaw.\n\nThat said, the paper is not sloppy in its overall approach. The sparse SVD theory, the plug-in analysis, and the empirical work are all coherent. The main missing piece is a correct statement and proof of the oracle theorem under appropriate restrictions on K. The conclusion overstates novelty slightly — claiming to be the first work on testing the structure of matrix-valued data ignores related rank tests in econometrics — but the high-dimensional matrix-mean setting is genuinely new.\n\nMy recommendation: this paper deserves a serious referee, not a desk reject. The idea is worth developing, and the flaw appears fixable. But in its current form, I would not accept it; the central asymptotic result is incorrect as stated. If the authors can supply a corrected theorem with a proper condition on K and update the plug-in theory accordingly, the paper could make a useful contribution.","headline":"A sensible rank test for high-dimensional matrix means that is undermined by an over-broad oracle theorem, but the underlying idea is worth a serious referee.","tokens_in":14481,"tokens_out":2200,"would_cite":false,"duration_ms":24109,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62H15","62F05","62G20","62H25"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proposes a U-statistic difference test for the rank of a high-dimensional matrix mean, proves asymptotic normality of its oracle and plug-in versions, and demonstrates the test on surveillance-video object detection.","keywords":["matrix-valued data","rank hypothesis testing","high-dimensional statistics","U-statistics","sparse singular value decomposition","minimum discrepancy test","surveillance video analysis"],"falsifier":"Take $q=p=K$, $\\Pi_0=0$, and $\\Sigma=I_{qp}$. Then with any full-rank oracle pair $U_0,V_0$ the statistic $V_n = U_n - \\sum_{k=1}^K T_k$ is identically zero for every sample, while the claimed scaling $2\\mathrm{tr}(\\Sigma^2)/n^2 = 2p^2/n^2$ is positive, so the central limit theorem cannot hold. A concrete check is to simulate $V_n$ in this setting, estimate its variance, and watch it vanish as $K$ approaches $\\min(q,p)$.","tokens_in":13374,"feed_emoji":"🎥","tokens_out":12846,"duration_ms":118639,"temperature":0.7,"pith_summary":"This paper tries to establish a valid test for whether the mean matrix of high-dimensional matrix-valued data has rank at most $K$, in the regime where the matrix dimensions $q$ and $p$ grow with or exceed the sample size $n$. The classical minimum discrepancy test works only when $\\max(q,p)=o(n^{1/4})$ and loses size control in high dimensions, so the authors build a different statistic: compare an estimate of the total squared signal in the mean matrix with an estimate of the sum of the $K$ largest squared singular values. Under the null these two quantities agree in expectation, and the paper proves that their difference, after scaling by an estimate of the noise covariance trace, is asymptotically standard normal, with known local power. The result matters because rank is a structural property of images and other array data, and this is the first test designed for that question when dimensions are large. The paper also supplies the plug-in machinery, a sparse SVD estimator for singular vectors and a ratio-consistent estimator of $\\mathrm{tr}(\\Sigma^2)$, and validates the method on surveillance videos.","feed_headline":"A new test checks matrix rank when dimensions dwarf the sample","feed_subtitle":"It compares total signal with top-K singular values, catching objects entering a surveillance scene.","key_machinery":"The load-bearing object is the difference of two U-statistics: $U_n = \\mathrm{tr}\\bigl(\\tfrac{1}{n(n-1)}\\sum_{i\\ne j}X_iX_j^\\top\\bigr)$ estimates the total squared Frobenius norm of the mean, and $T_k = \\tfrac{1}{n(n-1)}\\sum_{i\\ne j}\\mu_k^\\top X_i V_0V_0^\\top X_j^\\top \\mu_k$ estimates the $k$-th squared singular value using the oracle singular vectors. The plug-in statistic $\\hat G_n = (U_n - \\sum_{k=1}^K \\hat T_k)/\\sqrt{2\\hat S^2_{n1}/n^2}$ is built from a sparse SVD estimator that solves the penalized Frobenius problem (3.7) with a group concave penalty, and from a ratio-consistent estimator $\\hat S^2_{n1}$ of $\\mathrm{tr}(\\Sigma^2)$. The theoretical work is to show that the sparse SVD error is small enough (A10) that the plug-in statistic inherits the oracle normal limit, while the minimum discrepancy statistic fails when dimensions grow.","core_discovery":"The central claim is that the oracle statistic $V_n = U_n - \\sum_{k=1}^K T_k$, where $U_n$ estimates $\\mathrm{tr}(\\Pi_0\\Pi_0^\\top)$ and $T_k$ estimates the $k$-th squared singular value $\\sigma_k^2$, is asymptotically normal with variance $2\\mathrm{tr}(\\Sigma^2)/n^2$ under $H_0: \\mathrm{rank}(\\Pi_0)\\le K$, and has nontrivial power when $n\\sum_{i>K}\\sigma_i^2/\\sqrt{2\\mathrm{tr}(\\Sigma^2)}\\to\\delta$. The sample version $\\hat G_n$ replaces unknown singular vectors by a sparse SVD estimator and replaces $\\mathrm{tr}(\\Sigma^2)$ by a ratio-consistent estimator; under assumptions A1, A4, A9 and A10 the sample version has the same null limit and the same local power. Thus the paper asserts that rank of a high-dimensional matrix mean is testable with standard normal p-values, without estimating the inverse covariance matrix.","pith_inferences":["As an extension the paper does not state, the normal calibration requires $K$ to be small relative to $\\min(q,p)$; at $K=\\min(q,p)$ with a full-rank projection, $V_n$ can be identically zero, so a corrected theorem would either restrict $K$ or compute the variance of the subtracted part explicitly.","The tuning-parameter selection by sample splitting is a practical bottleneck; the theory assumes a well-behaved penalty, so a fully data-driven choice (for example an analog of BIC under non-identity noise) would be the natural next development.","The same difference-of-sums logic could be carried to tensor-valued data by replacing squared singular values with Tucker-core norms, although the paper mentions tensors only as motivation.","The video experiments suggest estimated rank could serve as a continuous event signal for counting or sizing objects; the paper itself reports only that rank rises when objects enter and falls when they leave."],"forward_implications":["A practitioner can test $H_0: \\mathrm{rank}(\\Pi_0)\\le K$ with standard normal critical values even when $q$ and $p$ are much larger than $n$, a regime where the minimum discrepancy statistic's $\\chi^2$ approximation breaks down.","Sequential testing over $K=0,1,\\ldots$ yields an estimate of the true rank, and the paper shows this estimate tracks the presence of pedestrians and cars in two surveillance-video datasets.","Under the sparsity assumptions, the sparse SVD estimator recovers the row and column support of the singular vectors at rate $(\\sqrt{k}+\\sqrt{l})/\\sqrt{n}$, so the plug-in test is valid whenever $n(k+l)=o(\\mathrm{tr}(\\Sigma^2))$.","Setting $K=0$ reduces the new statistic to the L2-type high-dimensional mean test from which the U-statistic construction is drawn, making the matrix rank test a generalization of an existing vector test.","The local power formula shows the test can detect alternatives with signal $\\delta = n\\sum_{i>K}\\sigma_i^2/\\sqrt{2\\mathrm{tr}(\\Sigma^2)}$, and that power drops predictably as the noise covariance trace grows."],"supporting_citations":[{"why":"Introduces the minimum discrepancy test whose fixed-dimension chi-squared theory is the baseline the paper shows breaks down in high dimensions.","marker":"Cragg & Donald (1997)"},{"why":"Supplies the high-dimensional mean-testing framework and the moment and sub-Gaussian assumptions collected in A1 that the new statistic's theory relies on.","marker":"Bai & Saranadasa (1996)"},{"why":"Provides the L2-type U-statistic approach for high-dimensional covariance tests that the new statistic reduces to when K=0.","marker":"Chen et al. (2010)"},{"why":"Gives the ratio-consistent estimator of tr(Sigma^2) used in the plug-in statistic's denominator.","marker":"Li & Chen (2012)"},{"why":"Proposes the penalized-regression sparse SVD algorithm that the paper adapts to non-identity noise.","marker":"Lee et al. (2010)"},{"why":"Establishes the rank-K Frobenius approximation whose penalized version is the paper's sparse SVD objective (3.7).","marker":"Eckart & Young (1936)"},{"why":"Provides rate-optimal sparse SVD theory for identity noise that the paper contrasts with its own general-covariance result.","marker":"Yang et al. (2016)"}],"fun_headline_variants":["Matrix rank test survives high dimensions without inverse covariance","Rank test for matrix means when dimensions dwarf sample","High-dim matrix rank test: no inverse covariance needed","New test checks matrix rank in high-dimensional settings","Sparse SVD powers rank test for high-dim matrix means"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that subtracting the estimated top-$K$ squared singular-value contributions does not change the variability of the test statistic, so that the normal approximation with variance $2\\mathrm{tr}(\\Sigma^2)/n^2$ is valid; this assumption is stated nowhere, and it fails when $K$ is close to the matrix dimensions.","fun_headline_variants_meta":{"raw":{"variants":["Matrix rank test survives high dimensions without inverse covariance","Rank test for matrix means when dimensions dwarf sample","High-dim matrix rank test: no inverse covariance needed","New test checks matrix rank in high-dimensional settings","Sparse SVD powers rank test for high-dim matrix means"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000258,"raw_usage":{"total_tokens":1567,"prompt_tokens":918,"completion_tokens":649,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":534,"completion_tokens_details":{"reasoning_tokens":584}},"tokens_in":534,"tokens_out":649,"duration_ms":6493,"temperature":1.0,"reasoning_tokens":584,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T18:22:24.820227+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take $q=p=K$, $\\Pi_0=0$, and $\\Sigma=I_{qp}$. Then with any full-rank oracle pair $U_0,V_0$ the statistic $V_n = U_n - \\sum_{k=1}^K T_k$ is identically zero for every sample, while the claimed scaling $2\\mathrm{tr}(\\Sigma^2)/n^2 = 2p^2/n^2$ is positive, so the central limit theorem cannot hold. A concrete check is to simulate $V_n$ in this setting, estimate its variance, and watch it vanish as $K$ approaches $\\min(q,p)$.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the minimum discrepancy test whose fixed-dimension chi-squared theory is the baseline the paper shows breaks down in high dimensions."},{"cited_title":"& Saranadasa, H","cited_arxiv_id":null,"evidence_quote":"Supplies the high-dimensional mean-testing framework and the moment and sub-Gaussian assumptions collected in A1 that the new statistic's theory relies on."},{"cited_title":"X., Zhang, L.-X","cited_arxiv_id":null,"evidence_quote":"Provides the L2-type U-statistic approach for high-dimensional covariance tests that the new statistic reduces to when K=0."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Proposes the penalized-regression sparse SVD algorithm that the paper adapts to non-identity noise."},{"cited_title":"& Young, G","cited_arxiv_id":null,"evidence_quote":"Establishes the rank-K Frobenius approximation whose penalized version is the paper's sparse SVD objective (3.7)."},{"cited_title":"& Buja, A","cited_arxiv_id":null,"evidence_quote":"Provides rate-optimal sparse SVD theory for identity noise that the paper contrasts with its own general-covariance result."}],"review_version":1}