{"id":"1dbede6b-a3d8-439b-ac88-8c32dc9df4a1","arxiv_id":"2501.15016","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A unified regularized MLE for generalized multi-response regression jointly exploits column homogeneity, low rank, and sparsity, with rates that improve as within-group column differences shrink.","lead":"This paper introduces a regression estimator that combines three structural assumptions at once: response-column homogeneity (clustering), low rank, and covariate sparsity, for generalized linear models with many responses. It proves convergence rates showing when column clustering improves estimation, and demonstrates the method on simulations and Barro Colorado Island tree species data.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Condition 1 is not merely strong: it is unsatisfiable for the paper's main GLMs, so Theorem 1's bound is vacuous as stated for Bernoulli, multinomial-logit, and Poisson responses without a bounded-parameter or local restricted-strong-convexity assumption.","rationale":"The reader's weakest assumption is exactly the one I find load-bearing: Condition 1 is stated over an unbounded set and fails for the GLM links the paper targets. The paper's construction, algorithm, and simulations provide real support for the heuristic that homogeneity helps, and the algorithmic gap to stationary points is secondary in comparison. Because Condition 1 can plausibly be repaired by a bounded-parameter or local-RSC condition, the appropriate disposition remains conditional rather than rejection. Since the reader already marked CONDITIONAL on this same issue, no verdict adjustment is needed.","tokens_in":16956,"tokens_out":6972,"duration_ms":73356,"concrete_test":"Analytical check: specialize to Bernoulli logistic regression with q = 1, z_i = 1, X absent, and take ~B1 = 0, ~B2 consisting of A = t. Compute Q(t) = [log(1+exp(t)) - log(2) - t/2] / t^2. As t -> infinity, Q(t) -> 0. Condition 1 requires Q(t) to be bounded below by a positive constant times kappa for all t, which is false. If the authors replace Condition 1 by a parameter-ball constraint, rerun the proof of Theorem 1 under that constraint and verify that the constant kappa is independent of n; if local restricted strong convexity around ~B* is intended instead, state the radius and an initialization condition that ensures the constrained global minimizer lies in that ball.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing assumption is Condition 1, which asks for a uniform restricted strong convexity constant kappa > 0 over the entire solution space ~Brsγ. This set is unbounded: A has no norm constraint, and B can be scaled arbitrarily along exactly homogeneous directions (where Pkm(B) = 0). For the multinomial-logit link b(theta) = log(1 + sum_j exp(theta_j)) and for Bernoulli (m = 1), the directional second derivative along theta = (t,0,...,0) is exp(t)/(1+exp(t))^2, which tends to 0 as t -> infinity. Consequently, the quadratic lower bound in Condition 1 cannot hold for pairs (~B1, ~B2) separated in the A-direction by t when z contains an intercept: the excess over the tangent divided by ||Δ||_F^2 tends to 0. The same failure occurs for the Poisson link b(theta) = exp(theta) along theta -> -infinity. Thus Theorem 1's stated conditions are incompatible with the very models emphasized in Section 2 and Table 1; only the normal/linear model has a constant Hessian. The proof needs either a bounded parameter set (e.g., ||~B||_F <= R, with the rate then depending on R) or a local restricted strong convexity near the true parameter together with an initialization guarantee. Without such a modification, the main statistical rate is vacuous for the target GLMs.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a unified regularized maximum likelihood framework for multi-response generalized linear models under three structural assumptions: column homogeneity, low rank, and row sparsity. The coefficient matrix B is factorized as UV^T, homogeneity is encouraged through a normalized Laplacian fusion penalty on the columns of B, sparsity is imposed by a hard row constraint on U, and the rank is fixed through V^T V = I_r. The main theoretical result, Theorem 1, gives a consistency rate for the constrained estimator (3.1) under Conditions 1-3, with the rate depending explicitly on the within-group similarity parameter γ, the number of groups K, the rank r, the sparsity level s, and the dimensions m and p. An iterative block-coordinate majorization-minimization algorithm is proposed and shown to be monotone in Theorem 2. Numerical experiments compare the proposed Laplacian SRRR method with SRRR and gLasso in linear and multinomial logistic settings, and the method is applied to tree species data from Barro Colorado Island.","tokens_in":17277,"tokens_out":12312,"duration_ms":112259,"significance":"If the main theorem can be made sound, this is a valuable contribution: it offers a single convergence rate that interpolates between sparse reduced-rank regression (large K, large γ) and a univariate-like regression rate (K = 1, γ = 0), and it extends existing reduced-rank methodology beyond the linear model to a broad GLM family. The rate's explicit dependence on γ and K is a genuinely falsifiable comparative prediction, and the numerical results in Tables 2-5 are consistent with the qualitative claim that stronger homogeneity improves estimation. The paper does not fit constants in the proof, and the algorithmic convergence guarantee in Theorem 2 is useful. The main weakness is that Condition 1 is stated too strongly for the very GLMs that the paper targets, so the central theoretical claim currently rests on an unsatisfiable assumption.","major_comments":[{"comment":"The restricted strong convexity condition is imposed uniformly over the entire solution set ~Brsγ, which is unbounded. A is completely unconstrained, and B may be scaled by an arbitrary c > 0 along any direction that is exactly homogeneous (then Pkm(B) = 0, rank(B) = 1, and ∥B∥_{2,0} ≤ s remain true for every c), so pairs in ~Brsγ can be arbitrarily far apart. For the Bernoulli log-partition b(θ) = log(1 + e^θ), b''(θ) = e^θ/(1 + e^θ)^2 → 0 as θ → ∞; for the Poisson log-partition b(θ) = e^θ, b''(θ) → 0 as θ → −∞; and for the multinomial logistic log-partition b(θ) = log(1 + ∑_j e^{θ_j}), the directional second derivative along θ = (t, 0, ..., 0) tends to 0 as t → ∞. Since Z contains an intercept and X takes values on R^p, these unbounded directions belong to ~Brsγ. Hence no fixed κ > 0 can satisfy the displayed inequality, and Theorem 1 is vacuous as stated for the Bernoulli and multinomial models emphasized in Section 2 and for the Poisson row of Table 1. The proof should either work with a bounded parameter set (e.g., ∥~B∥_F ≤ R, with the final rate then depending on R) or impose a local restricted strong convexity in a neighborhood of ~B* together with an initialization or restricted-step guarantee that keeps the iterates in that neighborhood.","section":"Section 3.1, Condition 1"},{"comment":"Condition 3 assumes that E = Y − E(Y) has sub-Gaussian tails, with P(|a^T e_i| > δ) ≤ 2 exp(−δ²/(τ∥a∥²)). This holds for bounded responses such as Bernoulli and multinomial indicators, but not for the Poisson responses listed in Table 1: a Poisson variable has P(|Y − λ| > t) decaying like exp(−ct log t), not exp(−ct²). Since the paper's GLM framework explicitly includes Poisson in Table 1, the theorem's scope is narrower than claimed; it covers the linear model and bounded-response GLMs but not Poisson. Please either remove the Poisson claim from the stated scope or replace Condition 3 with a tail or moment condition that Poisson satisfies, and adjust the theorem accordingly.","section":"Section 3.1, Condition 3 and Table 1"}],"minor_comments":[{"comment":"There are several typos and grammatical errors, including 'speficially' in Section 2.1, 'procruste' in the text after equation (2.9), 'seqe' in Section 4.2, and 'Whose the j-th row' in Section 2.2.2. The paper should be carefully proofread.","section":"Section 2.1 and Section 4.2"},{"comment":"In the (b) panels of Tables 2-5, the rows for gLasso, SRRR, and Laplacian are not labeled, so the reader must reconstruct the methods from the order in panel (a). Please repeat the method names in every panel.","section":"Tables 2-5"},{"comment":"The notation b(~B) in Condition 1 is not defined. In Section 2, b(θ_i) is the per-observation log-partition, whereas Condition 1 evaluates b at the whole parameter matrix and includes a factor nκ on the right-hand side. Please clarify whether b(~B) denotes the summed log-partition ∑_{i=1}^n b(θ_i) and define κ accordingly.","section":"Section 3.1, Condition 1"},{"comment":"The statement that for large γ and K the rate reduces to O(√(s log p) + √(r(s+m) log m)), 'the same as sparse reduced-rank regression excluding a logarithmic factor', is too strong. In Theorem 1(b), when K is close to m the bound still contains √(m log m) from √(s log p + m log K) and the third entropy term, so the comparison with the SRRR rate should be qualified.","section":"Section 3, discussion after Theorem 1"},{"comment":"For the Laplacian SRRR method, the rank and sparsity hyperparameters are taken from the cross-validation results of SRRR rather than being re-tuned jointly with K and λ. Please justify that this does not bias the comparison, since the proposed method is otherwise tuned on a larger grid.","section":"Section 4.2"}],"recommendation":"major_revision","confidential_remarks":"To the editor: The paper is on-topic for a methodology journal, but I recommend that the revision be returned with the explicit request that Condition 1 be repaired and that the supplementary proof be included in the review package. Please ask the authors to state how a boundedness parameter, if introduced, enters the rate, and to either remove the Poisson claim or support it with an appropriate tail condition."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a real contribution, not a repackaging. Fusing column-homogeneity pursuit into sparse reduced-rank regression for GLMs, with a rate that explicitly depends on within-group variance γ and group count K, is new relative to She et al. (row homogeneity) and Price & Sherwood (column clustering without rank/sparsity). The simulation study is careful—linear and multinomial, low/high-dim, varying within-group scatter—and the qualitative agreement with the rate intuition (smaller γ, smaller K help) is convincing. The BCI tree-species analysis is a nice applied illustration, even if the ecological story is a bit speculative.\n\nThe soft spots are real. The main one is Condition 1. Restricted strong convexity is asserted uniformly on the entire set ~Brsγ, which is unbounded: A is unconstrained and exactly homogeneous directions in B are scale-invariant. For Bernoulli, multinomial-logit, and Poisson links, the Hessian of the cumulant function decays to zero at extreme natural parameters, so no fixed κ > 0 can satisfy the quadratic lower bound. Theorem 1's stated conditions are therefore vacuous for the GLMs the paper emphasizes; only the linear model escapes because its Hessian is constant. That is a load-bearing problem, not a technicality. The fix is standard—impose a bounded parameter set or a local restricted strong convexity near the truth plus an initialization guarantee—but the rate will then need to track the extra radius or the initialization quality.\n\nA secondary gap: the consistency bound is for the constrained global minimizer in (3.1), while Algorithm 1 is only proved to converge to a critical point. No argument shows the CV-tuned output is near the global solution. This is common in nonconvex problems, but it should be stated explicitly and ideally given an initialization condition. Also, the proofs are deferred to a supplementary that is not present in the arXiv v1, so an editor should insist on seeing them.\n\nThe citation coverage looks fair; the self-citations are background and the positioning against She et al. is accurate.\n\nBottom line: the method is worth refereeing. The empirical evidence and the unification are solid; the theory needs a serious revision of Condition 1 and an honest discussion of the global-minimizer gap. A competent referee can fix this rather than kill it.","headline":"A genuinely useful unification of column homogeneity, low rank, and sparsity in GLMs, with simulations that back it up; the main rate theorem's stated conditions are unsound for the very GLMs it targets, so the theory needs a boundedness or local-RSC fix before publication.","tokens_in":17782,"tokens_out":3142,"would_cite":true,"duration_ms":30781,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62J12","62J07","62H30","62H25"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proves that when columns of a multi-response regression coefficient matrix form tight groups, the estimation error bound shrinks from the sparse reduced-rank rate toward the much faster univariate-response rate, with the…","keywords":["generalized linear model","reduced-rank regression","homogeneity pursuit","column clustering","high-dimensional regression","Laplacian penalty","variable selection","multinomial logistic regression"],"falsifier":"Take a Bernoulli logistic regression with a two-column coefficient matrix in the feasible set, and compute the smallest eigenvalue of the Hessian of the negative log-likelihood along a sequence of matrices whose norms grow to infinity while staying in the feasible set. If that eigenvalue tends to zero, Condition 1 cannot hold with a fixed $\\kappa>0$, and the claimed $O_P(\\xi_{m,K}(\\gamma)/\\sqrt n)$ rate for that setting lacks a valid premise.","tokens_in":16765,"feed_emoji":"📊","tokens_out":6087,"duration_ms":53530,"temperature":0.7,"pith_summary":"This paper tries to establish a unified theoretical picture of three structural assumptions used in high-dimensional multi-response regression: the coefficient matrix's columns cluster into a few similar groups (homogeneity), the matrix has low rank, and only a few covariates are relevant (sparsity). It proves an estimation-error bound for a regularized maximum likelihood estimator that enforces all three structures simultaneously, and the bound quantifies when column homogeneity helps. The central message is that homogeneity acts like an extra dimension-reduction device: when true columns are tightly grouped, the error rate approaches the much faster univariate-response regression rate, and when they are not, the estimator gracefully degrades to the sparse reduced-rank regression rate.","feed_headline":"Grouped responses shrink regression error to near-univariate rates","feed_subtitle":"Unified theory ties error rate to group similarity and group count, showing exactly when homogeneity helps.","key_machinery":"The mechanism is a Laplacian fusion penalty $P_{km}(B)=\\min_{\\mathcal G,\\mu}\\sum_{k=1}^K\\sum_{j\\in\\mathcal G_k}\\|\\beta_j-\\mu_k\\|_2^2$ that encourages the columns of the coefficient matrix to cluster, combined with the factorization $B=UV^\\top$ with $V^\\top V=I_r$ that enforces low rank and converts row sparsity on $B$ into row sparsity on $U$. Proposition 2 supplies the key equivalence $\\operatorname{trace}\\{B L B^\\top\\}=2\\sum_{k=1}^K\\sum_{j\\in\\mathcal G_k}\\|\\beta_j-\\mu_k\\|_2^2$, which turns the penalty into a K-means update on the rows of $V$. The algorithm iterates proximal gradient steps with hard thresholding on $U$, orthogonal Procrustes updates on $V$, and K-means updates on the group memberships, and Theorem 2 shows each blockwise update reduces the surrogate objective.","core_discovery":"The paper's central claim is Theorem 1: under restricted strong convexity of the link function, a bounded sparse-operator norm of the design, and sub-Gaussian noise, the constrained estimator defined by problem (3.1) satisfies $\\|\\hat A - A^*\\|_F + \\|\\hat B - B^*\\|_F = O_P(\\rho \\xi_{m,K}(\\gamma)/(\\kappa \\sqrt n))$, where $\\xi_{m,K}(\\gamma)$ combines three terms: $\\sqrt{r\\gamma(s+m)}$, $\\sqrt{s\\log p + m\\log K}$, and $\\sqrt{[s+\\min(K^2,m)]\\min(r,K)\\log m}$. The rate explicitly depends on the within-group similarity $\\gamma$ and the number of groups $K$, so the paper can state precisely when homogeneity improves accuracy: a small $\\gamma$ (relative to the noise floor $\\zeta_n^2$) and a small $K$ shrink the error, while a large $\\gamma$ or indistinct groups makes the bound collapse to the known sparse reduced-rank regression rate.","pith_inferences":["A practical rule implied by the theorem: pursue homogeneity only when the estimated within-group column spread is below the noise floor $\\zeta_n^2$; otherwise the SRRR rate already applies, and the extra penalty buys little.","Because the fusion penalty is equivalent to K-means on the rows of $V$, the same framework could be extended to two-way homogeneity, clustering rows and columns simultaneously, and to other losses such as quantile regression.","The theorem suggests that a data-driven estimator of $\\gamma$ and $K$ from the fitted $V$ could replace cross-validation on a two-dimensional grid, making the method more scalable in applications with many responses.","The BCI application indicates the clustering output has ecological interpretability, which would be strengthened by quantifying uncertainty in the estimated groups, a topic the paper leaves open."],"forward_implications":["When all columns are identical ($\\gamma=0$, $K=1$), the bound becomes $O_P(\\sqrt{s\\log(p\\vee m)}/\\sqrt n)$, matching the high-dimensional univariate-response regression rate.","When the group structure is absent or weak, the error bound reduces to the sparse reduced-rank regression rate, so the method does not lose accuracy compared with SRRR.","Both the within-group similarity $\\gamma$ and the group number $K$ enter the rate explicitly, giving a quantitative way to decide when column homogeneity will help.","The blockwise algorithm decreases the objective at every update and converges to a critical point, so the estimator is computationally reproducible.","Tuning $K$, $\\lambda$, rank, and sparsity by cross-validation lets practitioners compare homogeneity, low-rank, and sparsity effects on a common benchmark."],"supporting_citations":[{"why":"Establishes the joint variable-and-rank selection framework and restricted-eigenvalue rates that the sparse reduced-rank baselines are compared against.","marker":"Bunea et al. 2012"},{"why":"Supplies the sparse operator-norm device and the SOFAR estimator conditions that Theorem 1's Conditions 1-2 adapt.","marker":"Uematsu et al. 2019"},{"why":"Provides the fast stagewise sparse factor regression approach whose conditions and rate structure Theorem 1 extends to grouped columns.","marker":"Chen et al. 2022"},{"why":"Underpins the iterative hard-thresholding update for U and the descent argument used in Theorem 2.","marker":"Jain et al. 2014"},{"why":"Gives the selective factor extraction rate that the paper matches when the homogeneity term disappears.","marker":"She 2017"},{"why":"Models row homogeneity in reduced-rank regression, the closest prior unified treatment that the paper contrasts with column homogeneity.","marker":"She et al. 2022"},{"why":"Uses column homogeneity to improve multi-response prediction, the practical motivation for the fusion penalty.","marker":"Price and Sherwood 2018"},{"why":"Provides the semiparametric multinomial logistic model that the empirical Barro Colorado Island analysis builds on.","marker":"Hessellund et al. 2022b"}],"fun_headline_variants":["Group homogeneity shrinks regression error to univariate-like rates","Unified GLM shows when grouping beats sparse low-rank","Error bound scales with group similarity and group count","Homogeneity helps exactly when groups are tight and few"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole rate proof rests on the assumption that the log-likelihood stays strongly convex with a fixed positive curvature constant across the entire (unbounded) feasible set of coefficient matrices, but for logistic and Poisson models the curvature shrinks toward zero as the natural parameters grow without bound, so this uniformity can fail for exactly the models the paper targets.","fun_headline_variants_meta":{"raw":{"variants":["Group homogeneity shrinks regression error to univariate-like rates","Unified GLM shows when grouping beats sparse low-rank","Error bound scales with group similarity and group count","Homogeneity helps exactly when groups are tight and few"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001625,"raw_usage":{"total_tokens":6440,"prompt_tokens":895,"completion_tokens":5545,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":511,"completion_tokens_details":{"reasoning_tokens":5481}},"tokens_in":511,"tokens_out":5545,"duration_ms":55053,"temperature":1.0,"reasoning_tokens":5481,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T14:43:14.868975+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a Bernoulli logistic regression with a two-column coefficient matrix in the feasible set, and compute the smallest eigenvalue of the Hessian of the negative log-likelihood along a sequence of matrices whose norms grow to infinity while staying in the feasible set. If that eigenvalue tends to zero, Condition 1 cannot hold with a fixed $\\kappa>0$, and the claimed $O_P(\\xi_{m,K}(\\gamma)/\\sqrt n)$ rate for that setting lacks a valid premise.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the joint variable-and-rank selection framework and restricted-eigenvalue rates that the sparse reduced-rank baselines are compared against."},{"cited_title":"Sofar: Large-scale association network learning","cited_arxiv_id":null,"evidence_quote":"Supplies the sparse operator-norm device and the SOFAR estimator conditions that Theorem 1's Conditions 1-2 adapt."},{"cited_title":"On iterative hard thresholding methods for high-dimensional m-estimation","cited_arxiv_id":null,"evidence_quote":"Underpins the iterative hard-thresholding update for U and the descent argument used in Theorem 2."},{"cited_title":"Selective factor extraction in high dimensions","cited_arxiv_id":null,"evidence_quote":"Gives the selective factor extraction rate that the paper matches when the homogeneity term disappears."},{"cited_title":"A cluster elastic net for multivariate regression","cited_arxiv_id":null,"evidence_quote":"Uses column homogeneity to improve multi-response prediction, the practical motivation for the fusion penalty."}],"review_version":1}