{"id":"14cd9ca2-1201-434f-91a0-d5f26c2468a2","arxiv_id":"2501.01079","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"For inhomogeneous random matrices, the spectral radius is bounded by the variance row/column sums up to the optimal sparsity (log n)^{-1/2}.","lead":"Random matrices with uneven entry patterns can still have predictable largest eigenvalue magnitudes: this paper proves sharp upper bounds based only on the variance profile. The results apply to sparse, structured matrices used in neural-network and ecological stability models, and pin down the optimal sparsity scale.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 1.1's absence-of-outliers conclusion is false as stated: diagonal Gaussian matrices with sigma*=1/sqrt(n) satisfy all assumptions but have rho(A) ~ sqrt(2 log n) sigma.","rationale":"The reader's weakest assumption was Definition 1.3 (long-time control). That concern does not land as stated: for any nonnegative variance matrix S with row sums bounded by sigma^2, induction gives sup_i sum_j [S^k]_{ij} <= sigma^{2k}, and similarly for S^*, so long-time control with C=1 is automatic for the row-sum parameter sigma. The condition is only a device for choosing a smaller sigma when the variance profile has extra structure, not a restrictive structural assumption. The genuinely load-bearing issue is in Theorem 1.1's final conclusion. The proof itself is careful to split the regime sigma/sigma* >> sqrt(log n), where it gets the desired limsup bound, from an intermediate regime where it only controls rho by D(sigma + sigma* sqrt(log n)). But the theorem's 'In particular' drops the first condition and asserts P(limsup rho > sigma)=0 whenever sigma* = o(1/sqrt(log n)). The diagonal Gaussian example shows this assertion is false. It satisfies all hypotheses with sigma = sigma* = 1/sqrt(n), and the spectral radius is the maximum of n Gaussians divided by sqrt(n), so rho/sigma ~ sqrt(2 log n) -> infinity almost surely. Thus the central claim, as formulated, is falsified. A repaired statement should either assume sigma/sigma* >> sqrt(log n) or normalize sigma = O(1) and restrict to profiles with sigma of order one; the technical moment bounds may well support that repaired claim, but the current theorem overclaims. This is why the verdict should move from CONDITIONAL to REJECT: the false conclusion is part of the paper's headline contribution, not a minor gap in a peripheral lemma.","tokens_in":51032,"tokens_out":26311,"duration_ms":267215,"concrete_test":"Analytically verify the diagonal counterexample: for A = diag(g_1,...,g_n)/sqrt(n) with g_i i.i.d. N(0,1), compute P(rho(A) > sigma) = P(max_i |g_i| > 1) = 1 - (1 - 2(1 - Phi(1)))^n, which tends to 1. This directly refutes the 'In particular' part of Theorem 1.1 under sigma* = o(1/sqrt(log n)). Additionally, check the proof's regime split on page 28: the Borel-Cantelli argument for limsup rho <= sigma is only written for sigma/sigma* >> sqrt(log n), and the intermediate-regime bound contains an extra factor D(sigma + sigma* sqrt(log n)), so it cannot yield the stated conclusion.","verdict_should_be":"REJECT","load_bearing_attack":"The load-bearing flaw is in Theorem 1.1, not in the long-time control condition. The proof establishes P(limsup rho(A) > sigma)=0 only in the regime sigma/sigma* >> sqrt(log n); in the complementary regime it only obtains a bound with an extra factor D(sigma + sigma* sqrt(log n)). The theorem statement, however, claims this conclusion whenever sigma* = o(1/sqrt(log n)), which is much weaker. This claim is false: take A diagonal with independent entries N(0,1/n). Then sigma* = sigma = 1/sqrt(n), so sigma* sqrt(log n) = sqrt(log n/n) -> 0 and all assumptions of Theorem 1.1 hold. But rho(A) = (1/sqrt(n)) max_i |g_i| with g_i i.i.d. N(0,1), so rho(A)/sigma = max_i |g_i| ~ sqrt(2 log n) -> infinity almost surely. Hence P(limsup rho(A) > sigma)=1, contradicting the theorem. The moment comparison to a homogeneous matrix Y_r cannot detect this: the relevant diagonal self-loop shapes are not controlled by the same comparison when sigma^2 is much smaller than p, which is exactly the regime omitted by the proof. The theorem needs an added condition such as sigma/sigma* >> sqrt(log n), or sigma bounded below by a constant; without it, the central 'outliers absent' claim is not merely unproved but wrong.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the spectral radius ρ(A) of n×n random matrices with independent, mean-zero entries and variance profile S. It proves several upper bounds: Theorem 1.1 gives a non-asymptotic expectation bound and claims almost-sure absence of outliers (ρ(A) ≤ σ) when σ* = o(1/√log n); Theorem 1.6 establishes small-deviation bounds under a long-time control condition; Theorem 1.7 gives small-deviation bounds via ρ(S); Theorem 1.10 gives a large-deviation inequality for Gaussian doubly stochastic profiles; and Theorem 1.11 extends upper bounds to heavy-tailed entries with only 2+ε moments. The proofs rely on trace moment expansions, combinatorial counting, and comparisons to homogeneous random matrices.","tokens_in":51295,"tokens_out":9712,"duration_ms":81494,"significance":"The paper's technical core — the moment estimates in Section 2 and their inhomogeneous adaptation in Section 3 — is substantial and potentially useful for non-Hermitian random band matrices and inhomogeneous neural-network models. The long-time control parameter (Definition 1.3) is a noteworthy refinement over row/column variance sums. However, the main theorem as stated is false, and several proofs of load-bearing cases are only sketched or omitted. These issues must be addressed before the paper can be considered for publication.","major_comments":[{"comment":"The 'in particular' statement of Theorem 1.1 — that P(limsup ρ(A) > σ) = 0 whenever σ* = o(1/√log n) — is false as stated. Take A diagonal with b_ii = n^{-1/2} and b_ij = 0 for i ≠ j, with g_ij i.i.d. standard real Gaussians. Then σ = σ* = n^{-1/2}, σ/σ* = 1 ≤ n, and all assumptions on the entry distribution hold. But ρ(A) = n^{-1/2} max_i |g_i|, so ρ(A)/σ = max_i |g_i| ~ √(2 log n) → ∞ almost surely; hence P(limsup ρ(A) > σ) = 1. The proof in §3.1 explicitly derives the absence-of-outliers only under the additional condition σ/σ* ≫ √(log n) ('First suppose that σ/σ* ≫ √log n ... In the regime where C1√log n ≤ σ/σ* ≤ C2√log n ... only a bound with an extra factor is obtained'). Thus Theorem 1.1 needs an extra hypothesis such as σ/σ* ≫ √(log n) (or σ bounded below by a positive constant), and the abstract's claim (1) and the introduction's 'whenever σ*√log n → 0 then ρ(A) ≤ σ(1+ε)' must be amended accordingly. Without this correction, the central claim of the paper is incorrect.","section":"Theorem 1.1 and §3.1"},{"comment":"The complex case of Theorem 2.5 (and therefore of Theorem 1.1) is left unproved: the proof in §2.2.3 ends with 'The rest of the proof is the same and omitted,' and the proof of Theorem 1.1 for β = 2 in §3.1 says 'The details are omitted.' Since Theorem 1.1 is stated for both real and complex Gaussian entries, the complex case is part of the main result and requires a proof; referencing an argument that is 'essentially identical' is not sufficient for a load-bearing case.","section":"§2.2.3 and §3.1 (complex case)"},{"comment":"The removal of the symmetry assumption (Theorem 2.3) is deferred to Appendix C, and the adaptation to inhomogeneous profiles in Theorem 1.6 for non-symmetric distributions is only sketched ('The remaining computations are analogous', §3.2). The paper states that the gluing construction in Appendix C is lengthy, but since Theorem 1.6 is a claimed improvement over Theorem 1.1 (no symmetry needed), the non-symmetric case should be proved rather than indicated.","section":"Appendix C and §3.2"}],"minor_comments":[{"comment":"The abstract states 'up to the optimal sparsity σ_* ≫ (\\log n)^{-1/2}', but the condition in Theorem 1.1 is σ* = o((\\log n)^{-1/2}); the inequality direction appears to be a typo and should be corrected.","section":"Abstract"},{"comment":"The title contains a typographical error: 'CONCENTRA TION' should read 'CONCENTRATION'.","section":"Title"},{"comment":"The notation E[g_i]^{n_i(s)} in (3.3) is not precise; it should be clarified that g_i denotes an i.i.d. copy of the entry distribution and n_i(s) is the multiplicity of the corresponding directed edge in the shape s.","section":"Equation (3.3)"},{"comment":"The condition '√σ∗(log n)^{3/4} ≤ t^{1.5}/500' mixes the parameters t and n in a way that is hard to interpret; stating the equivalent regime for t in terms of σ* and n would improve readability.","section":"Theorem 1.10"},{"comment":"In Definition 1.3, the constant C is allowed to depend on σ, which is unusual for a definition of long-time control; the paper should specify whether the dependence is quantitative or merely existence of some C(σ).","section":"Definition 1.3"}],"recommendation":"major_revision","confidential_remarks":"The paper's main theorem is false as stated, but the counterexample is simple and the proof already contains the necessary missing condition (σ/σ* ≫ √log n). The technical core — the moment computations in Section 2 and their inhomogeneous adaptation in Section 3 — appears substantial. With a corrected Theorem 1.1 and completed proofs for the complex and non-symmetric cases, the paper could become publishable. I recommend major revision and a careful re-check of all passages marked 'details are omitted'."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"One thing you should know before reading this paper: the headline claim, Theorem 1.1, is false as stated. The proof only works in the regime sigma/sigma* >> sqrt(log n), but the theorem's \"in particular\" clause claims absence of outliers whenever sigma* = o(1/sqrt(log n)), with no condition on sigma/sigma*. The diagonal Gaussian example is clean: take b_ii = 1/sqrt(n), all other b_ij = 0. Then sigma = sigma* = 1/sqrt(n), so sigma* sqrt(log n) -> 0 and sigma/sigma* = 1, all assumptions hold. But rho(A) = max_i |g_i|/sqrt(n) ~ sqrt(2 log n) sigma, so P(limsup rho(A) > sigma) = 1, contradicting the theorem. The proof itself admits the gap: the absence-of-outliers conclusion is derived only when sigma/sigma* >> sqrt(log n), and in the intermediate regime the paper settles for a weaker bound with an extra D factor. This is a load-bearing error, not a typo.\n\nNow the credit. There is real substance here. The high-moment estimates in Section 2 are new and nontrivial, and the comparison with a homogeneous matrix is a clever way to handle inhomogeneous variance profiles. The long-time control parameter (Definition 1.3) is a useful idea, and the small deviation bound in Theorem 1.6 is plausible under that condition. The large deviation result for Gaussian doubly stochastic matrices (Theorem 1.10) appears to be the first of its kind, and the free probability argument is appropriate. The heavy-tailed extension (Theorem 1.11) is a natural generalization of Bordenave et al. and looks plausible.\n\nSoft spots: beyond the false Theorem 1.1, the complex case is stated with details omitted, and the non-symmetric case lives in an appendix. That would be acceptable in a paper under review, but the central theorem needs a fix. The abstract also has a reversed inequality (\"sigma_* >> (log n)^-1/2\" should be \"<<\"). The long-time control condition is restrictive, as the paper says, but that is a modeling choice, not a flaw.\n\nBottom line: substantial ideas, but a load-bearing incorrect statement. It deserves a serious referee only after the author corrects Theorem 1.1, either by adding sigma/sigma* >> sqrt(log n) to the absence-of-outliers claim or by restating the weaker bound. I would not cite it in its present form.","headline":"The headline claim of Theorem 1.1 is false as stated; the proof only covers sigma/sigma* >> sqrt(log n), and a diagonal Gaussian example is a clean counterexample.","tokens_in":51866,"tokens_out":4528,"would_cite":false,"duration_ms":39234,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60B20","60F10","60F15"],"pacs":[],"model":"deepseek-v4-flash","headline":"Inhomogeneous random matrices: spectral radius is governed by the variance profile, not the operator norm, down to optimal sparsity.","keywords":["spectral radius","inhomogeneous random matrices","variance profile","trace moment method","sparsity threshold","small deviations","large deviations","heavy-tailed entries"],"falsifier":"Construct a variance profile that provably fails the long-time control condition for any $\\sigma$ below the maximal row/column sum (for example a nonnegative matrix whose powers show transient amplification), scale the entries so that $\\sigma_*\\ll(\\log n)^{-1}$, and check numerically whether $P(\\rho(A)\\ge\\sigma(1+t\\sigma_*))$ stays below $C_0 n e^{-Ct}$; a violation would show the small-deviation theorem does not extend beyond its stated assumption.","tokens_in":50777,"feed_emoji":"🎲","tokens_out":16612,"duration_ms":140289,"temperature":0.7,"pith_summary":"This paper tries to establish that for square random matrices with independent, mean-zero, non-identically distributed entries, the spectral radius $\\rho(A)$ is controlled by the variance profile $S$, specifically by the maximal row/column variance sum $\\sigma$, rather than by the operator norm bound $\\|A\\|$, which would carry an extra factor of 2. It proves that after normalization, $\\rho(A) \\leq (1+\\epsilon)\\sigma$ with high probability whenever the largest entry standard deviation $\\sigma_*$ is $o((\\log n)^{-1/2})$, and that this sparsity scale is optimal. For finer fluctuations it introduces a 'long-time control' condition on $S$ and proves small-deviation bounds at the almost optimal scale $\\sigma_* \\log n$, along with a large-deviation bound for Gaussian entries with doubly stochastic variance and a heavy-tailed bound under only $2+\\epsilon$ moments. The motivation is that such inhomogeneous, often sparse matrices arise in neural networks, ecology, and non-Hermitian band matrices, where the variance profile is far from flat.","feed_headline":"Spectral radius tracks variance profile, not operator norm","feed_subtitle":"New bounds show a random matrix's spectral radius is governed by its variance sums, not by its operator norm, down to optimal sparsity.","key_machinery":"The proof rests on the trace moment method: for even $p$, $\\rho(A)^{2p}\\le \\operatorname{Tr}(A^p(A^*)^p)$, so bounding high moments bounds the spectral radius. The key is to compute $E[\\operatorname{Tr}(A^p(A^*)^p)]$ for $p$ growing with $n$—up to $p=O(n)$ for Theorem 1.1 and $p\\ll\\sqrt{n}$ for the small-deviation bounds—by enumerating closed paths, with independence forcing the dominant contributions to be non-backtracking or doubled paths. Comparison with a homogeneous matrix of effective size $\\lceil\\sigma^2\\rceil+p$ lets the variance profile enter only through $\\sigma$; the 'long-time control' parameter from Definition 1.3, meaning all powers of $S$ grow at most like $\\sigma^{2k}$, replaces $\\sigma$ at smaller sparsity scales, and a free-probability resolvent/Dyson equation handles the large-deviation case.","core_discovery":"The central claim is that for an $n\\times n$ matrix $A=(b_{ij}g_{ij})$ with independent mean-zero entries, the spectral radius $\\rho(A)$ is bounded, after normalization, by $(1+\\epsilon)\\sigma$, where $\\sigma$ is the largest row or column sum of the variance matrix $S=(b_{ij}^2)$, as long as the largest entry scale $\\sigma_*=\\max b_{ij}$ satisfies $\\sigma_*\\ll(\\log n)^{-1/2}$; in that regime almost surely no eigenvalues lie outside the disk of radius $\\sigma$. The paper also claims small-deviation control at the almost optimal scale: if $S$ is 'long-time controlled' by $\\sigma$, then $P(\\rho(A)\\ge\\sigma(1+t\\sigma_*))\\le C_0 n e^{-Ct}$ for sub-exponential entries, and, when $S$ is flat, a bound by $\\sqrt{\\rho(S)}$ with fluctuations of order $n^{-1/2}\\log n$. For Gaussian entries with doubly stochastic variance it claims a large-deviation inequality with the expected dependence $e^{-t^2/\\sigma_*^2}$, and for symmetric heavy-tailed entries with only $2+\\epsilon$ moments it claims boundedness by $\\sigma$ or $\\sqrt{\\rho(S)}$.","pith_inferences":["If these bounds are correct, stability analyses of linear systems driven by inhomogeneous random matrices can be carried out using only the variance profile's row/column sums, which may simplify criteria for balanced neural networks and ecological networks.","The long-time control condition suggests a testable dichotomy: variance profiles with transient amplification may exhibit spectral radius closer to the maximal row/column sum than to the long-time parameter, and simulating such profiles would reveal whether a new parameter is needed.","The sharp sparsity threshold $\\sigma_*\\ll(\\log n)^{-1/2}$ suggests a practical design rule: keep entry variances below $1/\\log n$ to avoid outlier eigenvalues in applications that rely on spectral radius estimates.","The large-deviation result for Gaussian doubly stochastic variances might extend to sub-Gaussian entries, but the paper notes its current method gives suboptimal rates; a direct numerical test would compare simulated tail probabilities against the $e^{-t^2/\\sigma_*^2}$ prediction."],"forward_implications":["For any variance profile with $\\sigma_*\\ll(\\log n)^{-1/2}$, the spectral radius of the inhomogeneous matrix is at most $(1+\\epsilon)\\sigma$ with high probability, so outliers are absent from the disk of radius $\\sigma$.","The sparsity scale is sharp: if entries are concentrated in diagonal blocks of size $d_n=o(\\log n)$, the spectral radius is eventually larger than any $a>1$, so $\\sigma_*\\ll(\\log n)^{-1/2}$ cannot be relaxed.","When the variance profile is long-time controlled, fluctuations of $\\rho(A)$ are at most of size $\\sigma\\sigma_*\\log n$ with high probability, i.e., near the scale predicted for Gaussian matrices.","For flat variance profiles, $\\rho(A)$ is bounded by $\\sqrt{\\rho(S)}$ plus fluctuations of order $n^{-1/2}(\\log n)^{1+c}$, improving previous bounds in both scale and tail probability.","For Gaussian entries with doubly stochastic variance, the upper tail of $\\rho(A)$ is sub-Gaussian with rate $\\sigma_*^{-2}$, the same dependence one would get if $\\rho(A)$ were Lipschitz in the entries."],"supporting_citations":[{"why":"supplies the operator-norm bound and the comparison argument that Theorem 1.1 sharpens by removing the factor 2.","marker":"[10]"},{"why":"prior result capturing optimal sparsity under a flatness assumption, which this paper extends to band-type inhomogeneous profiles.","marker":"[12]"},{"why":"earlier bound $\\rho(A)\\le\\sqrt{\\rho(S)}+n^{-1/2+\\epsilon}$ that Theorem 1.7 improves in scale and tail probability.","marker":"[3]"},{"why":"method for bounding the spectral radius without a fourth moment, adapted in Theorem 1.11 with an extra averaging step.","marker":"[14]"},{"why":"combinatorial reduction to diagrams used to compute high moments up to $p=O(n)$ for Theorem 1.1.","marker":"[28]"},{"why":"parity-counting scheme behind the $p\\ll\\sqrt{n}$ moment estimates in Theorem 1.6.","marker":"[50]"},{"why":"free-probability resolvent and matrix concentration framework used to prove the large-deviation bound Theorem 1.10.","marker":"[9]"},{"why":"Ginibre edge distribution used in Appendix A to show the sparsity scale $\\sigma_*\\ll(\\log n)^{-1/2}$ is optimal.","marker":"[47]"}],"fun_headline_variants":["Spectral radius pinned by variance sums, not operator norm","Inhomogeneous matrices: spectral radius tracks variance profile to optimal sparsity","New bounds: spectral radius controlled by entry variances down to log n scale","Spectral radius concentration for random matrices with non-identical entries","Variance profile determines spectral radius of inhomogeneous random matrices"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the entries have zero mean with sub-Gaussian (or sub-exponential) tails and that the variance profile admits a 'long-time control' parameter under which every power of $S$ grows at most like $\\sigma^{2k}$; the refined small-deviation bounds collapse if only the trivial choice $\\sigma$ equal to the maximal row/column sum is available.","fun_headline_variants_meta":{"raw":{"variants":["Spectral radius pinned by variance sums, not operator norm","Inhomogeneous matrices: spectral radius tracks variance profile to optimal sparsity","New bounds: spectral radius controlled by entry variances down to log n scale","Spectral radius concentration for random matrices with non-identical entries","Variance profile determines spectral radius of inhomogeneous random matrices"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000229,"raw_usage":{"total_tokens":1532,"prompt_tokens":1049,"completion_tokens":483,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":665,"completion_tokens_details":{"reasoning_tokens":395}},"tokens_in":665,"tokens_out":483,"duration_ms":4385,"temperature":1.0,"reasoning_tokens":395,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:35:08.400072+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Construct a variance profile that provably fails the long-time control condition for any $\\sigma$ below the maximal row/column sum (for example a nonnegative matrix whose powers show transient amplification), scale the entries so that $\\sigma_*\\ll(\\log n)^{-1}$, and check numerically whether $P(\\rho(A)\\ge\\sigma(1+t\\sigma_*))$ stays below $C_0 n e^{-Ct}$; a violation would show the small-deviation theorem does not extend beyond its stated assumption.","supporting_citations":[{"cited_title":"Sharp nonasymptotic bounds on the norm of random matrices with independent entries","cited_arxiv_id":null,"evidence_quote":"supplies the operator-norm bound and the comparison argument that Theorem 1.1 sharpens by removing the factor 2."},{"cited_title":"Spectral radii of sparse random matrices","cited_arxiv_id":null,"evidence_quote":"prior result capturing optimal sparsity under a flatness assumption, which this paper extends to band-type inhomogeneous profiles."},{"cited_title":"Spectral radius of random matri- ces with independent entries","cited_arxiv_id":null,"evidence_quote":"earlier bound $\\rho(A)\\le\\sqrt{\\rho(S)}+n^{-1/2+\\epsilon}$ that Theorem 1.7 improves in scale and tail probability."},{"cited_title":"On the spectral radius of a random matrix: an upper bound without fourth mosment","cited_arxiv_id":null,"evidence_quote":"method for bounding the spectral radius without a fourth moment, adapted in Theorem 1.11 with an extra averaging step."},{"cited_title":"A universality result for the smallest eigenvalues of certain sample covariance matrices","cited_arxiv_id":null,"evidence_quote":"combinatorial reduction to diagrams used to compute high moments up to $p=O(n)$ for Theorem 1.1."},{"cited_title":"Central limit theorem for traces of large random sym- metric matrices with independent matrix elements","cited_arxiv_id":null,"evidence_quote":"parity-counting scheme behind the $p\\ll\\sqrt{n}$ moment estimates in Theorem 1.6."},{"cited_title":"Matrix concen- tration inequalities and free probability","cited_arxiv_id":null,"evidence_quote":"free-probability resolvent and matrix concentration framework used to prove the large-deviation bound Theorem 1.10."},{"cited_title":"A limit theorem at the edge of a non-Hermitian random matrix ensemble","cited_arxiv_id":null,"evidence_quote":"Ginibre edge distribution used in Appendix A to show the sparsity scale $\\sigma_*\\ll(\\log n)^{-1/2}$ is optimal."}],"review_version":1}