{"id":"e3fbad74-0b6c-4d12-9887-04ce19b909f0","arxiv_id":"2502.15752","paper_version":4,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Under block dependence, m-dependence, and weak β-mixing, high-dimensional logistic regression risks are Gaussian-universal, and a new low-rank CGMT gives the exact asymptotic effect of data augmentation on test risk.","lead":"This paper proves that Gaussian universality still holds for high-dimensional logistic regression when training data are dependent, and it builds a new convex Gaussian min-max theorem that handles correlation across both data points and features. The tools yield an exact asymptotic account of how data augmentation, such as random permutation or sign flipping, changes test risk.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"DA and universality theorems hold for the S_p-constrained minimizer, but the actual logistic estimator is unconstrained; the needed ℓ∞ bound is explicitly deferred in Appendix D.1, so the headline claims are not yet proved for the estimator being used.","rationale":"The reader's weakest_assumption focused on Assumption 7 and the interior-solution condition for (DO), which is indeed unverified. But the S_p restriction is even more load-bearing because it blocks the DA theorem before Assumption 7 is even reached. Theorem 12 states its result for \\hat\\beta fitted via (OO) with S=S_p, while the DA procedure in Section 6 and equation (11) minimizes over R^p. Appendix D.1 explicitly defers the \\ell_\\infty bound that would identify the constrained and unconstrained minimizers, and it notes why the standard singular-value route fails under dependence. This is not a minor technicality: data augmentation introduces within-block dependence that can make the augmented design rank-deficient, so the usual proof mechanism cannot apply. The paper is otherwise strong: the dependent CGMT genuinely extends prior variants, the universality proofs adapt the Lindeberg interpolation to block dependence, and the simulations include non-sub-Gaussian stress tests. The issue is not internal inconsistency but a mismatch between the stated theorems and the proven statements. A conditional verdict remains appropriate: the central claims may be true, but the DA headline is not yet supported for the actual estimator. I therefore do not move the reader's verdict, but I identify a different principal gap.","tokens_in":86691,"tokens_out":6929,"duration_ms":70684,"concrete_test":"Analytically re-derive an \\ell_\\infty bound for the random-permutation DA scheme (k=2, one group) using the leave-one-block-out argument of Karoui (2013) / Han-Shen (2023). Show whether the bound requires a lower bound on \\sigma_min of the augmented design matrix. Since the augmented matrix contains two copies of each original row when the permutation is the identity on a block, \\sigma_min=0, so any proof relying on such a lower bound cannot work. If instead a valid \\ell_\\infty bound can be derived without this condition, the S_p gap is benign; if not, the constrained-to-unconstrained step and hence Theorem 12 are unproved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing gap is the S_p restriction. Section 2 defines S_p and says the restriction 'becomes equivalent to the unconstrained minimization when one proves that \\hat\\beta\\in S_p with high probability.' Theorem 15, the actual training-risk result, proves universality for any \\tilde S\\subseteq S_p; setting \\tilde S=S_p gives a constrained version of Theorem 2(5). Appendix D.1 admits that the needed \\ell_\\infty bound is not proved: 'To extend this into a proof covering all DA schemes will be left to future work, and for now we will work under the mild constraint that \\hat\\beta\\in S_p.' The DA estimator in (11) and Theorem 12, however, minimizes over R^p, not S_p. The standard leave-one-out route to \\ell_\\infty bounds requires \\sigma_min(XX^\\top)\\ge p/C (Montanari-Saeed), and D.1 notes this is 'nearly impossible' in dependent setups: identical rows in a block make \\sigma_min=0. Data augmentation creates exactly this kind of within-block dependence, so the unresolved S_p membership is not cosmetic. Until \\|\\hat\\beta\\|_\\infty\\le L p^{1/2-r} is shown for the DA minimizer, neither the training-risk universality nor the test-risk equation (Theorem 12) applies to the practical estimator, and the abstract's 'establish the impact of data augmentation' outruns the proof. Assumption 7 is a separate, secondary gap; the S_p issue already blocks the DA application.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper extends two workhorses of high-dimensional statistics — Gaussian universality and the convex Gaussian min-max theorem (CGMT) — from independent observations to dependent ones. For penalized logistic regression with labels generated by a logistic link, it proves universality of the training and test risks (matching to a Gaussian surrogate with the same covariance) under block dependence (Theorem 2), m-dependence, and a restricted class of β-mixing processes (Theorem 3). The CGMT part (Theorem 5/13) handles Gaussian design matrices whose covariance is a sum of Kronecker products of row- and column-side matrices (Assumption 10), with comparison probabilities off by a factor 2^M. The tools are applied to data augmentation schemes — random permutations, sign flipping, and cropping — culminating in a ten-equation fixed-point system (EQs) claimed to characterize the asymptotic test risk (Theorem 12). Simulations across coordinate distributions (Gaussian, uniform, gamma, exponential, t₃) confirm the predicted risk values and illustrate that full permutation invariance helps while partial structure knowledge does not.","tokens_in":86969,"tokens_out":12602,"duration_ms":109646,"significance":"If the main theorems hold, this is a substantial contribution: it removes the independence-of-rows assumption in a principled way for high-dimensional logistic regression, and Theorem 5 is a model-independent tool — the 2^M comparison bounds (Lemmas 45–46) and the low-rank Kronecker-factor covariance structure (Assumption 10) are likely to be reused beyond this paper. The proof work is real: the within-block Lindeberg control (Lemmas 24–28) is a genuine extension of the Montanari–Saeed template, and the deterministic ten-equation system (EQs) is anchored by Lemma 52, which recovers exactly Salehi et al.'s six equations in the isotropic no-augmentation limit. The authors ship reproducible code (footnote in Section 6), give detailed simulation settings (Appendix C), and are unusually candid about what is not proved (the S_p membership in Appendix D.1 and the non-universality of training trajectories in Section 9). The main caveat is that the data-augmentation application's headline claims are conditional on the deferred ℓ∞ bound and on unverified geometric properties of the (DO) limit; these gaps are load-bearing for the practical, unconstrained estimator.","major_comments":[{"comment":"The paper proves universality for the S_p-constrained minimizer (Theorem 15, for any S̃ ⊆ S_p), yet Theorem 2(5) and Theorem 3(7) are stated with the unconstrained min_β R̂_n(β;·), and the practical estimator (3) is the unconstrained minimizer. Section D.2 asserts that Theorem 15 implies (5) 'by setting S̃ = S_p', but this step requires the ℓ∞ membership \\hatβ ∈ S_p, i.e. ‖\\hatβ‖_∞ ≤ L p^{1/2−r}. Appendix D.1 explicitly leaves this to future work and explains that the standard leave-one-out singular-value route (σ_min(XX^T) ≥ p/C) is 'nearly impossible' when blocks contain identical rows — precisely the within-block dependence created by data augmentation. Until this membership is established for the DA estimators, the abstract's claim to 'establish the impact of data augmentation' overstates what is proved: Theorem 12 applies to the S_p-restricted (OO), not to the unconstrained estimator (3) that is fitted in the simulations. The authors should either supply the ℓ∞ bound for the DA schemes they study, or restate the main theorems and the DA conclusions in explicitly constrained form.","section":"Section 2 (Eq. 4), Section D.2, Appendix D.1"},{"comment":"Test-risk universality (Theorem 2(6)) requires Assumption 7, and the proof of Theorem 12 verifies Assumption 7 by invoking two premises: that the minimizer-maximizers of (DO) lie in the interior of the domain of optimization, and that restricting the β-domain to |(β^T Σ_new β)^{1/2} − χ̄| > ε changes the (DO) value by Θ(ε²). The first is stated only as an assumption of the theorem, and the second is asserted without derivation. These premises are exactly what makes the Gaussian training risk sharply minimized on the thin shell; no verification for the permutation, sign-flip, or crop schemes is provided. The conclusion |R_test(\\hatβ(X,X^Φ)) − R̄_test(χ̄²)| → 0 is therefore conditional on an unverified geometric property of the deterministic limit, and this should be stated as an open condition rather than presented as a completed verification.","section":"Section B.1, Theorem 12 and its proof"},{"comment":"The Lindeberg interpolation yields an error bound of order √k γ + √k δ + α^{−1} log(1/δ) + τ, so the result holds for each fixed block size k and contains no control uniform in k. In the data-augmentation application, k (the number of augmentations) is a free parameter that the simulations vary (Figures 1 and 2), and the text draws conclusions 'as more and more augmentations are used'. The theorems and the (EQs) derivation cover the regime where k is held fixed as p → ∞; the k → ∞ regime is outside the current proof. This limitation should be stated explicitly wherever the DA conclusions are drawn, since the figures' x-axis is exactly the parameter that the theory does not let grow.","section":"Section D.5, Eqs. (16)–(19) and the final display of the proof of Theorem 15"}],"minor_comments":[{"comment":"The phrase 'assumption that significantly limit its applicability' has a subject-verb agreement error and should read 'significantly limits'.","section":"Abstract"},{"comment":"The theorem statement says the test risk is characterized by the parameter set (r, θ, σ, τ), while the system (EQs) solves for ten parameters (α, σ_1, σ_2, τ_1, τ_2, ν_1, ν_2, r_1, r_2, θ); the notation should be aligned between the statement and the equations.","section":"Section B.1, Theorem 12"},{"comment":"The word 'estbalished' appears in the proof and should be corrected to 'established'.","section":"Section B.1, Theorem 12 proof"},{"comment":"Eq. (11) restricts the minimization to S_p while the original problem (3) is unconstrained; this distinction is only explained via Remark 1 and Appendix B.3, and it should be stated in the main text that all DA theorems concern the S_p-constrained problem, with the unconstrained version open (see Major Comment 1).","section":"Section 6, Eq. (11) and Section 1.1, Eq. (3)"},{"comment":"The symbol 'B' is used as a substitute for '≔' throughout the manuscript; it is nonstandard and should be defined at first use or replaced by standard notation.","section":"Section 2 and throughout"},{"comment":"The curves for r_perm = 0.8 and no augmentation appear to overlap, and one of the paper's central findings is that partial permutations are no better than no augmentation; the caption should report the trial counts (50 versus 200) and the error-bar convention so that the 'within error margins' claim is checkable.","section":"Section 6, Figure 1"}],"recommendation":"major_revision","confidential_remarks":"This is honest, high-quality work whose principal gap is explicit: the ℓ∞/S_p membership for dependent designs with duplicate or highly correlated blocks is not proved for the DA schemes, and the paper itself acknowledges this in Appendix D.1. I would encourage the editors to ask for either a proof of the S_p membership for at least the permutation scheme (where rows within a block are dependent but not identical, so a singular-value lower bound may be within reach) or a restructuring of the headline claims to the constrained statement. The 'for the first time' claims for the dependent CGMT should be checked carefully against Dhifallah–Lu (2021) and Akhtiamov et al. (2024a); Remark 8 does address the comparison, but the abstract is stronger than that remark. The paper is a good fit for the venue and the data-augmentation section will be of interest to the ML theory community."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The dependent CGMT (Theorem 5/13) is the real contribution here. It genuinely generalizes the standard CGMT, Dhifallah-Lu's multivariate version, and Akhtiamov et al., and the 2^M comparison bound looks like the right mechanism. The universality results under block dependence, m-dependence, and mixing are also new, and the proofs are serious: the Montanari-Saeed Lindeberg template with joint within-block control is a real extension, not a cosmetic tweak. The 10-equation system collapsing to Salehi et al. in the no-augmentation limit is a good sanity check, and the simulations include non-sub-Gaussian distributions, which is honest stress-testing.\n\nThe soft spot is the data augmentation application. The headline theorems hold for the S_p-constrained minimizer, but the DA estimator in (11) minimizes over R^p. Section 2 says the constraint becomes equivalent once one proves the unconstrained minimizer lies in S_p with high probability, and Appendix D.1 explicitly says that proof is not done: the needed infinity-norm bound is \"left to future work.\" This is not a technicality. Data augmentation creates within-block dependence where rows can be identical, making sigma_min(XX^T) exactly zero, so the standard leave-one-out route to the infinity-norm bound is unavailable. Until that is resolved, Theorem 12 does not apply to the estimator actually being fitted, and the abstract's claim to \"establish the impact of data augmentation\" outruns the proof. Assumption 7's interior-solution condition is a separate, real gap, but the S_p issue already blocks the DA application.\n\nMinor issues: the theory is per-fixed-k, so the cross-k benefit curves are simulations; figures lack error bars; the code URL appears malformed. None of these affects the core CGMT or universality results.\n\nThis paper deserves a serious referee. The dependent CGMT alone justifies referee time, and the universality theorems are substantial. A referee should focus on the S_p membership question and whether Assumption 7 can be verified without the interior condition. If the DA claims are softened to apply conditionally on S_p membership, the paper is strong. I would engage with it.","headline":"The dependent CGMT and universality results are real and important; the data augmentation application is not yet proved for the unconstrained estimator because the S_p restriction is left unresolved.","tokens_in":87678,"tokens_out":1869,"would_cite":true,"duration_ms":20558,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F12","62J12","60F05"],"pacs":[],"model":"deepseek-v4-flash","headline":"For penalized logistic regression in the proportional regime, Gaussian universality and the convex Gaussian min-max theorem hold even for dependent data, and the asymptotic risk depends only on the data's mean and covariance — with data…","keywords":["high-dimensional logistic regression","Gaussian universality","convex Gaussian min-max theorem","block dependence","m-dependence","beta-mixing","data augmentation","proportional asymptotics"],"falsifier":"Run the paper's own comparisons at block sizes and aspect ratios beyond the simulated range: if the excess test risk of penalized logistic regression on, say, shifted-gamma covariates differs from the matched-covariance Gaussian surrogate by more than the $\\sqrt{k}$ universality bound the authors derive, the claim that dependence enters only through the covariance fails. For the augmentation conclusion, check whether increasing the number of 80%-permutation augmentations eventually lowers the test risk at large $k$; if it does, the qualitative claim that partial invariance is as good as no augmentation would be overturned.","tokens_in":86304,"feed_emoji":"📈","tokens_out":12818,"duration_ms":100454,"temperature":0.7,"pith_summary":"This paper tries to establish that the two standard tools of exact high-dimensional statistics — Gaussian universality and the convex Gaussian min-max theorem (CGMT) — remain valid when the data are dependent rather than independent. For penalized logistic regression in the proportional regime ($p/n \\to \\kappa$), it proves that the asymptotic training and test risks under block dependence, $m$-dependence, and certain $\\beta$-mixing processes coincide with the risks of a Gaussian model with matching mean and covariance (Theorems 2 and 3). It then builds a new CGMT that tolerates correlated rows and columns as long as the covariance has a low-rank factored structure (Theorem 5), and applies both tools to data augmentation, producing a ten-equation system that characterizes the test risk exactly (Theorem 12). The concrete payoff: full random permutation of exchangeable coordinates markedly lowers the test risk, while an 80% permutation leaves it statistically unchanged from no augmentation. If the paper is right, dependence enters high-dimensional risk analysis only through the first two moments, and practitioners receive a quantitative warning that partial invariance knowledge buys almost nothing.","feed_headline":"Universality holds for dependent high-dimensional logistic regression","feed_subtitle":"Asymptotic risk is set by mean and covariance alone; augmentation pays off only with full structural knowledge.","key_machinery":"The machinery has three load-bearing parts. First, the block-dependent universality theorem (Theorem 2) uses a Lindeberg-style interpolation between the original data and a matched-covariance Gaussian surrogate, with Assumption 5 requiring joint Gaussian approximation of $k$-tuples of projections $X_{i+r}^{\\top}\\beta_r$ for $r \\le k$, weighted on the sphere, instead of the single projection needed in the independent case. Second, the dependent CGMT (Theorem 5) handles a Gaussian min-max problem whose design matrix $H$ has covariance $\\mathrm{Cov}[H_{ji},H_{j'i'}] = \\sum_{l=1}^{M} \\Sigma^{(l)}_{jj'}\\tilde{\\Sigma}^{(l)}_{ii'}$, a low-rank sum of factored row and column covariances; for data augmentation $M = 2$ suffices because the covariance of augmented copies is determined by $\\mathrm{Var}[\\phi_1(Z_1)]$ and $\\mathrm{Cov}[\\phi_1(Z_1), \\phi_2(Z_1)]$. Third, Assumption 7 — that the Gaussian training risk is sharply minimized on the thin shell $|(\\beta^{\\top}\\Sigma_{\\text{new}}\\beta)^{1/2} - \\bar{\\chi}| \\le \\epsilon$ — is what converts training-risk universality into test-risk universality, and the paper verifies it through the CGMT so long as the minimizer-maximizers of the deterministic auxiliary problem lie in the interior of its domain. The final output is the system of ten equations (EQs) whose solution pins down the test risk through the scalar $\\bar{\\chi}$.","core_discovery":"The central claim is that the asymptotic risk of high-dimensional penalized logistic regression is universal across data distributions once the first two moments are fixed, even when the observations are dependent. Theorems 2 and 3 show that the minimum training risk and the test risk of the estimator on block-dependent, $m$-dependent, or suitably $\\beta$-mixing data converge to those of a Gaussian dataset with the same covariance; a key consequence stated in the paper is that uncorrelated dependent data give the same asymptotic risk as independent data. Because the Gaussian surrogate may still have correlated rows and columns, the paper proves a dependent CGMT (Theorem 5) under an assumption that the design covariance factorizes as a sum of Kronecker products of row and column covariance matrices, and uses it to verify the sharp-shell condition that upgrades training-risk universality to test-risk universality. For data augmentation, the estimator trained on $k$ transformed copies of each observation is shown to have a test risk characterized by a deterministic system of ten scalar equations (Theorem 12); the numerical solution shows that full permutation of exchangeable coordinates reduces the test risk substantially, whereas $r_{\\text{perm}} = 0.8$ permutations, sign flipping, and cropping without full knowledge of the zero coordinates yield risks within error margins of no augmentation.","pith_inferences":["If the mean-variance reduction is as general as the proof suggests, the same universality pipeline should transfer to other classifiers whose loss depends on the data through one-dimensional projections, such as SVMs and other generalized linear models — the paper notes this extension is direct.","The $r_{\\text{perm}} = 0.8$ plateau suggests a testable design rule for practitioners: augmentation groups that cover only part of the true symmetry may be wasting compute, and symmetry-learning methods should beat hand-chosen partial augmentations even though the paper does not run that comparison.","The simulations' observation that training trajectories need different learning rates for t-distributed versus uniform data, even though global minima are universal, indicates that risk universality will not extend to optimization dynamics; proving trajectory non-universality would be a natural sequel."],"forward_implications":["Uncorrelated but dependent data inherit all previously derived independent-data results for logistic regression, since the asymptotic risk is governed by mean and covariance alone.","Correlated dependent data can be studied by replacing the data with a Gaussian model of the same covariance and applying the dependent CGMT, avoiding random-matrix theory.","Data augmentation under the full permutation group of exchangeable coordinate blocks improves the test risk as more augmentations are used.","Augmenting only a large fraction of the coordinates ($r_{\\text{perm}} = 0.8$) leaves the test risk statistically unchanged relative to no augmentation.","Sign flipping and cropping that do not know the exact zero coordinates of the signal give the same test risk as no augmentation."],"supporting_citations":[{"why":"Supplies the Lindeberg interpolation method and the $d_H$ distance metric that the block-dependent universality proof in Section D builds on.","marker":"Montanari and Saeed (2022)"},{"why":"Provides the CGMT analysis of regularized logistic regression whose six scalar equations are recovered in the no-augmentation isotropic limit, pinning down the system (EQs).","marker":"Salehi et al. (2019)"},{"why":"The standard CGMT (Theorem 3.3.1) whose proof structure Theorem 13 extends to dependent rows and columns.","marker":"Thrampoulidis (2016)"},{"why":"The Gaussian min-max comparison inequality underlying all CGMT-type equivalences in the paper.","marker":"Gordon (1985)"},{"why":"The multivariate CGMT with block-diagonal covariances, recovered as a special case of Assumption 10, and the earlier noise-injection augmentation analysis.","marker":"Dhifallah and Lu (2021)"},{"why":"The embedding result that couples $\\beta$-mixing blocks with independent blocks, used in the mixing proof in Section F.","marker":"Yu (1994)"},{"why":"Prior block-dependence universality within covariate vectors with independent rows; the $S_p$ restriction discussion in Section D.1 extends their leave-one-out perspective.","marker":"Lahiry and Sur (2024)"}],"fun_headline_variants":["Dependent data no longer blocks logistic regression universality","New CGMT handles dependent data, revealing when augmentation pays off","For dependent data, risk is universal; augmentation needs full structure"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the constrained minimizer on the set $S_p$ is the true minimizer and that, in the Gaussian surrogate, the training risk is uniquely minimized on a thin shell of constant $(\\beta^{\\top}\\Sigma_{\\text{new}}\\beta)^{1/2}$; the paper verifies the shell condition only when the auxiliary deterministic optimization has interior solutions, and a general $S_p$ bound for all data-augmentation schemes is explicitly left to future work.","fun_headline_variants_meta":{"raw":{"variants":["Dependent data no longer blocks logistic regression universality","New CGMT handles dependent data, revealing when augmentation pays off","For dependent data, risk is universal; augmentation needs full structure"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001148,"raw_usage":{"total_tokens":4779,"prompt_tokens":980,"completion_tokens":3799,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":596,"completion_tokens_details":{"reasoning_tokens":3745}},"tokens_in":596,"tokens_out":3799,"duration_ms":24066,"temperature":1.0,"reasoning_tokens":3745,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T14:30:12.638637+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the paper's own comparisons at block sizes and aspect ratios beyond the simulated range: if the excess test risk of penalized logistic regression on, say, shifted-gamma covariates differs from the matched-covariance Gaussian surrogate by more than the $\\sqrt{k}$ universality bound the authors derive, the claim that dependence enters only through the covariance fails. For the augmentation conclusion, check whether increasing the number of 80%-permutation augmentations eventually lowers the test risk at large $k$; if it does, the qualitative claim that partial invariance is as good as no augmentation would be overturned.","supporting_citations":[{"cited_title":"Recovering structured signals in high dimensions via non-smooth convex optimization: Precise performance analysis","cited_arxiv_id":null,"evidence_quote":"The standard CGMT (Theorem 3.3.1) whose proof structure Theorem 13 extends to dependent rows and columns."},{"cited_title":"Some inequalities for G aussian processes and applications","cited_arxiv_id":null,"evidence_quote":"The Gaussian min-max comparison inequality underlying all CGMT-type equivalences in the paper."},{"cited_title":"On the inherent regularization effects of noise injection during training","cited_arxiv_id":null,"evidence_quote":"The multivariate CGMT with block-diagonal covariances, recovered as a special case of Assumption 10, and the earlier noise-injection augmentation analysis."},{"cited_title":"Universality in block dependent linear models with applications to nonparametric regression","cited_arxiv_id":null,"evidence_quote":"Prior block-dependence universality within covariate vectors with independent rows; the $S_p$ restriction discussion in Section D.1 extends their leave-one-out perspective."}],"review_version":1}