Pith. sign in

REVIEW 6 minor 1 cited by

Phase transitions for the existence of unregularized M-estimators in single index models

T0 review · 0 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This paper proves that one scalar threshold, $\delta_\infty$, simultaneously controls whether an unregularized M-estimator exists and whether the asymptotic system characterizing it has a unique solution, across smooth single-index models.

desk verdict A solid, important paper that closes a real proof gap in high-dimensional M-estimator theory; the C^1 caveat is explicit and does not sink the central theorem. read the letter →

arxiv 2501.03163 v3 pith:NJSPULEZ submitted 2025-01-06 math.ST stat.TH

classification math.STstat.TH MSC 62F1262J12
keywords M-estimatorsphasetransitionsingle-indexmodelsGaussiandesignlogisticregressionPoissonbinomialproportionalasymptotics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper establishes that in high-dimensional single-index models with Gaussian covariates and smooth, strictly convex losses, a single scalar threshold $\delta_\infty$ controls two separate facts: whether the unregularized M-estimator exists, and whether the nonlinear system that describes its asymptotic behavior has a unique solution. The threshold is given explicitly as the reciprocal of the minimum of a one-dimensional convex function $\varphi(t)$ that depends only on the law of the response and the link. This unifies and generalizes the known logistic-regression phase transition to models such as Poisson and binomial regression. The proof closes a gap that had previously only been verified numerically outside the global null.

What carries the argument

The central object is the infinite-dimensional convex program $\min_{(a,v)\in\mathbb{R}\times H} \mathbb{E}[\ell_Y(aU + v)]$ subject to $\|v\| - \mathbb{E}[vG]/\sqrt{1-\delta^{-1}} \le 0$, where $H$ is the Hilbert space of square-integrable functions of $(G,U,Y)$ and $G$ is an independent standard normal. Theorem 3.1 shows that minimizers of this program correspond one-to-one with solutions of the nonlinear system whenever the Lagrange multiplier is positive, so the phase transition for the system reduces to the existence of this minimizer. The threshold $\delta_\infty$ enters through the cone $C$ of directions along which the loss can only be constant or decreasing; the variational representation $\delta_\infty^{-1} = \inf_{(t,p)\in C} \mathbb{E}[(G-p)^2]$ yields the direction $p^*$ used to show that no minimizer exists for $\delta \le \delta_\infty$, while the same $p^*$ gives coercivity for $\delta > \delta_\infty$. The smoothness of the loss is used in Lemma 3.2 to rule out the degenerate minimizer $v^* = 0$.

What would settle it

Choose the Poisson model with a nonzero signal, compute $\delta_\infty$ by minimizing $\varphi$ from (4), then numerically solve the system (2) over a grid of $\delta$ values; any solution found for $\delta < \delta_\infty$, or no solution found for $\delta > \delta_\infty$, would refute Theorem 2.7.

Watch

Extended reading notes

Core claim

For $n/p \to \delta$, define $\delta_\infty$ by $1/\delta_\infty = \inf_{t\in\mathbb{R}} \varphi(t)$, where $\varphi$ is a convex function built from the events where the loss is coercive, increasing, or decreasing. Theorem 2.6 states that the unregularized M-estimator exists with probability tending to one when $\delta > \delta_\infty$ and with probability tending to zero when $\delta < \delta_\infty$. Theorem 2.7 states that the nonlinear system of equations characterizing the asymptotic behavior of the estimator has no solution when $\delta \le \delta_\infty$ and a unique solution when $\delta > \delta_\infty$. The two theorems are proved through an infinite-dimensional convex optimization problem whose minimizer is equivalent to a solution of the system, with the technical core being a non-degeneracy lemma showing that the minimizer does not collapse to $v^* = 0$.

Load-bearing premise

The if-and-only-if result depends on every loss being $C^1$ and strictly convex; the paper itself notes that non-differentiable losses can introduce a different threshold, so if smoothness fails the central claim is not generic.

Editorial extensions

If this is right

  • For Poisson and binomial regression with Gaussian covariates and smooth losses, the threshold $\delta_\infty$ from (4) rigorously separates existence from nonexistence of the unregularized M-estimator.
  • Any asymptotic analysis of these M-estimators based on the Convex Gaussian Minmax Theorem can now invoke Theorem 2.7 to justify the required unique solution of the system whenever $\delta > \delta_\infty$.
  • The global-null logistic result $\delta_\infty = 2$ is contained as a special case, and the threshold is available for non-null single-index models.
  • In the non-existence regime $\delta \le \delta_\infty$, the proof exhibits an explicit ray along which the objective decreases, giving a constructive certificate that no minimizer exists.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The authors do not pursue non-Gaussian designs, but the cone-projection form of $\delta_\infty$ suggests an analogous variational formula for elliptical covariate distributions.
  • A consequence the authors leave implicit is that for smooth losses, existence of the estimator and solvability of its asymptotic system are the same event, which may be a general structural fact rather than a special property of the examples treated.
  • The remark on $\delta_{\text{perfect}}$ invites a concrete follow-up: for nonsmooth losses, test whether the phase transitions for the estimator and for the nonlinear system occur at different thresholds.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

0 major / 6 minor

Summary. The paper studies unregularized M-estimators in single-index models with Gaussian covariates under proportional asymptotics n/p → δ ∈ (1, ∞). It defines a threshold δ∞ as the infimum of an explicit convex functional φ(t) of the data distribution and proves two main results. Theorem 2.6 establishes a phase transition for the existence of the M-estimator at δ∞, generalizing Candès & Sur (2020) beyond binary logistic regression. Theorem 2.7 proves that the asymptotic nonlinear system (2) of Sur & Candès has a unique solution if and only if δ > δ∞, closing a gap noted in the literature. The proof proceeds through an equivalence (Theorem 3.1) between system (2) and an infinite-dimensional convex program (5), a non-degeneracy lemma (Lemma 3.2), a uniqueness lemma (Lemma 3.3), and a phase-transition theorem for the convex program (Theorem 3.4).

Significance. If correct, the results provide a comprehensive theoretical foundation for proportional-asymptotic analyses of M-estimators that assume existence of a solution to the asymptotic system. The threshold δ∞ is given by an explicit, parameter-free convex program, and both theorems are proved with complete arguments in the appendices. The main results are numerically verified for Poisson and binomial losses. The paper is careful about its scope: the smoothness assumption on the loss (Assumption 2.5(1)) is stated up front, and the authors explicitly note that the non-degeneracy conclusion changes for non-differentiable losses (Bellec & Koriyama, 2023). The central claim is internally consistent under the stated assumptions.

minor comments (6)
  1. [Section 3, paragraph after Theorem 2.6] The sentence referring to 'the nonlinear system (4)' should refer to 'the nonlinear system (2)'.
  2. [Appendix E, Lemma E.2] The displayed bound '∥v − ṽ∥ ≤ C^(2)(ξ)(1 + ∥v∥²)' should use a linear term '1 + ∥v∥' rather than '1 + ∥v∥²'. The subsequent inequalities (28)–(29) require the linear version, and the proof's earlier bound on |a| is indeed linear in ∥v∥.
  3. [Appendix A, Lemma A.1 proof] The computation of E[p_*²] contains a typographical clutter with repeated terms 'E[p_*²] + 2t_* E[p_*U] E[p_*U] = 0'; this should be simplified to a clean derivation.
  4. [Section 4, Numerical Simulation] The phrase 'using linear programming' for the Poisson loss is potentially confusing because the loss is exponential; presumably the authors solve the linear feasibility problem (12) to detect existence, and this should be stated explicitly.
  5. [Section 4, Figures 2 and 3] The figures would be more informative if they included error bars or standard errors over the 20 repetitions.
  6. [Theorem 2.6] The theorem does not address the boundary case δ = δ∞; a sentence noting that this case is not covered (and why) would improve precision.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the phase-transition threshold is an independently defined variational quantity and the system-solvability iff is proved, not fitted.

full rationale

The central claim, Theorem 2.7, is not circular. The threshold delta_infinity is defined in (4) by an explicit variational infimum over a convex function phi(t) built from the loss and the distribution of (U, Y); no unknown constant is fitted to the nonlinear system (2). Theorem 2.6 derives the M-estimator existence phase transition from this same threshold using Lemma A.2 and the Gaussian kinematic formula, and the proof shows the empirical distance to the relevant cone converges to 1/delta_infinity by the law of large numbers. Theorem 2.7 is proved through an independent chain: Theorem 3.1 proves a bijective equivalence between solutions of system (2) and minimizers of the infinite-dimensional program (5) with v* != 0; Lemma 3.2 proves non-degeneracy; Lemma 3.3 proves uniqueness; and Theorem 3.4 proves that (5) has a minimizer if and only if delta > delta_infinity. The two directions of Theorem 3.4 use Lemma A.1's representation delta_infinity^{-1} = inf_{(t,p) in C} E[(G-p)^2] and direct coercivity arguments; the threshold is not chosen to make either direction true. The paper's references to the authors' own prior work, e.g. 'The notation and setup of this infinite-dimensional convex optimization problem is heavily inspired by Bellec & Koriyama (2023)' and 'The argument in this proof is inspired by the proof of Lemma 2.6 in (Bellec & Koriyama, 2023)', are methodological or technical, not load-bearing for the main theorem; the core equivalence and phase-transition proofs are contained in the paper. The explicit acknowledgement that non-differentiable losses can produce a different threshold delta_perfect (Bellec & Koriyama, 2023) is a scope restriction stated up front, not a hidden circular step. No prediction reduces to a fitted value, and no theorem is imported from a self-citation as the sole justification for the central claim.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

The central claims rest on standard tools (Gaussian kinematic formula, convex analysis in Hilbert spaces) and on explicit domain assumptions: Gaussian design, single-index structure, C^1 strictly convex losses with controlled growth, and non-triviality conditions. There are no fitted constants and no invented entities; the threshold is derived from the data-generating model rather than fit to the system it is claimed to govern.

assumptions (6)
  • domain assumption Gaussian design and single-index model: x_i iid N(0, I_p), y_i depends on x_i only through x_i^T w with ||w||=1; (U,Y) = (x_i^T w, y_i) with U ~ N(0,1).
    This is the model definition (Section 1) and is used to apply rotational invariance and the Gaussian kinematic formula in the proof of Theorem 2.6 (Appendix A).
  • domain assumption Loss differentiability and strict convexity: l_y is C^1 and strictly convex for every y, and P(Omega_v) < 1 (Assumptions 2.4 and 2.5).
    Strict convexity is used for uniqueness of the minimizer and of the Lagrange multiplier (Lemma 3.3); differentiability is essential in Lemma 3.2 to rule out v* = 0. The authors note the non-differentiable case would require a different threshold delta_perfect, so this assumption is load-bearing.
  • domain assumption Affine one-sided lower bound on the loss with square-integrable intercept D(Y): l_Y(u) >= -D(Y) + (1/b) times (u, |u|, or -u depending on the event) (Assumption 2.5(4)).
    Used in Lemma E.2 to prove coercivity of the infinite-dimensional objective on the delta > delta_infinity side, which is necessary for existence of a minimizer.
  • standard math Gaussian kinematic formula and statistical dimension theory of Amelunxen et al. (2014).
    Used in the proof of Theorem 2.6: the phase transition for existence of the M-estimator is obtained by comparing the statistical dimension of the cone C(u,y) with p/n (Appendix A, after Lemma A.2).
  • standard math Convex analysis in Hilbert spaces: Slater's condition, KKT conditions, coercivity, and the relevant propositions from Bauschke and Combettes (2017).
    Used to prove the equivalence between the nonlinear system (2) and the infinite-dimensional program (5) (Lemmas B.1-B.3, Theorem 3.1, Lemma E.2).
  • domain assumption Non-triviality conditions of Assumption 2.4 (two positive expectations involving U^2 and the events).
    Ensures the p=1 problem is not degenerate and the threshold phi(t) is coercive; otherwise the M-estimator would fail to exist for all p >= 1 and the phase transition would be trivial (Section 2).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Phase transitions for the existence of unregularized M-estimators in single index models." pith.science (2026). https://pith.science/paper/NJSPULEZ

@misc{pith2026250103163,
  author       = {Pith},
  title        = {Pith review of: Phase transitions for the existence of unregularized M-estimators in single index models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NJSPULEZ}},
  note         = {Machine review of arXiv:2501.03163}
}
abstract

This paper studies phase transitions for the existence of unregularized M-estimators under proportional asymptotics where the sample size $n$ and feature dimension $p$ grow proportionally with $n/p \to \delta \in (1, \infty)$. We study the existence of M-estimators in single-index models where the response $y_i$ depends on covariates $x_i \sim N(0, I_p)$ through an unknown index ${w} \in \mathbb{R}^p$ and an unknown link function. An explicit expression is derived for the critical threshold $\delta_\infty$ that determines the phase transition for the existence of the M-estimator, generalizing the results of Cand\'es & Sur (2020) for binary logistic regression to other single-index models. Furthermore, we investigate the existence of a solution to the nonlinear system of equations governing the asymptotic behavior of the M-estimator when it exists. The existence of solution to this system for $\delta > \delta_\infty$ remains largely unproven outside the global null in binary logistic regression. We address this gap with a proof that the system admits a solution if and only if $\delta > \delta_\infty$, providing a comprehensive theoretical foundation for proportional asymptotic results that require as a prerequisite the existence of a solution to the system.

Figures

Figures reproduced from arXiv: 2501.03163 by the authors.

Figure 1
Figure 1. Three examples of loss functions. where G ∼ N(0, 1) is independent of (U, Y ), and (U, Y ) has the same distribution as (x ⊤ i w, yi). In particular U ∼ N(0, 1). The phase transition result of Candes & Sur ` (2020) for Gaussian design and binary logistic regression has been extended by Tang & Ye (2020) to elliptic covariate distributions and general binary response models. Han & Ren (2022) extended the phase transit… view at source ↗
Figure 2
Figure 2. Count of instances where the minimizer in (1) exists for varying p/n and signal strength. Simulation parameter: n = 1500, 20 repetitions, ℓy(u) = e u−yu is the Poisson loss, yi | xi satisfies the Poisson model (8). Our first result is that the threshold δ∞ characterizes the phase transition regarding whether the M-estimator exists or not with high-probability under the proportional regime n/p → δ. Theorem 2.6. As n,… view at source ↗
Figure 3
Figure 3. Count of instances where the minimizer in (1) exists for varying p/n and signal strength κ. Simulation parameter: n = 1000, 20 repetitions, yi | xi ∼ satisfies the binomial model Binomial(q, pi) as in (9). 4. Numerical Simulation We generate the covariates (xi) n i=1 iid∼ N(0p, Ip) and re￾sponses yi | xi according to the Poisson model ∀k ∈ N, P  yi = k | xi  = λ k i k! exp(−λi), (8) where λi = exp(−κe ⊤ 1 xi). Her… view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A High-Dimensional Statistical Theory for Convex and Nonconvex Matrix Sensing

    math.ST 2025-06 conditional novelty 8.0 of 10

    In Gaussian matrix sensing, nonconvex factorized least squares is asymptotically equivalent to matrix hard thresholding, while convex nuclear-norm regularization behaves like soft thresholding, making nonconvex no wor...

Reference graph

Works this paper leans on

23 extracted references · 17 canonical work pages · cited by 1 Pith paper

  1. [1]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...

  2. [2]

    B., and Tropp, J

    Amelunxen, D., Lotz, M., McCoy, M. B., and Tropp, J. A. Living on the edge: Phase transitions in convex programs with random data. Information and Inference: A Journal of the IMA, 3 0 (3): 0 224--294, 2014

  3. [3]

    Bauschke, H. H. and Combettes, P. L. Convex Analysis and Monotone Operator Theory in Hilbert Spaces. Springer, 2017

  4. [4]

    and Montanari, A

    Bayati, M. and Montanari, A. The lasso risk for gaussian matrices. IEEE Transactions on Information Theory, 58 0 (4): 0 1997--2017, 2011

  5. [5]

    Bellec, P. C. and Koriyama, T. Existence of solutions to the nonlinear equations characterizing the precise error of m-estimators. arXiv preprint arXiv:2312.13254, 2023

  6. [6]

    Concentration inequalities: A nonasymptotic theory of independence

    Boucheron, S., Lugosi, G., and Massart, P. Concentration inequalities: A nonasymptotic theory of independence. Oxford university press, 2013

  7. [7]

    Cand \`e s, E. J. and Sur, P. The phase transition for the existence of the maximum likelihood estimate in high-dimensional logistic regression. The Annals of Statistics, 48 0 (1): 0 27--42, 2020

  8. [8]

    The lasso with general gaussian designs with applications to hypothesis testing

    Celentano, M., Montanari, A., and Wei, Y. The lasso with general gaussian designs with applications to hypothesis testing. The Annals of Statistics, 51 0 (5): 0 2194--2220, 2023

Show all 23 references
  1. [9]

    Cover, T. M. Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition. IEEE Transactions on Electronic Computers, EC-14 0 (3): 0 326--334, 1965

  2. [10]

    and Derrida, B

    Gardner, E. and Derrida, B. Optimal storage properties of neural network models. Journal of Physics A: Mathematical and general, 21 0 (1): 0 271, 1988

  3. [11]

    Generalisation error in learning with random features and the hidden manifold model

    Gerace, F., Loureiro, B., Krzakala, F., M \'e zard, M., and Zdeborov \'a , L. Generalisation error in learning with random features and the hidden manifold model. In International Conference on Machine Learning, pp.\ 3452--3462. PMLR, 2020

  4. [12]

    and Ren, H

    Han, Q. and Ren, H. Gaussian random projections of convex cones: approximate kinematic formulae and applications. arXiv preprint arXiv:2212.05545, 2022

  5. [13]

    and M \'e zard, M

    Krauth, W. and M \'e zard, M. Storage capacity of memory networks with binary couplings. Journal de Physique, 50 0 (20): 0 3057--3066, 1989

  6. [14]

    and Sur, P

    Liang, T. and Sur, P. A precise high-dimensional asymptotic theory for boosting and minimum- _1 -norm interpolated classifiers. The Annals of Statistics, 50 0 (3): 0 1669--1695, 2022

  7. [15]

    Learning curves of generic features maps for realistic datasets with a teacher-student model

    Loureiro, B., Gerbelot, C., Cui, H., Goldt, S., Krzakala, F., Mezard, M., and Zdeborov \'a , L. Learning curves of generic features maps for realistic datasets with a teacher-student model. Advances in Neural Information Processing Systems, 34: 0 18137--18151, 2021

  8. [16]

    The role of regularization in classification of high-dimensional noisy gaussian mixture

    Mignacco, F., Krzakala, F., Lu, Y., Urbani, P., and Zdeborova, L. The role of regularization in classification of high-dimensional noisy gaussian mixture. In International conference on machine learning, pp.\ 6874--6883. PMLR, 2020

  9. [17]

    and Montanari, A

    Miolane, L. and Montanari, A. The distribution of the lasso: Uniform control over sparse balls and adaptive parameter tuning. The Annals of Statistics, 49 0 (4): 0 2313--2335, 2021

  10. [18]

    The generalization error of max-margin linear classifiers: High-dimensional asymptotics in the overparametrized regime

    Montanari, A., Ruan, F., Sohn, Y., and Yan, J. The generalization error of max-margin linear classifiers: High-dimensional asymptotics in the overparametrized regime. arXiv preprint arXiv:1911.01544, 2023

  11. [19]

    The impact of regularization on high-dimensional logistic regression

    Salehi, F., Abbasi, E., and Hassibi, B. The impact of regularization on high-dimensional logistic regression. Advances in Neural Information Processing Systems, 32, 2019

  12. [20]

    and Cand \`e s, E

    Sur, P. and Cand \`e s, E. J. A modern maximum-likelihood theory for high-dimensional logistic regression. Proceedings of the National Academy of Sciences, 116 0 (29): 0 14516--14525, 2019

  13. [21]

    Sur, P., Chen, Y., and Cand \`e s, E. J. The likelihood ratio test in high-dimensional logistic regression is asymptotically a rescaled chi-square. Probability theory and related fields, 175: 0 487--558, 2019

  14. [22]

    and Ye, Y

    Tang, W. and Ye, Y. The existence of maximum likelihood estimate in high-dimensional binary response generalized linear models. Electronic Journal of Statistics, 2020

  15. [23]

    Precise error analysis of regularized m -estimators in high dimensions

    Thrampoulidis, C., Abbasi, E., and Hassibi, B. Precise error analysis of regularized m -estimators in high dimensions. IEEE Transactions on Information Theory, 64 0 (8): 0 5592--5628, 2018

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.