REVIEW 6 minor 1 cited by
Phase transitions for the existence of unregularized M-estimators in single index models
T0 review · 0 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read This paper proves that one scalar threshold, $\delta_\infty$, simultaneously controls whether an unregularized M-estimator exists and whether the asymptotic system characterizing it has a unique solution, across smooth single-index models.
desk verdict A solid, important paper that closes a real proof gap in high-dimensional M-estimator theory; the C^1 caveat is explicit and does not sink the central theorem. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the infinite-dimensional convex program $\min_{(a,v)\in\mathbb{R}\times H} \mathbb{E}[\ell_Y(aU + v)]$ subject to $\|v\| - \mathbb{E}[vG]/\sqrt{1-\delta^{-1}} \le 0$, where $H$ is the Hilbert space of square-integrable functions of $(G,U,Y)$ and $G$ is an independent standard normal. Theorem 3.1 shows that minimizers of this program correspond one-to-one with solutions of the nonlinear system whenever the Lagrange multiplier is positive, so the phase transition for the system reduces to the existence of this minimizer. The threshold $\delta_\infty$ enters through the cone $C$ of directions along which the loss can only be constant or decreasing; the variational representation $\delta_\infty^{-1} = \inf_{(t,p)\in C} \mathbb{E}[(G-p)^2]$ yields the direction $p^*$ used to show that no minimizer exists for $\delta \le \delta_\infty$, while the same $p^*$ gives coercivity for $\delta > \delta_\infty$. The smoothness of the loss is used in Lemma 3.2 to rule out the degenerate minimizer $v^* = 0$.
What would settle it
Choose the Poisson model with a nonzero signal, compute $\delta_\infty$ by minimizing $\varphi$ from (4), then numerically solve the system (2) over a grid of $\delta$ values; any solution found for $\delta < \delta_\infty$, or no solution found for $\delta > \delta_\infty$, would refute Theorem 2.7.
Extended reading notes
Core claim
For $n/p \to \delta$, define $\delta_\infty$ by $1/\delta_\infty = \inf_{t\in\mathbb{R}} \varphi(t)$, where $\varphi$ is a convex function built from the events where the loss is coercive, increasing, or decreasing. Theorem 2.6 states that the unregularized M-estimator exists with probability tending to one when $\delta > \delta_\infty$ and with probability tending to zero when $\delta < \delta_\infty$. Theorem 2.7 states that the nonlinear system of equations characterizing the asymptotic behavior of the estimator has no solution when $\delta \le \delta_\infty$ and a unique solution when $\delta > \delta_\infty$. The two theorems are proved through an infinite-dimensional convex optimization problem whose minimizer is equivalent to a solution of the system, with the technical core being a non-degeneracy lemma showing that the minimizer does not collapse to $v^* = 0$.
Load-bearing premise
The if-and-only-if result depends on every loss being $C^1$ and strictly convex; the paper itself notes that non-differentiable losses can introduce a different threshold, so if smoothness fails the central claim is not generic.
Editorial extensions
If this is right
- For Poisson and binomial regression with Gaussian covariates and smooth losses, the threshold $\delta_\infty$ from (4) rigorously separates existence from nonexistence of the unregularized M-estimator.
- Any asymptotic analysis of these M-estimators based on the Convex Gaussian Minmax Theorem can now invoke Theorem 2.7 to justify the required unique solution of the system whenever $\delta > \delta_\infty$.
- The global-null logistic result $\delta_\infty = 2$ is contained as a special case, and the threshold is available for non-null single-index models.
- In the non-existence regime $\delta \le \delta_\infty$, the proof exhibits an explicit ray along which the objective decreases, giving a constructive certificate that no minimizer exists.
Reading between the lines
- The authors do not pursue non-Gaussian designs, but the cone-projection form of $\delta_\infty$ suggests an analogous variational formula for elliptical covariate distributions.
- A consequence the authors leave implicit is that for smooth losses, existence of the estimator and solvability of its asymptotic system are the same event, which may be a general structural fact rather than a special property of the examples treated.
- The remark on $\delta_{\text{perfect}}$ invites a concrete follow-up: for nonsmooth losses, test whether the phase transitions for the estimator and for the nonlinear system occur at different thresholds.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies unregularized M-estimators in single-index models with Gaussian covariates under proportional asymptotics n/p → δ ∈ (1, ∞). It defines a threshold δ∞ as the infimum of an explicit convex functional φ(t) of the data distribution and proves two main results. Theorem 2.6 establishes a phase transition for the existence of the M-estimator at δ∞, generalizing Candès & Sur (2020) beyond binary logistic regression. Theorem 2.7 proves that the asymptotic nonlinear system (2) of Sur & Candès has a unique solution if and only if δ > δ∞, closing a gap noted in the literature. The proof proceeds through an equivalence (Theorem 3.1) between system (2) and an infinite-dimensional convex program (5), a non-degeneracy lemma (Lemma 3.2), a uniqueness lemma (Lemma 3.3), and a phase-transition theorem for the convex program (Theorem 3.4).
Significance. If correct, the results provide a comprehensive theoretical foundation for proportional-asymptotic analyses of M-estimators that assume existence of a solution to the asymptotic system. The threshold δ∞ is given by an explicit, parameter-free convex program, and both theorems are proved with complete arguments in the appendices. The main results are numerically verified for Poisson and binomial losses. The paper is careful about its scope: the smoothness assumption on the loss (Assumption 2.5(1)) is stated up front, and the authors explicitly note that the non-degeneracy conclusion changes for non-differentiable losses (Bellec & Koriyama, 2023). The central claim is internally consistent under the stated assumptions.
minor comments (6)
- [Section 3, paragraph after Theorem 2.6] The sentence referring to 'the nonlinear system (4)' should refer to 'the nonlinear system (2)'.
- [Appendix E, Lemma E.2] The displayed bound '∥v − ṽ∥ ≤ C^(2)(ξ)(1 + ∥v∥²)' should use a linear term '1 + ∥v∥' rather than '1 + ∥v∥²'. The subsequent inequalities (28)–(29) require the linear version, and the proof's earlier bound on |a| is indeed linear in ∥v∥.
- [Appendix A, Lemma A.1 proof] The computation of E[p_*²] contains a typographical clutter with repeated terms 'E[p_*²] + 2t_* E[p_*U] E[p_*U] = 0'; this should be simplified to a clean derivation.
- [Section 4, Numerical Simulation] The phrase 'using linear programming' for the Poisson loss is potentially confusing because the loss is exponential; presumably the authors solve the linear feasibility problem (12) to detect existence, and this should be stated explicitly.
- [Section 4, Figures 2 and 3] The figures would be more informative if they included error bars or standard errors over the 20 repetitions.
- [Theorem 2.6] The theorem does not address the boundary case δ = δ∞; a sentence noting that this case is not covered (and why) would improve precision.
Circularity Check
No significant circularity: the phase-transition threshold is an independently defined variational quantity and the system-solvability iff is proved, not fitted.
full rationale
The central claim, Theorem 2.7, is not circular. The threshold delta_infinity is defined in (4) by an explicit variational infimum over a convex function phi(t) built from the loss and the distribution of (U, Y); no unknown constant is fitted to the nonlinear system (2). Theorem 2.6 derives the M-estimator existence phase transition from this same threshold using Lemma A.2 and the Gaussian kinematic formula, and the proof shows the empirical distance to the relevant cone converges to 1/delta_infinity by the law of large numbers. Theorem 2.7 is proved through an independent chain: Theorem 3.1 proves a bijective equivalence between solutions of system (2) and minimizers of the infinite-dimensional program (5) with v* != 0; Lemma 3.2 proves non-degeneracy; Lemma 3.3 proves uniqueness; and Theorem 3.4 proves that (5) has a minimizer if and only if delta > delta_infinity. The two directions of Theorem 3.4 use Lemma A.1's representation delta_infinity^{-1} = inf_{(t,p) in C} E[(G-p)^2] and direct coercivity arguments; the threshold is not chosen to make either direction true. The paper's references to the authors' own prior work, e.g. 'The notation and setup of this infinite-dimensional convex optimization problem is heavily inspired by Bellec & Koriyama (2023)' and 'The argument in this proof is inspired by the proof of Lemma 2.6 in (Bellec & Koriyama, 2023)', are methodological or technical, not load-bearing for the main theorem; the core equivalence and phase-transition proofs are contained in the paper. The explicit acknowledgement that non-differentiable losses can produce a different threshold delta_perfect (Bellec & Koriyama, 2023) is a scope restriction stated up front, not a hidden circular step. No prediction reduces to a fitted value, and no theorem is imported from a self-citation as the sole justification for the central claim.
Assumptions & free parameters
assumptions (6)
- domain assumption Gaussian design and single-index model: x_i iid N(0, I_p), y_i depends on x_i only through x_i^T w with ||w||=1; (U,Y) = (x_i^T w, y_i) with U ~ N(0,1).
- domain assumption Loss differentiability and strict convexity: l_y is C^1 and strictly convex for every y, and P(Omega_v) < 1 (Assumptions 2.4 and 2.5).
- domain assumption Affine one-sided lower bound on the loss with square-integrable intercept D(Y): l_Y(u) >= -D(Y) + (1/b) times (u, |u|, or -u depending on the event) (Assumption 2.5(4)).
- standard math Gaussian kinematic formula and statistical dimension theory of Amelunxen et al. (2014).
- standard math Convex analysis in Hilbert spaces: Slater's condition, KKT conditions, coercivity, and the relevant propositions from Bauschke and Combettes (2017).
- domain assumption Non-triviality conditions of Assumption 2.4 (two positive expectations involving U^2 and the events).
Cite this review
Pith. "Pith review of Phase transitions for the existence of unregularized M-estimators in single index models." pith.science (2026). https://pith.science/paper/NJSPULEZ
@misc{pith2026250103163,
author = {Pith},
title = {Pith review of: Phase transitions for the existence of unregularized M-estimators in single index models},
year = {2026},
howpublished = {\url{https://pith.science/paper/NJSPULEZ}},
note = {Machine review of arXiv:2501.03163}
}
abstract
This paper studies phase transitions for the existence of unregularized M-estimators under proportional asymptotics where the sample size $n$ and feature dimension $p$ grow proportionally with $n/p \to \delta \in (1, \infty)$. We study the existence of M-estimators in single-index models where the response $y_i$ depends on covariates $x_i \sim N(0, I_p)$ through an unknown index ${w} \in \mathbb{R}^p$ and an unknown link function. An explicit expression is derived for the critical threshold $\delta_\infty$ that determines the phase transition for the existence of the M-estimator, generalizing the results of Cand\'es & Sur (2020) for binary logistic regression to other single-index models. Furthermore, we investigate the existence of a solution to the nonlinear system of equations governing the asymptotic behavior of the M-estimator when it exists. The existence of solution to this system for $\delta > \delta_\infty$ remains largely unproven outside the global null in binary logistic regression. We address this gap with a proof that the system admits a solution if and only if $\delta > \delta_\infty$, providing a comprehensive theoretical foundation for proportional asymptotic results that require as a prerequisite the existence of a solution to the system.
Figures
Forward citations
Cited by 1 Pith paper
-
A High-Dimensional Statistical Theory for Convex and Nonconvex Matrix Sensing
In Gaussian matrix sensing, nonconvex factorized least squares is asymptotically equivalent to matrix hard thresholding, while convex nuclear-norm regularization behaves like soft thresholding, making nonconvex no wor...
Reference graph
Works this paper leans on
-
[1]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...
-
[2]
Amelunxen, D., Lotz, M., McCoy, M. B., and Tropp, J. A. Living on the edge: Phase transitions in convex programs with random data. Information and Inference: A Journal of the IMA, 3 0 (3): 0 224--294, 2014
work page 2014
-
[3]
Bauschke, H. H. and Combettes, P. L. Convex Analysis and Monotone Operator Theory in Hilbert Spaces. Springer, 2017
work page 2017
-
[4]
Bayati, M. and Montanari, A. The lasso risk for gaussian matrices. IEEE Transactions on Information Theory, 58 0 (4): 0 1997--2017, 2011
work page 1997
-
[5]
Bellec, P. C. and Koriyama, T. Existence of solutions to the nonlinear equations characterizing the precise error of m-estimators. arXiv preprint arXiv:2312.13254, 2023
arXiv 2023
-
[6]
Concentration inequalities: A nonasymptotic theory of independence
Boucheron, S., Lugosi, G., and Massart, P. Concentration inequalities: A nonasymptotic theory of independence. Oxford university press, 2013
work page 2013
-
[7]
Cand \`e s, E. J. and Sur, P. The phase transition for the existence of the maximum likelihood estimate in high-dimensional logistic regression. The Annals of Statistics, 48 0 (1): 0 27--42, 2020
work page 2020
-
[8]
The lasso with general gaussian designs with applications to hypothesis testing
Celentano, M., Montanari, A., and Wei, Y. The lasso with general gaussian designs with applications to hypothesis testing. The Annals of Statistics, 51 0 (5): 0 2194--2220, 2023
work page 2023
Show all 23 references
-
[9]
Cover, T. M. Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition. IEEE Transactions on Electronic Computers, EC-14 0 (3): 0 326--334, 1965
1965
-
[10]
and Derrida, B
Gardner, E. and Derrida, B. Optimal storage properties of neural network models. Journal of Physics A: Mathematical and general, 21 0 (1): 0 271, 1988
1988
-
[11]
Generalisation error in learning with random features and the hidden manifold model
Gerace, F., Loureiro, B., Krzakala, F., M \'e zard, M., and Zdeborov \'a , L. Generalisation error in learning with random features and the hidden manifold model. In International Conference on Machine Learning, pp.\ 3452--3462. PMLR, 2020
2020
-
[12]
and Ren, H
Han, Q. and Ren, H. Gaussian random projections of convex cones: approximate kinematic formulae and applications. arXiv preprint arXiv:2212.05545, 2022
2022
-
[13]
and M \'e zard, M
Krauth, W. and M \'e zard, M. Storage capacity of memory networks with binary couplings. Journal de Physique, 50 0 (20): 0 3057--3066, 1989
1989
-
[14]
and Sur, P
Liang, T. and Sur, P. A precise high-dimensional asymptotic theory for boosting and minimum- _1 -norm interpolated classifiers. The Annals of Statistics, 50 0 (3): 0 1669--1695, 2022
2022
-
[15]
Learning curves of generic features maps for realistic datasets with a teacher-student model
Loureiro, B., Gerbelot, C., Cui, H., Goldt, S., Krzakala, F., Mezard, M., and Zdeborov \'a , L. Learning curves of generic features maps for realistic datasets with a teacher-student model. Advances in Neural Information Processing Systems, 34: 0 18137--18151, 2021
2021
-
[16]
The role of regularization in classification of high-dimensional noisy gaussian mixture
Mignacco, F., Krzakala, F., Lu, Y., Urbani, P., and Zdeborova, L. The role of regularization in classification of high-dimensional noisy gaussian mixture. In International conference on machine learning, pp.\ 6874--6883. PMLR, 2020
2020
-
[17]
and Montanari, A
Miolane, L. and Montanari, A. The distribution of the lasso: Uniform control over sparse balls and adaptive parameter tuning. The Annals of Statistics, 49 0 (4): 0 2313--2335, 2021
2021
-
[18]
The generalization error of max-margin linear classifiers: High-dimensional asymptotics in the overparametrized regime
Montanari, A., Ruan, F., Sohn, Y., and Yan, J. The generalization error of max-margin linear classifiers: High-dimensional asymptotics in the overparametrized regime. arXiv preprint arXiv:1911.01544, 2023
1911 arXiv
-
[19]
The impact of regularization on high-dimensional logistic regression
Salehi, F., Abbasi, E., and Hassibi, B. The impact of regularization on high-dimensional logistic regression. Advances in Neural Information Processing Systems, 32, 2019
2019
-
[20]
and Cand \`e s, E
Sur, P. and Cand \`e s, E. J. A modern maximum-likelihood theory for high-dimensional logistic regression. Proceedings of the National Academy of Sciences, 116 0 (29): 0 14516--14525, 2019
2019
-
[21]
Sur, P., Chen, Y., and Cand \`e s, E. J. The likelihood ratio test in high-dimensional logistic regression is asymptotically a rescaled chi-square. Probability theory and related fields, 175: 0 487--558, 2019
2019
-
[22]
and Ye, Y
Tang, W. and Ye, Y. The existence of maximum likelihood estimate in high-dimensional binary response generalized linear models. Electronic Journal of Statistics, 2020
2020
-
[23]
Precise error analysis of regularized m -estimators in high dimensions
Thrampoulidis, C., Abbasi, E., and Hassibi, B. Precise error analysis of regularized m -estimators in high dimensions. IEEE Transactions on Information Theory, 64 0 (8): 0 5592--5628, 2018
2018
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.