Pith. sign in

REVIEW 2 cited by

Escaping mediocrity: how two-layer networks learn hard generalized linear models with SGD

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.18502 v2 pith:3JNLGQBZ submitted 2023-05-29 stat.ML cs.LG

classification stat.MLcs.LG
keywords escapinggeneralizedlearnlinearmediocritynetworksprocessscenario
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

This study explores the sample complexity for two-layer neural networks to learn a generalized linear target function under Stochastic Gradient Descent (SGD), focusing on the challenging regime where many flat directions are present at initialization. It is well-established that in this scenario $n=O(d \log d)$ samples are typically needed. However, we provide precise results concerning the pre-factors in high-dimensional contexts and for varying widths. Notably, our findings suggest that overparameterization can only enhance convergence by a constant factor within this problem class. These insights are grounded in the reduction of SGD dynamics to a stochastic process in lower dimensions, where escaping mediocrity equates to calculating an exit time. Yet, we demonstrate that a deterministic approximation of this process adequately represents the escape time, implying that the role of stochasticity may be minimal in this scenario.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning Regime

    cs.LG 2025-09 conditional novelty 6.0 of 10

    For diagonal and quadratic two-layer networks, training maps to LASSO and matrix compressed sensing, yielding a full phase diagram of excess-risk scaling exponents and a spectral characterization of the trained weights.

  2. Optimal Spectral Transitions in High-Dimensional Multi-Index Models

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Two linearized message-passing spectral estimators achieve the optimal weak-recovery threshold in Gaussian multi-index models, with a sharp BBP-like spectral phase transition at the critical sample complexity.

Pith tools