Pith. sign in

REVIEW 5 cited by

Saddle-to-Saddle Dynamics in Deep Linear Networks: Small Initialization Training, Symmetry, and Sparsity

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.15933 v2 pith:FHXPVUW7 submitted 2021-06-30 stat.ML cs.LG

classification stat.MLcs.LG
keywords gammadynamicsvariancecasegloballinearminimumtheta
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

The dynamics of Deep Linear Networks (DLNs) is dramatically affected by the variance $\sigma^2$ of the parameters at initialization $\theta_0$. For DLNs of width $w$, we show a phase transition w.r.t. the scaling $\gamma$ of the variance $\sigma^2=w^{-\gamma}$ as $w\to\infty$: for large variance ($\gamma<1$), $\theta_0$ is very close to a global minimum but far from any saddle point, and for small variance ($\gamma>1$), $\theta_0$ is close to a saddle point and far from any global minimum. While the first case corresponds to the well-studied NTK regime, the second case is less understood. This motivates the study of the case $\gamma \to +\infty$, where we conjecture a Saddle-to-Saddle dynamics: throughout training, gradient descent visits the neighborhoods of a sequence of saddles, each corresponding to linear maps of increasing rank, until reaching a sparse global minimum. We support this conjecture with a theorem for the dynamics between the first two saddles, as well as some numerical experiments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Singular perturbations and hierarchical learning in two-layer neural networks

    cs.LG 2026-07 accept novelty 7.0 of 10

    Constant and linear Hermite components of a misspecified single-index target are recovered at the conjectured singular-perturbation timescales; quadratic learning remains coupled to them via an auxiliary constrained flow.

  2. Adaptive kernel predictors from feature-learning infinite limits of neural networks

    cs.LG 2025-02 conditional novelty 7.0 of 10

    Feature-learning infinite-width neural networks are kernel machines with data-dependent kernels, defined by a min-max saddle point (Bayesian/Langevin) or a DMFT fixed point (gradient flow with weight decay).

  3. Over-Alignment vs Over-Fitting: The Role of Feature Learning Strength in Generalization

    cs.LG 2026-01 conditional novelty 6.0 of 10

    Training neural networks with an intermediate feature learning strength generalizes best; the paper derives this from a trade-off between over-alignment to the empirical class mean and over-fitting from a large hypoth...

  4. Gradient Flow Equations for Deep Linear Neural Networks: A Survey from a Network Perspective

    cs.LG 2025-11 conditional novelty 6.0 of 10

    The adjacency-matrix reformulation of deep-linear gradient flow reveals a quotient-space structure of the loss landscape, but the proof that arcs between critical points determine stable and unstable manifolds is incomplete.

  5. Optimal Learning Rate Scaling Depends on Data in Deep Scalar Linear Networks

    cs.LG 2026-07 accept novelty 5.0 of 10

    Optimal depth-wise learning-rate scaling in deep scalar linear networks is data-dependent, so data-agnostic rules fail to transfer while the data-aware rule yields depth-independent linear convergence.

Pith tools