REVIEW 5 cited by
Saddle-to-Saddle Dynamics in Deep Linear Networks: Small Initialization Training, Symmetry, and Sparsity
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
The dynamics of Deep Linear Networks (DLNs) is dramatically affected by the variance $\sigma^2$ of the parameters at initialization $\theta_0$. For DLNs of width $w$, we show a phase transition w.r.t. the scaling $\gamma$ of the variance $\sigma^2=w^{-\gamma}$ as $w\to\infty$: for large variance ($\gamma<1$), $\theta_0$ is very close to a global minimum but far from any saddle point, and for small variance ($\gamma>1$), $\theta_0$ is close to a saddle point and far from any global minimum. While the first case corresponds to the well-studied NTK regime, the second case is less understood. This motivates the study of the case $\gamma \to +\infty$, where we conjecture a Saddle-to-Saddle dynamics: throughout training, gradient descent visits the neighborhoods of a sequence of saddles, each corresponding to linear maps of increasing rank, until reaching a sparse global minimum. We support this conjecture with a theorem for the dynamics between the first two saddles, as well as some numerical experiments.
Forward citations
Cited by 5 Pith papers
-
Singular perturbations and hierarchical learning in two-layer neural networks
Constant and linear Hermite components of a misspecified single-index target are recovered at the conjectured singular-perturbation timescales; quadratic learning remains coupled to them via an auxiliary constrained flow.
-
Adaptive kernel predictors from feature-learning infinite limits of neural networks
Feature-learning infinite-width neural networks are kernel machines with data-dependent kernels, defined by a min-max saddle point (Bayesian/Langevin) or a DMFT fixed point (gradient flow with weight decay).
-
Over-Alignment vs Over-Fitting: The Role of Feature Learning Strength in Generalization
Training neural networks with an intermediate feature learning strength generalizes best; the paper derives this from a trade-off between over-alignment to the empirical class mean and over-fitting from a large hypoth...
-
Gradient Flow Equations for Deep Linear Neural Networks: A Survey from a Network Perspective
The adjacency-matrix reformulation of deep-linear gradient flow reveals a quotient-space structure of the loss landscape, but the proof that arcs between critical points determine stable and unstable manifolds is incomplete.
-
Optimal Learning Rate Scaling Depends on Data in Deep Scalar Linear Networks
Optimal depth-wise learning-rate scaling in deep scalar linear networks is data-dependent, so data-agnostic rules fail to transfer while the data-aware rule yields depth-independent linear convergence.
Discussion (0). Continue with ORCID to comment.