Pith. sign in

REVIEW 10 cited by

The large learning rate phase of deep learning: the catapult mechanism

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.02218 v1 pith:XPHPVTNM submitted 2020-03-04 stat.ML cs.LG

classification stat.MLcs.LG
keywords learninglargeratedeepdynamicsnetworksphaserates
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The choice of initial learning rate can have a profound effect on the performance of deep networks. We present a class of neural networks with solvable training dynamics, and confirm their predictions empirically in practical deep learning settings. The networks exhibit sharply distinct behaviors at small and large learning rates. The two regimes are separated by a phase transition. In the small learning rate phase, training can be understood using the existing theory of infinitely wide neural networks. At large learning rates the model captures qualitatively distinct phenomena, including the convergence of gradient descent dynamics to flatter minima. One key prediction of our model is a narrow range of large, stable learning rates. We find good agreement between our model's predictions and training dynamics in realistic deep learning settings. Furthermore, we find that the optimal performance in such settings is often found in the large learning rate phase. We believe our results shed light on characteristics of models trained at different learning rates. In particular, they fill a gap between existing wide neural network theory, and the nonlinear, large learning rate, training dynamics relevant to practice.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 60 citations worldwide. Full citation record

  1. A Defense of the Quadratic Model

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Local Taylor-expanded quadratic models reproduce a 150M-parameter LLM's validation loss for up to 10% of training late in the run, and LLM pretraining operates within a factor of 2 of a stochastic or deterministic edg...

  2. The Map Behind the Flow: Finite-Step Gradient Descent as a Dynamical System

    cs.LG 2026-07 accept novelty 7.0 of 10

    The edge of stability is the first bifurcation of the finite-step gradient map; residual oscillations then drive balancing and representation selection beyond that edge.

  3. Implicit Bias of SGD in Multivariate ReLU Networks: Effective Width Collapse

    cs.LG 2026-07 accept novelty 7.0 of 10

    Noisy SGD in the mean-field regime forces wide multivariate ReLU networks to an effective width of at most 2P-1, yielding a continuous piecewise-affine predictor whose hyperplanes are non-redundant with respect to the...

  4. Zeroth-Order Optimization at the Edge of Stability

    cs.LG 2026-04 unverdicted novelty 7.0 of 10

    Zeroth-order methods achieve mean-square stability when the step size satisfies a condition involving the entire Hessian spectrum, with full-batch ZO optimizers operating at the edge of stability and large steps regul...

  5. The Fourth Quadrant: A Stylized View of Benign Misfitting

    cs.LG 2026-08 conditional novelty 6.0 of 10

    In a stylized single-spike linear model, useful span predictors in the window d/gamma^2 << n << d/gamma are forced to overshoot the training labels, so good test error comes together with large training error.

  6. Directional Curvature from Armijo Backtracking: A Low-Cost Sharpness Probe and a Calibration-Free Learning-Rate Safeguard for Adam

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A single Armijo backtracking line search at initialization estimates local Hessian sharpness and caps Adam's learning rate to prevent divergence, with fixed safety factor κ=2 across nine architectures.

  7. (How) Learning Rates Regulate Catastrophic Overtraining

    cs.LG 2026-04 unverdicted novelty 6.0 of 10

    Pretraining LR decay sharpens LLMs, and that sharpness—not just token count—drives catastrophic forgetting during supervised fine-tuning.

  8. Adaptive Preconditioners Trigger Loss Spikes in Adam

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Loss spikes in Adam occur when its second-moment memory decays faster than gradients grow, briefly removing the adaptive brake; a single Hessian-vector product along the gradient direction can flag the onset.

  9. Weight-norm Criticality: A Mechanism for Loss Spikes Induced by the Normalization and Weight Decay

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Weight decay on scale-invariant weights creates a norm-dependent sharpness boundary; crossing it predicts loss spikes in normalized networks.

  10. Right Time to Learn:Promoting Generalization via Bio-inspired Spacing Effect in Knowledge Distillation

    cs.LG 2025-02 conditional novelty 5.0 of 10

    Training the teacher a small number of steps ahead of the student and freezing it during distillation improves student generalization by up to 3.4% on image benchmarks.

Pith tools