Pith. sign in

REVIEW 2 major objections 4 minor 69 references

Dynamics of Gradient Descent with Large Step Size Near a Manifold of Flat Minima

T0 review · 2 major / 4 minor · reviewed 2026-07-10 · grok-4.5

Pith's one-line read Near a manifold of flat minima, large-step gradient descent splits into sharpness descent along the solution set plus a flip bifurcation and contraction in the normal directions, with three distinct convergence regimes.

desk verdict Solid local normal-form extension of large-step GD theory to vector outputs and flat-minima manifolds; critical rate is conditional on a stated parabolic-foliation conjecture except in special spectral cases. read the letter →

arxiv 2607.08380 v1 pith:65VBXAJ2 submitted 2026-07-09 cs.LG math.DSmath.OC

classification cs.LGmath.DSmath.OC
keywords gradientdescentedgeofstabilitysharpnessflatminimanormalformoverparametrisedleastsquaresmatrixfactorisationMorse-Bott
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Classical gradient-descent theory demands a step size smaller than twice the reciprocal of the sharpness, yet deep networks routinely violate that bound and still train. This paper shows that, for overparametrised least-squares problems of any output dimension, the same large-step dynamics remain well-behaved once one is near a manifold of flat minima rather than an isolated flat point. In suitably chosen coordinates the iteration becomes Riemannian gradient descent on the sharpness along the solution manifold, a flip bifurcation in the top Hessian eigendirection, and linear contraction in the remaining normal directions. The relative size of the step size to the flat-manifold sharpness then produces three regimes: exponential approach to a suboptimally flat minimiser, polynomial approach to the flat manifold itself, or exponential approach to a stable period-two orbit. The same geometric hypotheses hold for deep matrix factorisation, where the flat minima form a fibre bundle over a product of spheres and the sharpness is Morse-Bott. The result therefore supplies a local dynamical picture that covers the multi-output, multi-observation setting actually used in practice.

What carries the argument

The C1 normal form of Theorem 4.1 (and its precise version Theorem B.7), obtained by successive tubular, centre-manifold, singular-PDE and strong-stable-foliation coordinate changes; it makes the three-regime analysis possible.

What would settle it

Run large-step gradient descent on a multi-output least-squares problem whose flat minima form a manifold that violates the constant-multiplicity or invariance hypotheses of Assumption D.12 and check whether the critical-regime distance to the manifold still decays as Theta(t^{-1/2}); a clear power different from -1/2 would refute the general claim.

Watch

Extended reading notes

Core claim

Under four geometric assumptions on an overparametrised least-squares loss, large-step gradient descent near a manifold of flat minima admits a C1 normal form that decouples into Riemannian gradient descent on sharpness along the solution manifold, a cubic flip bifurcation in the top eigendirection, and exponential contraction in the remaining normal directions; the three classical step-size regimes relative to twice the reciprocal of the flat-manifold sharpness then yield the three corresponding convergence theorems.

Load-bearing premise

The general critical-regime rate proof needs either a constant multiple of the identity for the Hessian of sharpness normal to the flat manifold, or an invariance assumption plus an unproven parabolic stable-foliation conjecture.

Editorial extensions

If this is right

  • Local large-step training dynamics near flat minima are governed by sharpness descent plus a one-dimensional flip bifurcation, independent of output dimension.
  • Deep matrix-factorisation landscapes possess an explicit fibre-bundle structure of flat minima over products of spheres, with Morse-Bott sharpness.
  • Subcritical, critical and supercritical step-size choices produce, respectively, exponential suboptimal-flat convergence, t^{-1/2} approach to the flat manifold, and exponential approach to a period-two orbit.
  • The same normal-form programme can in principle be applied to other overparametrised models once the four geometric assumptions are verified.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the normal form is local to the flat manifold, progressive sharpening and the bulk of the edge-of-stability phase remain outside its present scope and will need a global geometric theory.
  • The fibre-bundle description of flat minima in matrix factorisation suggests that not all flat points are equivalent for generalisation; distinguished sub-bundles may be statistically preferred.
  • The novel singular-PDE technique used to straighten the centre manifold may be reusable for other resonant bifurcations that occur along submanifolds rather than isolated points.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper extends the large-step-size GD analysis of MacDonald et al. (2025) from codimension-1 least squares with an isolated flat minimum to arbitrary-codimension overparametrised least squares near a manifold F of flat minima. Under geometric Assumptions 3.1–3.5 it constructs a C^{1} normal form (Theorem 4.1 / B.7) in which the dynamics split into Riemannian GD on sharpness along the solution manifold M (step size controlled by y_{1}^{2}), a flip bifurcation in the top Hessian eigendirection y_{1}, and linear contraction in the remaining normal directions y_{2:q}. From this form it proves three local convergence regimes: subcritical exponential convergence to a suboptimal flat minimiser (Theorem 5.1), critical Θ(t^{-1/2}) convergence to F (Theorem 5.2, under extra spectral/invariance hypotheses or Conjecture D.7), and supercritical exponential convergence to a period-2 orbit of amplitude Θ((ηλ_{1}|F−2)^{1/2}) along span(ν_{1}|F) (Theorem 5.3). The framework is verified for deep matrix factorisation, where F is shown to be a fibre bundle over a product of spheres and λ_{1} is Morse-Bott along F (Appendix E).

Significance. The work supplies a rigorous dynamical-systems account of large-step GD near manifolds of flat minima for vector-valued least squares, substantially enlarging the scope of the earlier codimension-1 theory. The normal form, the centre-manifold reduction, and especially the novel solution of the singular PDE (Theorem C.4) are technically substantial and may be of independent interest. The matrix-factorisation structural results (fibre-bundle description of F, Morse-Bott property of sharpness) are new and concrete. Subcritical and supercritical theorems are fully proved from the normal form; numerical experiments on 2 imes2 three-layer factorisation are consistent with all three regimes. The main limitation is that the general critical-rate claim remains conditional on an unproven parabolic foliation conjecture (or a restrictive spectral hypothesis).

major comments (2)
  1. Theorem 5.2 / D.13: the claimed Θ(t^{-1/2}) rate for general F is unconditional only when ∇^{2}_M λ_{1}|ν_M F is a constant multiple of the identity. Otherwise the proof reduces the system to the abstract form (D.206)–(D.214) and then invokes both Assumption D.12 (constant smallest eigenvalue of constant multiplicity, local invariance of the bottom-eigenspace–ν_{1} submanifold, z-independent (u,y_{1}) updates) and the unproven Conjecture D.7 (Lipschitz strong-stable foliation for the normally parabolic system). Assumption D.12 is verified only for two-layer matrix factorisation (Proposition E.7); the foliation conjecture is left open. The manuscript should either prove Conjecture D.7, restrict the critical theorem to the cases already covered, or state the critical claim more carefully as conditional.
  2. Section 6 and the abstract present the three convergence theorems as parallel generalisations of [39]. Because the critical regime is not fully proved for general manifolds of flat minima, the packaging overstates the unconditional reach of the theory. A clearer separation between the fully rigorous subcritical/supercritical results and the conditional critical result would better match the proofs.
minor comments (4)
  1. Assumption 3.5 (positivity and vanishing derivative of α_η on F) is stated abstractly; a short geometric interpretation of why these conditions are natural for least-squares models would help the reader.
  2. Figures 3–5 show only 2 imes2 three-layer factorisation. A brief remark on whether the same qualitative behaviour is expected for larger dimensions or deeper factorisations would strengthen the experimental section.
  3. The centre-manifold preprint [38] is cited as arXiv:2604.18202; ensure the reference remains accessible or supply the needed statement inline if the preprint is not yet public.
  4. Notation for the normal-form coordinates (x,y_{1},y_{2:q}) and the successive maps Φ_{1}–Φ_{4} is dense; a short schematic of the coordinate changes would improve readability of Section 4 / Appendix B.

Circularity Check

1 steps flagged · score 1.5 of 10

No significant circularity: normal form and rates are derived from geometric assumptions plus a novel singular-PDE solution; self-citations to [39] and the centre-manifold preprint [38] are foundational tools, not definitional loops that force the claimed rates.

  1. self citation load bearing [Section 4 / Lemma B.3 and the paragraph preceding Thm 4.1]
    "a dimension reduction must be undertaken using a centre manifold theorem [38] to reduce the effective dimension of the problem from q > 1 to q = 1 … The lemma is an immediate application of [38, Theorem 1.2] to the problem at hand."

    The reduction of the orthogonal dynamics to a one-dimensional centre manifold (essential for the subsequent normal-form PDE and all three convergence theorems) rests on a centre-manifold theorem for maps along manifolds of fixed points that appears only as the author’s own concurrent preprint [38]. While the paper still performs substantial independent work (singular PDE, higher-order estimates, matrix-factorisation geometry), this step is a self-citation that is load-bearing for the normal form; it is not, however, a definitional loop that forces the claimed rates by construction.

full rationale

The paper is a pure dynamical-systems extension of large-step GD analysis. The normal form (Thm 4.1/B.7) is obtained by an explicit sequence of coordinate changes (tubular neighbourhood, centre-manifold reduction via [38], singular PDE solved in Thm C.4 by a new approximate-series + variation-of-constants argument, then strong-stable foliation). Subcritical and supercritical rates follow from that normal form by standard invariant-neighbourhood and contraction arguments (D.2–D.5, D.14–D.15). The critical rate is stated conditionally on either a spectral hypothesis or (Assumption D.12 + open Conjecture D.7); the paper does not claim an unconditional proof and verifies D.12 only for two-layer matrix factorisation. Matrix-factorisation structural results (fibre-bundle structure of F, Morse-Bott property of λ1) are proved from first principles under the stated rank/singular-value hypotheses. There is no data fitting, no parameter estimated from one quantity and re-used as a “prediction,” and no uniqueness theorem imported solely to forbid alternatives. Self-citations to the codimension-1 predecessor [39] and the author’s centre-manifold preprint [38] supply the base case and a technical lemma; the new content (higher-codimension normal form, singular PDE, three regimes on a manifold, matrix-factorisation geometry) is independently derived. Numerical plots are purely confirmatory. Hence circularity is at most minor and non-load-bearing for the central claims.

Assumptions & free parameters 0 free parameters · 7 assumptions · 2 invented entities

Load-bearing content is geometric assumptions on the least-squares landscape plus one dynamical-systems conjecture for the general critical case. No numerical free parameters are fitted into the theorems. The fibre-bundle description of flat minima is proved, not postulated. The main novelty is the derivation under those assumptions, not new physical entities.

assumptions (7)
  • domain assumption Assumption 3.1: target τ is a regular value of the model f, so the solution set M is a smooth manifold.
    Standard regular-value hypothesis ensuring a smooth solution manifold; equivalent to full-rank NTK along M.
  • domain assumption Assumption 3.3: the set where λ1=λ2 is not all of M, so λ1 is simple and smooth on a nonempty open set.
    Needed for a smooth top eigenvector field ν1 and for the distinguished normal coordinate y1.
  • domain assumption Assumption 3.4: λ1 is Morse-Bott along a submanifold F of local minima, and span(ν1|F) is GD-invariant near zero.
    Replaces geodesic strong convexity about an isolated flat minimum; invariance is technical for higher-order decay.
  • domain assumption Assumption 3.5: the PDE coefficient α_η is positive on F near η=2/λ1|F and has vanishing derivative on F.
    Guarantees a C1 solution of the singular normal-form PDE (Theorem C.4); verified for matrix factorisation in Prop. E.6.
  • ad hoc to paper Conjecture D.7: invariant Lipschitz strong-stable foliation for the normally parabolic reduced system in the critical regime.
    Unproven parabolic analogue of Hirsch–Pugh–Shub stable foliation; required for general critical convergence when ∇²_M λ1 is not a scalar multiple of the identity.
  • ad hoc to paper Assumption D.12: constant smallest eigenvalue of ∇²_M λ1|F of constant multiplicity, with local invariance of the bottom-eigenspace–ν1 submanifold and z-independent (u,y1) updates.
    Extra structural hypothesis for critical regime; proved for two-layer matrix factorisation (Prop. E.7), not in full generality.
  • standard math Standard centre-manifold, normal-hyperbolicity, and invariant-manifold theorems for maps (cited [27],[38],[7]).
    Used for dimension reduction and final strong-stable foliation coordinates in the normal form.
invented entities (2)
  • η-dependent normal-form coordinates (x,y1,y2:q) for large-step GD near F independent evidence
    purpose: Make GD dynamics readable as RGD on sharpness plus flip bifurcation plus contraction.
    Coordinate construction, not a new physical object; existence is the theorem content.
  • Fibre-bundle model of flat minima F for deep matrix factorisation over a product of spheres independent evidence
    purpose: Describe geometry of flat minima and compute Morse-Bott spectrum of λ1.
    Proved structure theorem (Prop. E.2), not an extra postulate; independent of the GD dynamics theorems once the landscape is fixed.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Dynamics of Gradient Descent with Large Step Size Near a Manifold of Flat Minima." pith.science (2026). https://pith.science/paper/65VBXAJ2

@misc{pith2026260708380,
  author       = {Pith},
  title        = {Pith review of: Dynamics of Gradient Descent with Large Step Size Near a Manifold of Flat Minima},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/65VBXAJ2}},
  note         = {Machine review of arXiv:2607.08380}
}
read the original abstract

An important quantity in the theory of gradient descent (GD) is the \emph{sharpness}, defined as the largest eigenvalue of the objective Hessian. Classical analyses typically require the step size to be uniformly smaller than twice the reciprocal of the sharpness, but this condition is frequently violated in the training of deep neural networks. Recent work bridges this gap in the setting of overparametrised least-squares with a \emph{single scalar output}, providing a normal form for large-step GD in a neighbourhood of an \emph{isolated} flat minimum and establishing three corresponding convergence results. In this paper, we extend this theory in two directions: (1) to overparametrised least-squares with \emph{vector-valued outputs} (including regression with arbitrarily many observations), and (2) to a neighbourhood of a \emph{manifold} of flat minima (which we show is essential for applications such as matrix factorisation). We generalise both the normal form and all three convergence theorems of \cite{macdonaldeos} to this broader setting, overcoming several technical challenges, including the solution of a singular partial differential equation via a novel method that may be of independent interest. We further show that our framework applies to deep matrix factorisation under mild assumptions, yielding several new structural results. In particular, we prove that the set of flat minima forms a fibre bundle over a product of spheres, and that the sharpness is Morse-Bott along this manifold.

Figures

Figures reproduced from arXiv: 2607.08380 by the authors.

Figure 1
Figure 1. Roles of the assumptions in establishing the [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. The base (left), connected component of fibre (centre) and connected component of total space (right) of F for 2-layer matrix factorisation with d0 = d1 = d2 = 2. The base is the circle S 1 , while the fibre is an open subset of the so￾lution manifold of the factorisation problem of 1 dimension lower (which, for d0 = d1 = d2 = 2, is simply a 1-dimensional hyperbola). The total space is obtained by attaching a copy o… view at source ↗
Figure 3
Figure 3. Log y-scale plots of ∥yt∥ (left) and dM(xt, F) (right) for 5 independent trials of gradient descent in the subcritical regime on a 3 layer, 2 × 2 matrix factorisation problem. Initial instability (rising ∥yt∥) is overcome in finite time followed by exponential convergence to a suboptimally flat minimum. theorem for this regime is the hardest to prove. The more general case we consider is, however, even more difficul… view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Log-log plots of ∥yt∥ (left, solid) and dM(xt, F) (right, solid) for 5 independent trials of gradient descent in the critical regime on a 3 layer, 2 × 2 matrix factorisation problem. Dotted lines show t −1/2 passing through the final values of each trial for reference.…
Figure 5
Figure 5. Figure 5: Log y-scale plots of ∥yt∥ (left) and dM(xt, F) (right) for 5 independent trials of gradient descent in the supercritical regime on a 3 layer, 2 × 2 matrix factorisation problem. All trials exhibit the same exponential convergence rate to the claimed period-2 cycle alon…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

69 extracted references · 69 canonical work pages

  1. [39]

    L. E. MacDonald, H. Min, L. Palma, S. Tarmoun, Z. Xu, and R. Vidal. Convergence Rates for Gradient Descent on the Edge of Stability for Overparametrised Least Squares. InNeurIPS, 2025

  2. [1]

    Second-order regression models exhibit progressive sharpening to the edge of stability

    Atish Agarwala, Fabian Pedregosa, and Jeffrey Pennington. Second-order regression models exhibit progressive sharpening to the edge of stability. InICML, 2023

  3. [2]

    K. Ahn, J. Zhang, and S. Sra. Understanding the unstable convergence of gradient descent. In ICML, 2022

  4. [3]

    Allen-Zhu, Y

    Z. Allen-Zhu, Y . Li, and Z. Song. A Convergence Theory for Deep Learning via Over- Parameterization. InICML, pages 242–252, 2019

  5. [4]

    J. M. Altschuler and P. Parrilo. Acceleration by Stepsize Hedging: Multi-Step Descent and the Silver Stepsize Schedule.Journal of the ACM, 2023

  6. [5]

    J. M. Altschuler and P. Parrilo. Acceleration by stepsize hedging: Silver Stepsize Schedule for smooth convex optimization.Mathematical Programming, pages 1–14, 2024

  7. [6]

    Arora, Z

    S. Arora, Z. Li, and A. Panigrahi. Understanding Gradient Descent on Edge of Stability in Deep Learning. InICML, 2022

  8. [7]

    Baldomá and E

    I. Baldomá and E. Fontich. Stable manifolds associated to fixed points with linear part equal to identity.Journal of Differential Equations, 197:45–72, 2004

Show all 69 references
  1. [8]

    Bombari, M

    S. Bombari, M. H. Amani, and M. Mondelli. Memorization and Optimization in Deep Neural Networks with Minimum Over-parameterization. InNeurIPS, 2022

  2. [9]

    Y . Cai, J. Wu, S. Mei, M. Lindsey, and P. L. Bartlett. Large Stepsize Gradient Descent for Non-Homogeneous Two-Layer Networks: Margin Improvement and Fast Optimization. In NeurIPS, 2024

  3. [10]

    Méthode générale pour la résolution des systèmes d’équations simul- tanées.Comptes Rendus Hebdomadaires des Séances de l’Académie des Sciences, 25:536–538, 1847

    Augustin-Louis Cauchy. Méthode générale pour la résolution des systèmes d’équations simul- tanées.Comptes Rendus Hebdomadaires des Séances de l’Académie des Sciences, 25:536–538, 1847

  4. [11]

    Beyond the Edge of Stability via Two-step Gradient Updates

    Lei Chen and Joan Bruna. Beyond the Edge of Stability via Two-step Gradient Updates. In ICML, 2023

  5. [12]

    Chizat, E

    L. Chizat, E. Oyallon, and F. Bach. On Lazy Training in Differentiable Programming . In NeurIPS, 2019

  6. [13]

    Cohen, S

    J. Cohen, S. Kaur, Y . Li, J. Zico Kolter, and A. Talwalkar. Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability. InICLR, 2021

  7. [14]

    Jeremy Cohen, Alex Damian, Ameet Talwalkar, J Zico Kolter, and Jason D. Lee. Understanding Optimization in Deep Learning with Central Flows. InICLR, 2025

  8. [15]

    Damian, E

    A. Damian, E. Nichani, and J. Lee. Self-Stabilization: The Implicit Bias of Gradient Descent at the Edge of Stability. InICLR, 2023

  9. [16]

    Universal Sharpness Dynamics in Neural Network Training: Fixed Point Analysis, Edge of Stability, and Route to Chaos

    Dayal Singh Kalra and Tianyu He and Maissam Barkeshli. Universal Sharpness Dynamics in Neural Network Training: Fixed Point Analysis, Edge of Stability, and Route to Chaos. In ICLR, 2025

  10. [17]

    S. S. Du, J. Lee, H. Li, L. Wang, and X. Zhai. Gradient Descent Finds Global Minima of Deep Neural Networks. InICML, pages 1675–1685, 2019

  11. [18]

    S. S. Du, X. Zhai, B. Poczos, and A. Singh. Gradient Descent Provably Optimizes Over- parameterized Neural Networks. InICLR, 2019. 10

  12. [19]

    Frobenius

    G. Frobenius. Uber die Integration der linearen Differentialgleichungen durch Reihen.Journal für die reine und angewandte Mathematik, 76:214–235, 1873

  13. [20]

    Gérard and H

    R. Gérard and H. Tahara. Holomorphic and Singular Solutions of Nonlinear Singular First Order Partial Differential Equations.Publ. RIMS, Kyoto Univ., 26:979–1000, 1990

  14. [21]

    Learning dynamics of deep matrix factorization beyond the edge of stability

    Avrajit Ghosh, Soo Min Kwon, Rongrong Wang, Saiprasad Ravishankar, and Qing Qu. Learning dynamics of deep matrix factorization beyond the edge of stability. InICLR, 2025

  15. [22]

    Granziol

    D. Granziol. Flatness is a False Friend. arXiv:2006.09091, 2020

  16. [23]

    B. Grimmer. Provably faster gradient descent via long steps.SIAM Journal on Optimization, 34:2588–2608, 2024

  17. [24]

    Grimmer, K

    B. Grimmer, K. Shu, and A. L. Wang. Accelerated gradient descent via long steps. arXiv:2309.09961, 2023

  18. [25]

    Grimmer, K

    B. Grimmer, K. Shu, and A. L. Wang. Accelerated objective gap and gradient norm convergence for gradient descent via long steps.INFORMS Journal on Optimization, 7:156–169, 2025

  19. [26]

    M. W. Hirsch.Differential Topology. Springer, 1976

  20. [27]

    M. W. Hirsch, C. C. Pugh, and M. Schub.Invariant Manifolds. Springer, 1977

  21. [28]

    Hochreiter and J

    S. Hochreiter and J. Schmidhuber. Flat minima.Neural Computation, 9(1):1–42, 1997

  22. [29]

    Jacot, F

    A. Jacot, F. Gabriel, and C. Hongler. Neural Tangent Kernel: Convergence and Generalization in Neural Networks. InNeurIPS, pages 8571–8580, 2018

  23. [30]

    Karimi, J

    H. Karimi, J. Nutini, and M. Schmidt. Linear Convergence of Gradient and Proximal-Gradient Methods Under the Polyak-Łojasiewicz Condition. InECML PKDD, pages 795—-811, 2016

  24. [31]

    N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, P. Tak, and P. Tang. On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima. InICLR, 2017

  25. [32]

    Gradient descent mono- tonically decreases the sharpness of gradient flow solutions in scalar networks and beyond

    Itai Kreisler, Mor Shpigel Nacson, Daniel Soudry, and Yair Carmon. Gradient descent mono- tonically decreases the sharpness of gradient flow solutions in scalar networks and beyond. In ICML, 2023

  26. [33]

    Y . A. Kuznetsov.Elements of Applied Bifurcation Theory, Fourth Edition. Springer, 2023

  27. [34]

    J. Lee, L. Xiao, S. Schoenholtz, Y . Bahri, R. Novak, J. Sohl-Dickstein, and J. Pennington. Wide neural networks of any depth evolve as linear models under gradient descent.NeurIPS, 2019

  28. [35]

    J. D. Lee, M. Simchowitz, M. I. Jordan, and B. Recht. Gradient Descent Only Converges to Minimizers. InCOLT, 2016

  29. [36]

    Lee and C

    S. Lee and C. Jang. A new characterization of the edge of stability based on a sharpness measure aware of batch gradient distribution. InICLR, 2023

  30. [37]

    A minimalist example of edge-of-stability and progressive sharpening

    Liming Liu, Zixuan Zhang, Simon Du, and Tuo Zhao. A minimalist example of edge-of-stability and progressive sharpening. InNeurIPS, 2025

  31. [38]

    L. E. MacDonald. Centre manifold theorem for maps along manifolds of fixed points. arXiv:2604.18202, 2026

  32. [40]

    Marion and L

    P. Marion and L. Chizat. Deep linear networks for regression are implicitly regularized towards flat minima. InNeurIPS, 2024

  33. [41]

    Mengi, E

    E. Mengi, E. A. Yildirim, and M. Kilic. Numerical Optimization of Eigenvalues of Hermitian Matrix Functions.SIAM Journal on Matrix Analysis and Applications, 35, 2014. 11

  34. [42]

    Mulayoff and T

    R. Mulayoff and T. Michaeli. Unique Properties of Flat Minima in Deep Networks. InICML, 2020

  35. [43]

    Q. Nguyen. On the Proof of Global Convergence of Gradient Descent for Deep ReLU Networks with Linear Widths. InNeurIPS, 2021

  36. [44]

    Nguyen and M

    Q. Nguyen and M. Mondelli. Global Convergence of Deep Networks with One Wide Layer Followed by Pyramidal Topology. InNeurIPS, 2020

  37. [45]

    Nguyen, M

    Q. Nguyen, M. Mondelli, and G. Montufar. Tight Bounds on the Smallest Eigenvalue of the Neural Tangent Kernel for Deep ReLU Networks. InICML, 2021

  38. [46]

    H. Tahara. Solvability of partial di¤erential equations of nonlinear totally characteristic type with resonances.J. Math. Soc. Japan, 55:1095–1113, 2003

  39. [47]

    Large Learning Rate Tames Homo- geneity: Convergence and Balancing Effect

    Yuqing Wang, Minshuo Chen, Tuo Zhao, and Molei Tao. Large Learning Rate Tames Homo- geneity: Convergence and Balancing Effect. InICLR, 2022

  40. [48]

    Good regularity creates large learning rate implicit biases: edge of stability, balancing, and catapult

    Yuqing Wang, Zhenghao Xu, Tuo Zhao, and Molei Tao. Good regularity creates large learning rate implicit biases: edge of stability, balancing, and catapult. InNeurIPS 2023 Workshop on Mathematics of Modern Machine Learning, 2023

  41. [49]

    Z. Wang, Z. Li, and J. Li. Analyzing Sharpness along GD Trajectory: Progressive Sharpening and Edge of Stability. InNeurIPS, 2022

  42. [50]

    K. Wen, Z. Li, and T. Ma. Sharpness Minimization Algorithms Do Not Only Minimize Sharpness To Achieve Better Generalization. InNeurIPS, 2023

  43. [51]

    J. Wu, P. L. Bartlett, M. Telgarsky, and B. Yu. Large Stepsize Gradient Descent for Logistic Loss: Non-Monotonicity of the Loss Improves Optimization Efficiency. InCOLT, 2024

  44. [52]

    J. Wu, V . Braverman, and J. Lee. Implicit Bias of Gradient Descent for Logistic Regression at the Edge of Stability. InNeurIPS, 2023

  45. [53]

    J. Wu, P. Marion, and P. L. Bartlett. Large Stepsizes Accelerate Gradient Descent for Regularized Logistic Regression. InNeurIPS, 2025

  46. [54]

    L. Wu, C. Ma, and W. E. How SGD Selects the Global Minima in Over-parameterized Learning: A Dynamical Stability Perspective. InNeurIPS, 2018

  47. [55]

    and Song, M

    Yoo, G. and Song, M. and Yun, C. Understanding Sharpness Dynamics in NN Training with a Minimalist Example: The Effects of Dataset Difficulty, Depth, Stochasticity, and More. In ICML, 2025

  48. [56]

    Yoshino and A

    M. Yoshino and A. Shirai. Singular solutions of nonlinear partial differential equations with resonances.J. Math. Soc. Japan, 60:237–263, 2008

  49. [57]

    X. Zhu, Z. Wang, X. Wang, M. Zhou, and R. Ge. Understanding Edge-of-Stability Training Dynamics with a Minimalist Example. InICLR, 2023. 12 A Additional notation Given vector bundles E and E′ over a space X denote by Hom(E, E ′) the vector bundle over X whose fibre over x∈X is...

  50. [58]

    Intersects R×M⊂R×R p in a neighbourhoodeV of {2/λ1|F } ×M which is open relative toR×M

  51. [59]

    Is tangent at each point(η, x)∈ eVtoT (η,x)(R×M)⊕span(ν 1(x)). 15

  52. [60]

    regular singular

    Is invariant underGD, i.e.GD(W c)⊂W c. We now demonstrate how the invariant manifold of Lemma B.3 can be applied to reduce the orthogonal dynamics to being essentially 1-dimensional. In the (η, x, z) coordinates of Lemma B.2, there is an open subset V ′′ 1 ⊂R containing 0 such...

  53. [61]

    The setV ρ,∆ :=I ρ ×U ρ,∆ ⊂Vis invariant underGD

  54. [62]

    Then for any ρ >0 sufficiently small, the set Vρ,∆ as defined in the statement is contained inV ′, hence in V , and satisfies β|Vρ,∆ ≤ −∆

    For any(η, x, y)∈V ρ,∆ one has ∥Iq−1 −ηΛ 2:q(x)∥+C(|y 1|+∥y 2:q∥)−min{|1−ηλ 1(x)|,1} ≤ −∆.(160) Proof.Consider the continuous functionβ:V→Rdefined by β(η, x, y) :=∥I−ηΛ 2:q(x)∥+C |y1|+∥y 2:q∥ −min{|1−ηλ 1(x)|,1}.(161) Since ¯x∈F one has ∥Λ2:q(¯x)∥< λ1|F , and since Λ2:q(¯x)is ...

  55. [63]

    Finally, we assume that for any ¯z∈Rdz and any sufficiently small ρ >0 , there is an open neighbourhood Vρ of(¯z,0,0,0)of diameter at mostρwhich is invariant underT

    +R y1(z, v, u, y),(209) Ty2:q(z, v, u, y) =B(z, v, u) 2y2:q +R y2:q(z, v, u, y),(210) whereR z, Rv, Ru, Ry1 andR y2:q areC 1, with Rz(z, v, u, y), R v(z, v, u, y) =O(y 3 1∥v∥, y2 1∥v∥(∥v∥+∥u∥)),(211) Ru(z, v, u, y) =O(y 3 1(∥u∥+∥v∥), y 2 1(∥u∥+∥v∥) 2),(212) Ry1(z, v, u, y) =O(...

  56. [64]

    For anyϵ >0, there isδ >0such thatLip(ϕ| {u∈U:∥u∥<δ} )≤κ+ϵ. 30

  57. [65]

    weak contraction

    The graphW ϕ :={(z ′,0, u, ϕ(u),0) :u∈U, z ′ nearz} ⊂Wis invariant underT. Proof.We will apply [7, Theorem 3.1]. This requires the coordinate transformation y1 =κ∥u∥+ey 1.(216) With respect to the new coordinates (z, v, u,ey1, y2:q), denoting ε:=∥u∥+∥v∥+ey 1 for notational eas...

  58. [66]

    +R y1(u, y1),(226) 31 Ty2:q(z, v, u, y) =B(z, v, u) 2y2:q +R y2:q(z, v, u, y),(227) where Rv(z, v, u, y) =O(y 3 1∥v∥, y2 1∥v∥(∥u∥+∥v∥)),(228) Ru(u, y1) =O(y 3 1∥u∥, y2 1∥u∥2),(229) Ry1(u, y1) =O(y 4 1, y1∥u∥3)(230) and Ry2:q(z, v, u, y) =O(y 1∥y2:q∥,∥y 2:q∥2, y2 1(∥u∥+∥v∥)∥y 2...

  59. [67]

    Thus |Φ◦T u,y1(u, y1)| ≤ |Φ(u, y 1)|(1−2cy 2 1).(249) We now turn to lower-boundingT u(u, y1)/∥u∥

    + 2a(κ+ϵ)∥u∥(ϕ(u) +y 1) +O(∥u∥ 3, y3 1,∥u∥ 2y1, y2 1∥u∥) (246) ≤1−2cy 2 1 −κ(c−a)∥u∥y 1 + 6κϵ(a+c)∥u∥ 2 (247) ≤1−2cy 2 1 (248) by taking ρ yet smaller if necessary to obtain the third line and using y1 ≥(κ/ √ 2)∥u∥ together with (240) to obtain the fourth. Thus |Φ◦T u,y1(u, y1...

  60. [68]

    Consider the ratio q:=y 1/∥u∥

    can be met from any sufficiently small, nonzero initial condition in at most O(∥u∥−2) iterations. Consider the ratio q:=y 1/∥u∥. Setting γ:= ( √ 2−1)κ/(2 √

  61. [69]

    and shrinking ρ yet further if necessary so that |ϕ(u)/∥u∥ −κ|< γ/2 , either q∈[κ−γ/2, κ+γ/2] , in which case we can set τ:= 0 , or q is outside of[κ−γ/2, κ+γ/2]. In the latter case, observe that q◦T u,y1(u, y1) =q 1 + 2b∥u∥2 −2cq 2∥u∥2 +O(∥u∥ 3) 1−2aq 2∥u∥2 +O(∥u∥ 3) (252) =q...

Pith tools

Reviewed July 10, 2026 · model on record in the stance chip above.