REVIEW 3 major objections 5 minor 68 references
Explaining Near-Zero Hessian Eigenvalues Through Approximate Symmetries in Neural Networks
T0 review · 3 major / 5 minor · reviewed 2026-07-10 · grok-4.5
Pith's one-line read The bulk of near-zero Hessian eigenvalues in neural nets are weakly broken continuous symmetries of the architecture.
desk verdict Clean eigenvector-level account of the Hessian bulk as weakly broken architectural symmetries; the math and diagnostics hold up. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Explicit orthogonal generators of the GL-type interlayer symmetries (built from singular vectors of consecutive weight matrices) together with the eigenvector-overlap diagnostic that measures how much each Hessian mode of a nonlinear network lies inside that linear symmetry subspace.
What would settle it
Train a multilayer ReLU network to a genuine critical point, compute the leading Hessian eigenvectors, and check whether their overlap with the linear-symmetry subspace remains near zero for the high-curvature modes and near one for the bulk; a clear failure of that two-tier structure would falsify the claim.
Extended reading notes
Core claim
The bulk of near-zero Hessian eigenvalues consists of the weakly lifted pseudo-Goldstone modes of the continuous symmetries of the network parametrization. In linear networks the symmetries are exact and their generators form an explicit orthogonal basis of the null space; a ReLU nonlinearity breaks them weakly, so that high-curvature eigenvectors stay orthogonal to the symmetry subspace while the bulk eigenvectors remain almost entirely inside it.
Load-bearing premise
That the null space of a linear comparison model (nonlinearities removed, low-variance input directions zeroed) remains a faithful diagnostic of the bulk even after training on real data, without a general guarantee that training or non-Gaussian inputs do not mix the subspaces beyond the residual already measured.
Editorial extensions
If this is right
- The bulk of near-zero modes is largely architectural and therefore persists across data sets and training algorithms that preserve the same continuous symmetries.
- Because the same directions remain zero modes of the Fisher matrix at arbitrary parameters, natural-gradient and second-order methods automatically ignore or treat them specially.
- Any architecture containing fully connected or convolutional blocks inherits an analogous bulk whose size is fixed by layer widths and over-parametrization count.
- The adiabatic connection from linear to ReLU spectra supplies a controlled starting point for analytic approximations of the bulk eigenvalues.
Reading between the lines
- If the bulk is mostly architectural, pruning or regularization that deliberately targets the symmetry subspace may remove far more parameters than curvature-based pruning alone suggests.
- The same diagnostic should apply to the fully connected blocks inside transformers; verifying the two-tier overlap there would test whether the mechanism survives attention and residual pathways.
- Because the residual term in the Hessian can reintroduce curvature off critical points, the pseudo-Goldstone picture may degrade late in training when gradients no longer vanish.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that the bulk of near-zero Hessian eigenvalues of neural-network training losses consists of weakly lifted pseudo-Goldstone modes of continuous architectural symmetries of the parametrization. In multilayer linear networks these symmetries are exact; the authors construct an explicit orthogonal basis of generators (via SVDs of the weight matrices) that spans the null space of the Hessian/Fisher and matches the known finite-eigenvalue count. A Leaky-ReLU deformation is treated as an explicit symmetry-breaking perturbation: the two-layer Gaussian student–teacher Fisher splits exactly into linear and absolute-value blocks, bulk eigenvalues rise as ε², and Kato perturbation theory supplies a confinement criterion (vanishing ε⁴ correction iff eigenvectors remain in the symmetry subspace). Eigenvector overlaps confirm that high-curvature modes are orthogonal to the symmetry subspace while the bulk lies inside it. The same diagnostic is applied to a three-layer student–teacher model, a trained three-layer ReLU MLP on CIFAR-10 (with whitening and cumulative coverage Ok), and a minimal convolutional network with nonlinear (circulant) generators.
Significance. If the eigenvector-level claim holds, it supplies a single architectural origin for the Hessian bulk that has been missing from the literature on outliers, Fisher spectra, and flat minima. The work goes beyond counting zero modes: it constructs the generators explicitly, derives the ε² lifting and confinement criterion analytically for the two-layer Gaussian case, and measures overlaps and cumulative coverage on both idealized and trained models. Code and data are released. The mechanism is expected to organize Fisher/Gauss–Newton spectra as well, with direct consequences for natural-gradient and second-order methods, and the convolutional example shows the diagnostic is not limited to fully connected layers. These are concrete, falsifiable contributions at the level of eigenvectors rather than eigenvalue counts alone.
major comments (3)
- End Matter C and SM §II establish the clean ε² law and confinement criterion only for the two-layer Gaussian student–teacher Fisher (exact block split, vanishing cross term by parity). In the three-layer SM case (Fig. S1) the ε² trend already bends before ε=1 and leading-mode overlaps level off near oi≈0.2 rather than near zero. The main-text claim that the bulk “lies almost entirely within” the symmetry subspace therefore needs an explicit scope statement: for deeper nets the adiabatic connection is only approximate and residual leakage is O(1). A short quantitative bound or additional depth-controlled experiment would make the generality claim load-bearing rather than extrapolative.
- CIFAR-10 section and SM §IV: the comparison subspace P mixes architectural GL-type symmetries with data-covariance flat directions obtained by zeroing low-variance input components. Dimension counting separates the two contributions (383 232 vs 26 748), and Ok≈0.95 bounds residual leakage of the unmeasured tail, but the non-whitened run (Fig. S2 right) shows that covariance-induced spread entangles the two mechanisms and smooths the overlap transition. The paper should state more sharply which fraction of the observed bulk is architectural versus data-driven, and whether the dramatic suppression of the leading Neff Ceff modes survives when the linear comparison Fisher is built without artificially zeroing variances.
- Eq. (1) and the Mexican-hat appendix correctly note that symmetry generators are exact Hessian zero modes only at critical points; off criticality the residual term can produce finite curvature. The CIFAR experiment evaluates the training-loss Hessian at a non-critical endpoint and compares to the Fisher null space of the linear model. While the observed two-tier structure is still striking, a brief check that the residual contribution along the measured bulk directions remains small (or a comparison to the Gauss–Newton matrix itself) would close the gap between the critical-point theory and the trained-network diagnostic.
minor comments (5)
- Fig. 2 caption and main text: the N rescaling modes that remain exactly zero for all ε are stated to lie “below the plotted range”; a short inset or explicit note that they are omitted would avoid the impression that the bulk starts above zero.
- Notation for the projector P and the overlap oi (Eq. 8) is introduced cleanly, but the random baseline orandom = rank(P)/d is quoted with different numerical values in different figures; a single consistent formula and the associated standard deviation for the CIFAR case would help the reader.
- The convolutional construction (SM §V) uses nonlinear generators involving C(w(1))⁻¹; a one-sentence remark in the main text that these are still continuous symmetries of the function (hence still produce Fisher zeros when exact) would clarify why the more general ϕ is needed.
- References [46,47] already count finite eigenvalues of deep linear networks; the novelty claim is correctly placed on the eigenvectors and the nonlinear extension, but a slightly sharper sentence distinguishing the present work from those counts would help.
- Typographical: “parametrization” is used consistently in the abstract/title; a few places in the SM switch to “parameterization.” Standardize.
Circularity Check
No circularity: symmetry generators defined independently of Hessian; overlaps and coverage are measured diagnostics, not fitted predictions.
full rationale
The derivation chain is self-contained and non-circular. Continuous symmetries of the linear network are defined by function-preserving transformations (GL+ action inserting M and M^{-1} between layers, Eq. 4), yielding generators φ'_A (Eq. 5) constructed via SVD of the weight matrices; these are shown to span the exact null space of H (and of F at arbitrary parameters) by direct verification of mutual orthogonality and dimension counting. The nonlinear case treats ReLU as an explicit perturbation of the linear network (Leaky-ReLU family, Eq. 6); the ε^{2} bulk scaling and confinement follow from the exact block decomposition of the Fisher (End Matter C, Eqs. C4–C6) under Gaussian parity, not from any fit. Overlaps o_i (Eq. 8) and cumulative coverage O_k (Eq. 9) are post-hoc measurements of the already-computed Hessian eigenvectors against the independently constructed linear symmetry projector P; the linear comparison model is obtained simply by deleting nonlinearities from the trained weights (or zeroing low-variance input directions), never by optimizing parameters to match the spectrum. The convolutional generators are likewise derived from the requirement that the transformed circulant remain circulant (SM V). No self-definitional loop, no fitted parameter re-labeled as prediction, and no load-bearing self-citation of an unverified uniqueness claim appear. The GitHub reference is solely for data availability.
Assumptions & free parameters
free parameters (2)
- Neff (whitening cutoff) =
100
- Leaky-ReLU leakage ε
assumptions (4)
- standard math Continuous symmetries of the network function generate exact zero modes of the Fisher (and of the Hessian at critical points) via the identity H ϕ' + [∂θ(ϕ')ᵀ] ∂θ L = 0.
- standard math Kato analytic perturbation theory for eigenvalues leaving a degenerate level applies to the Fisher family F(ε).
- domain assumption For positively homogeneous activations the diagonal rescaling subgroup remains an exact symmetry for all ε.
- domain assumption At an interpolating student-teacher minimum the residual term in the Gauss-Newton decomposition vanishes, so H = F.
invented entities (1)
-
pseudo-Goldstone modes of architectural symmetries
independent evidence
Cite this review
Pith. "Pith review of Explaining Near-Zero Hessian Eigenvalues Through Approximate Symmetries in Neural Networks." pith.science (2026). https://pith.science/paper/OA4P2EIV
@misc{pith2026260707845,
author = {Pith},
title = {Pith review of: Explaining Near-Zero Hessian Eigenvalues Through Approximate Symmetries in Neural Networks},
year = {2026},
howpublished = {\url{https://pith.science/paper/OA4P2EIV}},
note = {Machine review of arXiv:2607.07845}
}
read the original abstract
The Hessian of the training loss governs the local geometry of the loss landscape, yet despite existing explanations for its largest eigenvalues, the origin of the vast multitude of vanishingly small eigenvalues remains elusive. We argue that the bulk consists of the weakly lifted pseudo-Goldstone modes of the continuous symmetries of the network parametrization. In deep linear networks these symmetries are exact: they generate flat directions and hence exact zero modes, whose eigenvectors we construct explicitly. Introducing a ReLU nonlinearity as a perturbation, we show that it breaks these symmetries weakly and explicitly. Resolving the spectrum at the level of eigenvectors, we find that the high-curvature directions are orthogonal to the symmetry subspace, while the bulk lies almost entirely within it. We demonstrate the mechanism in a two-layer ReLU student--teacher model and in a network trained on CIFAR-10. A convolutional example demonstrates that the same diagnostic extends beyond fully connected layers. Together, these results link the Hessian bulk to weakly broken symmetries and clarify the origin of near-zero modes.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
- [1]
- [2]
- [3]
-
[4]
A. Atanasov, A. Meterez, J. Simon, and C. Pehlevan, The Optimization Landscape of SGD Across the Fea- ture Learning Strength, inInternational Conference on Learning Representations(2025)
work page 2025
- [5]
-
[6]
Amari, Natural Gradient Works Efficiently in Learn- ing, Neural Computation10, 251 (1998)
S. Amari, Natural Gradient Works Efficiently in Learn- ing, Neural Computation10, 251 (1998)
work page 1998
-
[7]
J. Martens, Deep learning via Hessian-free optimization, inProceedings of the 27th International Conference on International Conference on Machine Learning(Omni- press, 2010) p. 735–742
work page 2010
-
[8]
J. Martens and R. Grosse, Optimizing Neural Networks with Kronecker-factored Approximate Curvature, inPro- ceedings of the 32nd International Conference on Ma- chine Learning, Vol. 37 (PMLR, 2015) pp. 2408–2417
work page 2015
Show all 68 references
-
[9]
Z. Yao, A. Gholami, S. Shen, M. Mustafa, K. Keutzer, and M. Mahoney, Adahessian: An adaptive second or- der optimizer for machine learning, inproceedings of the AAAI conference on artificial intelligence, Vol. 35 (2021) pp. 10665–10673
2021
-
[10]
Hochreiter and J
S. Hochreiter and J. Schmidhuber, Flat Minima, Neural Computation9, 1 (1997)
1997
-
[11]
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyan- skiy, and P. T. P. Tang, On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima, inInternational Conference on Learning Representations (2017)
2017
-
[12]
Foret, A
P. Foret, A. Kleiner, H. Mobahi, and B. Neyshabur, Sharpness-aware Minimization for Efficiently Improving Generalization, inInternational Conference on Learning Representations(2021)
2021
-
[13]
Chaudhari, A
P. Chaudhari, A. Choromanska, S. Soatto, Y. Le- Cun, C. Baldassi, C. Borgs, J. Chayes, L. Sagun, and R. Zecchina, Entropy-SGD: biasing gradient descent into wide valleys, Journal of Statistical Mechanics: Theory and Experiment , 124018 (2019)
2019
-
[14]
Jastrzębski, Z
S. Jastrzębski, Z. Kenton, D. Arpit, N. Ballas, A. Fis- cher, Y. Bengio, and A. Storkey, Three Factors Influenc- ing Minima in SGD (2018), arXiv:1711.04623
2018 arXiv
-
[15]
LeCun, J
Y. LeCun, J. Denker, and S. Solla, Optimal Brain Dam- age, inAdvances in Neural Information Processing Sys- 6 tems, Vol. 2 (Morgan-Kaufmann, 1989)
1989
-
[16]
Hassibi and D
B. Hassibi and D. Stork, Second order derivatives for network pruning: Optimal Brain Surgeon, inAdvances in Neural Information Processing Systems, Vol. 5 (Morgan- Kaufmann, 1992)
1992
-
[17]
S. P. Singh and D. Alistarh, WoodFisher: Efficient Second-Order Approximation for Neural Network Com- pression, inAdvances in Neural Information Process- ing Systems, Vol. 33 (Curran Associates, Inc., 2020) pp. 18098–18109
2020
-
[18]
Frantar, S
E. Frantar, S. Ashkboos, T. Hoefler, and D. Alis- tarh, OPTQ: Accurate Quantization for Generative Pre- trained Transformers, inThe Eleventh International Conference on Learning Representations(2023)
2023
-
[19]
Z. Yao, A. Gholami, Q. Lei, K. Keutzer, and M. W. Ma- honey, Hessian-based Analysis of Large Batch Training and Robustness to Adversaries, inAdvances in Neural Information Processing Systems, Vol. 31 (Curran Asso- ciates, Inc., 2018)
2018
-
[20]
Moosavi-Dezfooli, A
S.-M. Moosavi-Dezfooli, A. Fawzi, J. Uesato, and P. Frossard, Robustness via Curvature Regularization, and Vice Versa, in2019 IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR)(2019)pp. 9070–9078
2019
-
[21]
Singla and S
S. Singla and S. Feizi, Second-Order Provable Defenses against Adversarial Attacks, inProceedings of the 37th International Conference on Machine Learning, Proceed- ings of Machine Learning Research, Vol. 119 (PMLR,
-
[22]
Garipov, P
T. Garipov, P. Izmailov, D. Podoprikhin, D. P. Vetrov, and A. G. Wilson, Loss Surfaces, Mode Connectivity, and Fast Ensembling of DNNs, inAdvances in Neural Infor- mation Processing Systems, Vol. 31, edited by S. Bengio, H.Wallach, H.Larochelle, K.Grauman, N.Cesa-Bianchi, and ...
2018
-
[23]
J. Brea, B. Simsek, B. Illing, and W. Gerstner, Weight- space symmetry in deep networks gives rise to permuta- tion saddles, connected by equal-loss valleys across the loss landscape (2019), arXiv:1907.02911
2019 arXiv
-
[24]
Ainsworth, J
S. Ainsworth, J. Hayase, and S. Srinivasa, Git Re-Basin: Merging Models modulo Permutation Symmetries, in The Eleventh International Conference on Learning Rep- resentations(2023)
2023
-
[25]
A. Ito, M. Yamada, and A. Kumagai, Linear Mode Con- nectivity between Multiple Models modulo Permutation Symmetries, inProceedings of the 42nd International Conference on Machine Learning, Proceedings of Ma- chine Learning Research, Vol. 267 (PMLR, 2025) pp. 26611–26626
2025
-
[26]
Sagun, L
L. Sagun, L. Bottou, and Y. LeCun, Eigenvalues of the Hessian in Deep Learning: Singularity and Beyond (2017), arXiv:1611.07476
2017 arXiv
-
[27]
Sagun, U
L. Sagun, U. Evci, V. U. Guney, Y. Dauphin, and L. Bottou, Empirical Analysis of the Hessian of Over- ParametrizedNeuralNetworks(2018),arXiv:1706.04454
2018 arXiv
-
[28]
Ghorbani, S
B. Ghorbani, S. Krishnan, and Y. Xiao, An Investiga- tionintoNeuralNetOptimizationviaHessianEigenvalue Density, inProceedings of the 36th International Confer- ence on Machine Learning, Vol. 97 (PMLR, 2019) pp. 2232–2241
2019
-
[29]
Papyan, Measurements of Three-Level Hierarchical StructureintheOutliersintheSpectrumofDeepnetHes- sians, inProceedings of the 36th International Conference on Machine Learning, Vol
V. Papyan, Measurements of Three-Level Hierarchical StructureintheOutliersintheSpectrumofDeepnetHes- sians, inProceedings of the 36th International Conference on Machine Learning, Vol. 97 (2019) pp. 5012–5021
2019
-
[30]
Papyan, Traces of Class/Cross-Class Structure Per- vade Deep Learning Spectra, Journal of Machine Learn- ing Research21, 1 (2020)
V. Papyan, Traces of Class/Cross-Class Structure Per- vade Deep Learning Spectra, Journal of Machine Learn- ing Research21, 1 (2020)
2020
-
[31]
Arjevani and M
Y. Arjevani and M. Field, Analytic Characterization of the Hessian in Shallow ReLU Models: A Tale of Symme- try, inAdvances in Neural Information Processing Sys- tems, Vol. 33 (Curran Associates, Inc., 2020) pp. 5441– 5452
2020
-
[32]
Arjevani and M
Y. Arjevani and M. Field, Analytic Study of Families of Spurious Minima in Two-Layer ReLU Neural Networks: A Tale of Symmetry II, inAdvances in Neural Informa- tion Processing Systems,Vol.34(CurranAssociates, Inc.,
-
[33]
Pennington and P
J. Pennington and P. Worah, The Spectrum of the Fisher Information Matrix of a Single-Hidden-Layer Neural Net- work, inAdvances in Neural Information Processing Sys- tems, Vol. 31 (Curran Associates, Inc., 2018)
2018
-
[34]
Karakida, S
R. Karakida, S. Akaho, and S.-i. Amari, Universal Statis- tics of Fisher Information in Deep Neural Networks: Mean Field Approach, inProceedings of the Twenty- Second International Conference on Artificial Intelli- gence and Statistics, Proceedings of Machine Learning Research...
2019
-
[35]
Karakida, S
R. Karakida, S. Akaho, and S.-i. Amari, Pathological Spectra of the Fisher Information Metric and Its Vari- ants in Deep Neural Networks, Neural Computation33, 2274 (2021)
2021
-
[36]
Fort and S
S. Fort and S. Ganguli, Emergent properties of the local geometry of neural loss landscapes (2019), arXiv:1910.05929
2019 arXiv
-
[37]
Gur-Ari, D
G. Gur-Ari, D. A. Roberts, and E. Dyer, Gradi- ent Descent Happens in a Tiny Subspace (2018), arXiv:1812.04754
2018 arXiv
-
[38]
S. S. Du, W. Hu, and J. D. Lee, Algorithmic regulariza- tion in learning deep homogeneous models: Layers are automatically balanced, Advances in neural information processing systems31(2018)
2018
-
[39]
Simsek, F
B. Simsek, F. Ged, A. Jacot, F. Spadaro, C. Hongler, W. Gerstner, and J. Brea, Geometry of the Loss Land- scape in Overparameterized Neural Networks: Symme- tries and Invariances, inProceedings of the 38th Interna- tional Conference on Machine Learning, Proceedings of Machine ...
2021
-
[40]
Kunin, J
D. Kunin, J. Sagastuy-Brena, S. Ganguli, D. L. Yamins, and H. Tanaka, Neural Mechanics: Symmetry and Bro- ken Conservation Laws in Deep Learning Dynamics, in International Conference on Learning Representations (2021)
2021
-
[41]
Tanaka and D
H. Tanaka and D. Kunin, Noether’s Learning Dynamics: Role of Symmetry Breaking in Neural Networks, inAd- vances in Neural Information Processing Systems, edited by A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan (2021)
2021
-
[42]
B. Zhao, N. Dehmamy, R. Walters, and R. Yu, Symmetry Teleportation for Accelerated Optimization, inAdvances in Neural Information Processing Systems, Vol. 35 (Cur- ran Associates, Inc., 2022) pp. 16679–16690
2022
-
[43]
Marcotte, R
S. Marcotte, R. Gribonval, and G. Peyré, Abide by the law and follow the flow: conservation laws for gradient flows, inThirty-seventh Conference on Neural Informa- tion Processing Systems(2023)
2023
-
[44]
Watanabe,Algebraic geometry and statistical learning theory, Vol
S. Watanabe,Algebraic geometry and statistical learning theory, Vol. 25 (Cambridge university press, 2009). 7
2009
-
[45]
K.FukumizuandS.Amari,Localminimaandplateausin hierarchical structures of multilayer perceptrons, Neural Networks13, 317 (2000)
2000
-
[46]
S. P. Singh, G. Bachmann, and T. Hofmann, Analytic Insights into Structure and Rank of Neural Network Hes- sian Maps, inAdvances in Neural Information Process- ing Systems, Vol. 34 (Curran Associates, Inc., 2021) pp. 23914–23927
2021
-
[47]
Bernacchia, M
A. Bernacchia, M. Lengyel, and G. Hennequin, Exact natural gradient in deep linear networks and its applica- tion to the nonlinear case, inAdvances in Neural Infor- mation Processing Systems, Vol. 31 (Curran Associates, Inc., 2018)
2018
-
[48]
C. M. Bishop and N. M. Nasrabadi,Pattern recognition and machine learning, Vol. 4 (Springer, 2006)
2006
-
[49]
Heskes, On “natural” learning and pruning in multi- layered perceptrons, Neural Computation12, 881 (2000)
T. Heskes, On “natural” learning and pruning in multi- layered perceptrons, Neural Computation12, 881 (2000)
2000
-
[50]
J.Martens,NewInsightsandPerspectivesontheNatural Gradient Method, Journal of Machine Learning Research 21, 1 (2020)
2020
-
[51]
See Supplemental Material for derivations, model exten- sions, and numerical details
-
[52]
Kawaguchi, Deep Learning without Poor Local Min- ima, inAdvances in Neural Information Processing Sys- tems, Vol
K. Kawaguchi, Deep Learning without Poor Local Min- ima, inAdvances in Neural Information Processing Sys- tems, Vol. 29 (Curran Associates, Inc., 2016)
2016
-
[53]
A. L. Maas, A. Y. Hannun, A. Y. Ng,et al., Rectifier nonlinearities improve neural network acoustic models, inProceedings of the 30th International Conference on Machine Learning, Vol. 28 (2013)
2013
-
[54]
Karakida, S
R. Karakida, S. Akaho, and S.-i. Amari, The Normal- ization Method for Alleviating Pathological Sharpness in Wide Neural Networks, inAdvances in Neural Informa- tion Processing Systems,Vol.32(CurranAssociates, Inc., 2019)
2019
-
[55]
Krizhevsky,Learning multiple layers of features from tiny images, Tech
A. Krizhevsky,Learning multiple layers of features from tiny images, Tech. Rep. (University of Toronto, Toronto, Ontario, 2009)
2009
-
[56]
Zhang, S
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals, Understanding deep learning (still) requires rethinking generalization, Communications of the ACM64(2021)
2021
-
[57]
independent compo- nents
A. J. Bell and T. J. Sejnowski, The “independent compo- nents” of natural scenes are edge filters, Vision Research 37, 3327 (1997)
1997
-
[58]
J. Lee, S. Schoenholz, J. Pennington, B. Adlam, L. Xiao, R. Novak, and J. Sohl-Dickstein, Finite Versus Infinite Neural Networks: an Empirical Study, inAdvances in Neural Information Processing Systems, Vol. 33 (Curran Associates, Inc., 2020) pp. 15156–15172
2020
-
[59]
Fukushima, Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position, Biological cybernetics36, 193 (1980)
K. Fukushima, Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position, Biological cybernetics36, 193 (1980)
1980
-
[60]
LeCun, L
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, Gradient-based learning applied to document recogni- tion, Proceedings of the IEEE86, 2278 (1998)
1998
-
[61]
Goodfellow, Y
I. Goodfellow, Y. Bengio, and A. Courville,Deep Learn- ing(MIT press Cambridge, MA, USA, 2016)
2016
-
[62]
R.PascanuandY.Bengio,Revisitingnaturalgradientfor deep networks, arXiv preprint arXiv:1301.3584 (2013)
2013 arXiv
-
[63]
Kühn and B
M. Kühn and B. Rosenow, Github repos- itory,https://github.com/RosenowGroup/ approximate-symmetries-hessian(2026)
2026
-
[64]
Explaining Near-Zero Hessian Eigenvalues Through Approximate Symmetries in Neural Networks
T. Kato,Perturbation theory for linear operators, 2nd ed. (Springer, 1995). END MA TTER Appendix A: Curvature despite symmetry in the Mexican-hat picture—The mechanism of Eq. (1) is seen most simply in the rotation-invariant potential L(x, y) = (x2 +y 2 −1) 2 (Fig. 6). At the ...
1995
-
[65]
For all (k,n )∈{1,...,M}2, at least one of the two blocks in Eq
M≤min(N,C ). For all (k,n )∈{1,...,M}2, at least one of the two blocks in Eq. (S6) is nonzero. Hence all M2 generators are finite and linearly independent, yielding exactly M2 symmetries. This holds in particular in bottleneck configurations, where the naive parameter-counting...
-
[66]
In this case one would naively obtainM2 generators
M≥max(N,C ). In this case one would naively obtainM2 generators. However, onlyM2−(M−N)(M−C) = MN +MC−NC of them are independent. Indeed, for (k,n ) such that boths(1) n = 0 ands(2) k = 0, the vector Eq. (S6) vanishes. Removing these (M−N)(M−C) zero vectors leaves precisely the...
-
[67]
independent components
C < M < N (and analogously N < M < C). Here the construction above yields M2 generators, while parameter counting predicts an additional ( N−M)(M−C) symmetry directions. These arise from additional generators that emerge from our definition via the singular value decomposition...
2000
-
[68]
15156–15172
pp. 15156–15172. [S6] M. K¨ uhn and B. Rosenow, Github repository, https://github.com/RosenowGroup/approximate-symmetries-hessian (2026). 10 140 150 160 170 180 190 Index (i) 0 5 Eigenvalue ( i) 140 150 160 170 180 190 Index (i) 0.0 0.5 1.0 Overlap (oi) d K random baseline FIG...
2026
Reviewed July 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.