Pith. sign in

REVIEW 3 cited by

Dissecting Hessian: Understanding Common Structure of Hessian in Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.04261 v6 pith:QPHARSYS submitted 2020-10-08 cs.LG cs.NEstat.ML

classification cs.LGcs.NEstat.ML
keywords networkshessianneuralstructurematricesconjecturedecouplinghessians
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Hessian captures important properties of the deep neural network loss landscape. Previous works have observed low rank structure in the Hessians of neural networks. In this paper, we propose a decoupling conjecture that decomposes the layer-wise Hessians of a network as the Kronecker product of two smaller matrices. We can analyze the properties of these smaller matrices and prove the structure of top eigenspace random 2-layer networks. The decoupling conjecture has several other interesting implications - top eigenspaces for different models have surprisingly high overlap, and top eigenvectors form low rank matrices when they are reshaped into the same shape as the corresponding weight matrix. All of these can be verified empirically for deeper networks. Finally, we use the structure of layer-wise Hessian to get better explicit generalization bounds for neural networks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 12 citations worldwide. Full citation record

  1. Decoupling Variance and Scale-Invariant Updates in Adaptive Gradient Descent for Unified Vector and Matrix Optimization

    cs.LG 2026-02 conditional novelty 6.0 of 10

    DeVA_S8 reweights Muon's matrix-sign update in the matrix's eigenbasis with a singular-value signal-to-noise ratio, reaching target LLM validation perplexity with ~6.6% fewer tokens than Muon.

  2. On the Convergence Analysis of Muon

    stat.ML 2025-05 unverdicted novelty 6.0 of 10

    Muon's convergence rate depends on an average Hessian curvature along its update directions, which can be much smaller than the worst-case Lipschitz constant when Hessians are low-rank.

  3. Elucidating Subspace Perturbation in Zeroth-Order Optimization: Theory and Practice at Scale

    cs.LG 2025-01 conditional novelty 6.0 of 10

    A unified analysis of subspace perturbations in zero-order optimization identifies subspace alignment as the key driver of convergence, and leads to MeZO-BCD, a block-coordinate method with up to 2.77x wall-clock spee...

Pith tools