REVIEW 3 cited by
A Unifying View on Implicit Bias in Training Linear Neural Networks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
abstract
We study the implicit bias of gradient flow (i.e., gradient descent with infinitesimal step size) on linear neural network training. We propose a tensor formulation of neural networks that includes fully-connected, diagonal, and convolutional networks as special cases, and investigate the linear version of the formulation called linear tensor networks. With this formulation, we can characterize the convergence direction of the network parameters as singular vectors of a tensor defined by the network. For $L$-layer linear tensor networks that are orthogonally decomposable, we show that gradient flow on separable classification finds a stationary point of the $\ell_{2/L}$ max-margin problem in a "transformed" input space defined by the network. For underdetermined regression, we prove that gradient flow finds a global minimum which minimizes a norm-like function that interpolates between weighted $\ell_1$ and $\ell_2$ norms in the transformed input space. Our theorems subsume existing results in the literature while removing standard convergence assumptions. We also provide experiments that corroborate our analysis.
Forward citations
Cited by 3 Pith papers
-
A Classical View on Benign Overfitting: The Role of Sample Size
The paper proves high-probability, non-asymptotic bounds showing that kernel ridge regression and two-layer ReLU networks in the NTK regime can achieve both arbitrarily small training and test error without assuming t...
-
The late-stage training dynamics of (stochastic) subgradient descent on homogeneous neural networks
Normalized stochastic subgradient descent iterates converge, after perfect classification, to critical points of the normalized margin for homogeneous neural networks.
-
Understanding Nonlinear Implicit Bias via Region Counts in Input Space
Region count, the number of connected same-label regions along random input-space lines, correlates strongly with the generalization gap and is proposed as a reparameterization-invariant measure of nonlinear implicit bias.
Discussion (0). Continue with ORCID to comment.