Pith. sign in

REVIEW 1 cited by

Feature Learning and Generalization in Deep Networks with Orthogonal Weights

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.07765 v2 pith:6XTXJIGY submitted 2023-10-11 cs.LG hep-phhep-thstat.ML

classification cs.LGhep-phhep-thstat.ML
keywords networksdepthdeepwidthorthogonaltrainingweightscomparable
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Fully-connected deep neural networks with weights initialized from independent Gaussian distributions can be tuned to criticality, which prevents the exponential growth or decay of signals propagating through the network. However, such networks still exhibit fluctuations that grow linearly with the depth of the network, which may impair the training of networks with width comparable to depth. We show analytically that rectangular networks with tanh activations and weights initialized from the ensemble of orthogonal matrices have corresponding preactivation fluctuations which are independent of depth, to leading order in inverse width. Moreover, we demonstrate numerically that, at initialization, all correlators involving the neural tangent kernel (NTK) and its descendants at leading order in inverse width -- which govern the evolution of observables during training -- saturate at a depth of $\sim 20$, rather than growing without bound as in the case of Gaussian initializations. We speculate that this structure preserves finite-width feature learning while reducing overall noise, thus improving both generalization and training speed in deep networks with depth comparable to width. We provide some experimental justification by relating empirical measurements of the NTK to the superior performance of deep nonlinear orthogonal networks trained under full-batch gradient descent on the MNIST and CIFAR-10 classification tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Kernels to Features: A Multi-Scale Adaptive Theory of Feature Learning

    cond-mat.dis-nn 2025-02 conditional novelty 7.0 of 10

    A multi-scale adaptive theory shows that kernel rescaling and directional feature adaptation are two approximations of the same posterior distribution, with differences appearing in output covariances and in non-linea...

Pith tools