Pith. sign in

REVIEW 2 cited by

The Influence of Learning Rule on Representation Dynamics in Wide Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.02157 v2 pith:U2Z3YZPT submitted 2022-10-05 stat.ML cond-mat.dis-nncond-mat.stat-mechcs.LG

classification stat.MLcond-mat.dis-nncond-mat.stat-mechcs.LG
keywords learningnetworksalignmentdynamicsfeedbackhebbkernellazy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

It is unclear how changing the learning rule of a deep neural network alters its learning dynamics and representations. To gain insight into the relationship between learned features, function approximation, and the learning rule, we analyze infinite-width deep networks trained with gradient descent (GD) and biologically-plausible alternatives including feedback alignment (FA), direct feedback alignment (DFA), and error modulated Hebbian learning (Hebb), as well as gated linear networks (GLN). We show that, for each of these learning rules, the evolution of the output function at infinite width is governed by a time varying effective neural tangent kernel (eNTK). In the lazy training limit, this eNTK is static and does not evolve, while in the rich mean-field regime this kernel's evolution can be determined self-consistently with dynamical mean field theory (DMFT). This DMFT enables comparisons of the feature and prediction dynamics induced by each of these learning rules. In the lazy limit, we find that DFA and Hebb can only learn using the last layer features, while full FA can utilize earlier layers with a scale determined by the initial correlation between feedforward and feedback weight matrices. In the rich regime, DFA and FA utilize a temporally evolving and depth-dependent NTK. Counterintuitively, we find that FA networks trained in the rich regime exhibit more feature learning if initialized with smaller correlation between the forward and backward pass weights. GLNs admit a very simple formula for their lazy limit kernel and preserve conditional Gaussianity of their preactivations under gating functions. Error modulated Hebb rules show very small task-relevant alignment of their kernels and perform most task relevant learning in the last layer.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Balancing structure and randomness: maximum entropy networks for context-dependent computations

    q-bio.NC 2026-05 unverdicted novelty 7.0 of 10

    Maximum entropy inference on weight distributions under context-dependent task constraints produces neuron populations with contextual gain modulation whose connectivity matches gradient-descent trained networks, with...

  2. Can Biologically Plausible Temporal Credit Assignment Rules Match BPTT for Neural Similarity? E-prop as an Example

    cs.NE 2025-06 conditional novelty 6.0 of 10

    At matched task accuracy, e-prop trained RNNs reach neural data similarity comparable to BPTT trained RNNs on Mante 2013 and Sussillo 2015 datasets, with initialization and architecture influencing similarity more tha...

Pith tools