REVIEW 9 cited by
Neural Networks as Kernel Learners: The Silent Alignment Effect
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Neural networks in the lazy training regime converge to kernel machines. Can neural networks in the rich feature learning regime learn a kernel machine with a data-dependent kernel? We demonstrate that this can indeed happen due to a phenomenon we term silent alignment, which requires that the tangent kernel of a network evolves in eigenstructure while small and before the loss appreciably decreases, and grows only in overall scale afterwards. We show that such an effect takes place in homogenous neural networks with small initialization and whitened data. We provide an analytical treatment of this effect in the linear network case. In general, we find that the kernel develops a low-rank contribution in the early phase of training, and then evolves in overall scale, yielding a function equivalent to a kernel regression solution with the final network's tangent kernel. The early spectral learning of the kernel depends on the depth. We also demonstrate that non-whitened data can weaken the silent alignment effect.
Forward citations
Cited by 9 Pith papers
-
A Defense of the Quadratic Model
Local Taylor-expanded quadratic models reproduce a 150M-parameter LLM's validation loss for up to 10% of training late in the run, and LLM pretraining operates within a factor of 2 of a stochastic or deterministic edg...
-
Implicit Bias of SGD in Multivariate ReLU Networks: Effective Width Collapse
Noisy SGD in the mean-field regime forces wide multivariate ReLU networks to an effective width of at most 2P-1, yielding a continuous piecewise-affine predictor whose hyperplanes are non-redundant with respect to the...
-
Universal One-third Time Scaling in Learning Peaked Distributions
Softmax + cross-entropy on peaked targets yields loss ∼ t^{−1/3}, giving an architecture-driven explanation for power-law LLM training time without power-law data.
-
Adaptive kernel predictors from feature-learning infinite limits of neural networks
Feature-learning infinite-width neural networks are kernel machines with data-dependent kernels, defined by a min-max saddle point (Bayesian/Langevin) or a DMFT fixed point (gradient flow with weight decay).
-
Parameter Symmetry Potentially Unifies Deep Learning Theory
This position paper argues that parameter symmetry breaking and restoration unify three hierarchies in deep learning: learning dynamics, model complexity, and representation formation.
-
Dataset Distillation by Influence Matching
Inf-Match distills datasets by matching estimated parameter influence of real and synthetic data, reporting SOTA classification and retrieval, but with an unsupported theoretical core.
-
Unifying Learning Dynamics and Generalization in Transformers Scaling Law
Claims a two-stage transformer scaling law (exponential then C^{-1/6}) with matching bounds, but the lower bounds are missing, the exponent is inconsistent (-1/7 vs -1/6), and the law is an artifact of hand-set M = Θ(...
-
Beyond Scaling Curves: Internal Dynamics of Neural Networks Through the NTK Lens
Using NTK trace and effective rank, this paper shows that model and data scaling improve test loss at similar rates but drive internal dynamics in opposite directions, and estimates a feature-learning width limit well...
-
Random at First, Fast at Last: NTK-Guided Fourier Pre-Processing for Tabular DL
Fixed random Fourier projections on tabular inputs are claimed to bound the NTK, speed up gradient descent, and improve accuracy across four architectures and eight benchmarks.
Discussion (0). Continue with ORCID to comment.