Pith. sign in

REVIEW 10 cited by

Towards Understanding the Spectral Bias of Deep Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1912.01198 v3 pith:ACUFIXYO submitted 2019-12-03 cs.LG stat.ML

Towards Understanding the Spectral Bias of Deep Learning

classification cs.LG stat.ML
keywords neuralbiasspectralnetworkslearningcertaindataexplanation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

An intriguing phenomenon observed during training neural networks is the spectral bias, which states that neural networks are biased towards learning less complex functions. The priority of learning functions with low complexity might be at the core of explaining generalization ability of neural network, and certain efforts have been made to provide theoretical explanation for spectral bias. However, there is still no satisfying theoretical result justifying the underlying mechanism of spectral bias. In this paper, we give a comprehensive and rigorous explanation for spectral bias and relate it with the neural tangent kernel function proposed in recent work. We prove that the training process of neural networks can be decomposed along different directions defined by the eigenfunctions of the neural tangent kernel, where each direction has its own convergence rate and the rate is determined by the corresponding eigenvalue. We then provide a case study when the input data is uniformly distributed over the unit sphere, and show that lower degree spherical harmonics are easier to be learned by over-parameterized neural networks. Finally, we provide numerical experiments to demonstrate the correctness of our theory. Our experimental results also show that our theory can tolerate certain model misspecification in terms of the input data distribution.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Neural Spectral Bias and Conformal Correlators I: Introduction and Applications

    hep-th 2026-04 unverdicted novelty 8.0

    Neural networks optimized solely on crossing symmetry reconstruct CFT correlators from minimal input data to few-percent accuracy across generalized free fields, minimal models, Ising, N=4 SYM, and AdS diagrams.

  2. Frequency Shift Physics-Informed Extreme Learning Machine for Solving High-Frequency Partial Differential Equations

    cs.LG 2026-07 unverdicted novelty 7.0

    FS-PIELM shifts the mean of Gaussian weights (variance fixed at 1) in PIELM to bound frequency variance and achieve 1-5 orders of magnitude better accuracy on high-frequency PDE benchmarks while retaining single linea...

  3. Deciphering Neural Reparameterized Full-Waveform Inversion with Neural Sensitivity Kernel and Wave Tangent Kernel

    physics.geo-ph 2026-05 unverdicted novelty 7.0

    Neural tangent kernel from neural reparameterization modulates sensitivity and wave tangent kernels to produce spectral filtering, wavenumber modulation, and frequency bias that improve NeurFWI convergence.

  4. The Global Empirical NTK: Self-Referential Bias and Dimensionality of Gradient Descent Learning

    cs.LG 2026-05 unverdicted novelty 7.0

    The global empirical NTK for finite-width networks has a universal Kronecker-core form that makes it structurally low-rank and biases gradient descent toward dominant modes of joint input-hidden activity.

  5. A Theory on Flow Matching with Neural Networks

    cs.LG 2026-06 unverdicted novelty 6.0

    Establishes convergence guarantees for overparameterized 2-layer ReLU networks in flow matching, generalization bounds for the velocity-field objective, and Wasserstein guarantees for generated samples, using multi-ta...

  6. Fourier Feature Pyramids for Physics-Informed Neural Networks

    cs.LG 2026-05 unverdicted novelty 6.0

    beignet replaces random Fourier feature embeddings in PINNs with a trainable multi-resolution Fourier feature pyramid, achieving higher accuracy on PDE benchmarks with fewer parameters and near machine precision resid...

  7. State-Space NTK Collapse Near Bifurcations

    cs.LG 2026-05 unverdicted novelty 6.0

    Bifurcations cause sNTK to reduce to a dominant rank-one channel matching normal forms, collapsing effective rank and funneling gradient descent into critical dynamical directions.

  8. Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

    cs.LG 2024-01 unverdicted novelty 6.0

    SPIN lets weak LLMs become strong by self-generating training data from previous model versions and training to prefer human-annotated responses over its own outputs, outperforming DPO even with extra GPT-4 data on be...

  9. SpiS-GAN: Spiral-Modulated Handwriting Synthesis with Star Operation

    cs.CV 2026-07 conditional novelty 5.0

    A GAN with elliptical-spiral feature mixing, star-operation blocks, and Sobel edge loss produces more realistic synthetic handwriting and lowers HTR error rates on English and Vietnamese datasets.

  10. Spectral methods: crucial for machine learning, natural for quantum computers?

    quant-ph 2026-03 unverdicted novelty 5.0

    Quantum computers may enable more natural manipulation of Fourier spectra in ML models via the Quantum Fourier Transform, potentially leading to resource-efficient spectral methods.