Pith. sign in

REVIEW 2 cited by

On the Activation Function Dependence of the Spectral Bias of Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.04924 v3 pith:HE27HZUS submitted 2022-08-09 cs.LG

classification cs.LG
keywords functionactivationnetworksneuralbiasspectralrelutheory
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Neural networks are universal function approximators which are known to generalize well despite being dramatically overparameterized. We study this phenomenon from the point of view of the spectral bias of neural networks. Our contributions are two-fold. First, we provide a theoretical explanation for the spectral bias of ReLU neural networks by leveraging connections with the theory of finite element methods. Second, based upon this theory we predict that switching the activation function to a piecewise linear B-spline, namely the Hat function, will remove this spectral bias, which we verify empirically in a variety of settings. Our empirical studies also show that neural networks with the Hat activation function are trained significantly faster using stochastic gradient descent and ADAM. Combined with previous work showing that the Hat activation function also improves generalization accuracy on image classification tasks, this indicates that using the Hat activation provides significant advantages over the ReLU on certain problems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Leveraging KANs for Expedient Training of Multichannel MLPs via Preconditioning and Geometric Refinement

    cs.LG 2025-05 conditional novelty 5.0 of 10

    Training in a B-spline KAN basis is equivalent to preconditioned gradient descent on a multichannel ReLU MLP, and geometric refinement plus trainable knots accelerate and improve training.

  2. Separated-Variable Spectral Neural Networks: A Physics-Informed Learning Approach for High-Frequency PDEs

    cs.LG 2025-08 conditional novelty 4.0 of 10

    A separable Fourier-feature neural network with learnable frequencies and a three-level frequency sampler is reported to solve high-frequency PDEs with far fewer parameters than vanilla PINNs.

Pith tools