Pith. sign in

REVIEW 1 cited by

A theory of data variability in Neural Network Bayesian inference

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.16695 v2 pith:CWOWQN5J submitted 2023-07-31 cond-mat.dis-nn cs.LGstat.ML

classification cond-mat.dis-nncs.LGstat.ML
keywords datapropertiesgeneralizationkernelvariabilityinferenceinfinitelylearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Bayesian inference and kernel methods are well established in machine learning. The neural network Gaussian process in particular provides a concept to investigate neural networks in the limit of infinitely wide hidden layers by using kernel and inference methods. Here we build upon this limit and provide a field-theoretic formalism which covers the generalization properties of infinitely wide networks. We systematically compute generalization properties of linear, non-linear, and deep non-linear networks for kernel matrices with heterogeneous entries. In contrast to currently employed spectral methods we derive the generalization properties from the statistical properties of the input, elucidating the interplay of input dimensionality, size of the training data set, and variability of the data. We show that data variability leads to a non-Gaussian action reminiscent of a ($\varphi^3+\varphi^4$)-theory. Using our formalism on a synthetic task and on MNIST we obtain a homogeneous kernel matrix approximation for the learning curve as well as corrections due to data variability which allow the estimation of the generalization properties and exact results for the bounds of the learning curves in the case of infinitely many training data points.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Kernels to Features: A Multi-Scale Adaptive Theory of Feature Learning

    cond-mat.dis-nn 2025-02 conditional novelty 7.0 of 10

    A multi-scale adaptive theory shows that kernel rescaling and directional feature adaptation are two approximations of the same posterior distribution, with differences appearing in output covariances and in non-linea...

Pith tools