Pith. sign in

REVIEW 4 cited by

Tensor Programs I: Wide Feedforward or Recurrent Neural Networks of Any Architecture are Gaussian Processes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.12478 v3 pith:TIEOIC7K submitted 2019-10-28 cs.NE cond-mat.dis-nncs.LGmath-phmath.MP

classification cs.NEcond-mat.dis-nncs.LGmath-phmath.MP
keywords networksneuralgaussianprocessdeepfeedforwardnetworknormalization
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Wide neural networks with random weights and biases are Gaussian processes, as originally observed by Neal (1995) and more recently by Lee et al. (2018) and Matthews et al. (2018) for deep fully-connected networks, as well as by Novak et al. (2019) and Garriga-Alonso et al. (2019) for deep convolutional networks. We show that this Neural Network-Gaussian Process correspondence surprisingly extends to all modern feedforward or recurrent neural networks composed of multilayer perceptron, RNNs (e.g. LSTMs, GRUs), (nD or graph) convolution, pooling, skip connection, attention, batch normalization, and/or layer normalization. More generally, we introduce a language for expressing neural network computations, and our result encompasses all such expressible neural networks. This work serves as a tutorial on the *tensor programs* technique formulated in Yang (2019) and elucidates the Gaussian Process results obtained there. We provide open-source implementations of the Gaussian Process kernels of simple RNN, GRU, transformer, and batchnorm+ReLU network at github.com/thegregyang/GP4A.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Cost of Discretization in Functional Linear Regression: Minimax Rates and Adaptation

    math.ST 2026-07 accept novelty 7.0 of 10

    Matching minimax prediction rates for discretely observed functional linear regression are n^{-ν/(ν+1)}+(nm)^{-ν/κ} under independent design, and those two terms plus m^{-ν}+m^{-4α} under common design.

  2. On the Stability of the Jacobian Matrix in Deep Neural Networks

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A general random matrix stability theorem shows when sparse or correlated neural-network weights still preserve stable Jacobians, and derives the required weight-rescaling.

  3. Pre-Strings Lectures on Artificial Intelligence

    hep-th 2026-07 accept novelty 5.5 of 10

    Lecture notes define neural-network field theory and survey how it recovers known QFT/string results plus applied AI techniques for string problems.

  4. Reduced Order Models and Conditional Expectation -- Analysing Parametric Low-Order Approximations

    cs.LG 2024-12 conditional novelty 4.0 of 10

    Parametric reduced-order models built by least-squares projection, including POD, reduced basis methods, and Gaussian process emulation, can be viewed as conditional expectations in a Bayesian updating framework.

Pith tools