Pith. sign in

REVIEW 6 cited by

The Principles of Deep Learning Theory

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.10165 v2 pith:HQEJJINJ submitted 2021-06-18 cs.LG cs.AIhep-thstat.ML

classification cs.LGcs.AIhep-thstat.ML
keywords networkslearningexplainflownetworkratioaspectdeep
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This book develops an effective theory approach to understanding deep neural networks of practical relevance. Beginning from a first-principles component-level picture of networks, we explain how to determine an accurate description of the output of trained networks by solving layer-to-layer iteration equations and nonlinear learning dynamics. A main result is that the predictions of networks are described by nearly-Gaussian distributions, with the depth-to-width aspect ratio of the network controlling the deviations from the infinite-width Gaussian description. We explain how these effectively-deep networks learn nontrivial representations from training and more broadly analyze the mechanism of representation learning for nonlinear models. From a nearly-kernel-methods perspective, we find that the dependence of such models' predictions on the underlying learning algorithm can be expressed in a simple and universal way. To obtain these results, we develop the notion of representation group flow (RG flow) to characterize the propagation of signals through the network. By tuning networks to criticality, we give a practical solution to the exploding and vanishing gradient problem. We further explain how RG flow leads to near-universal behavior and lets us categorize networks built from different activation functions into universality classes. Altogether, we show that the depth-to-width ratio governs the effective model complexity of the ensemble of trained networks. By using information-theoretic techniques, we estimate the optimal aspect ratio at which we expect the network to be practically most useful and show how residual connections can be used to push this scale to arbitrary depths. With these tools, we can learn in detail about the inductive bias of architectures, hyperparameters, and optimizers.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Statistics of correlations in nonlinear recurrent neural networks

    q-bio.NC 2025-10 unverdicted novelty 7.0 of 10

    Derives exact correlation statistics for nonlinear RNNs in the large-N limit with Gaussian quenched disorder using path integrals, generalizing linear results and adding 1/N corrections.

  2. Interpreting learning dynamics of autoencoders: Transient scaling and emerging concepts of the Ising model

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Unsupervised autoencoders on Ising configurations form magnetization then energy representations in two dynamical regimes, with recursive error flow fields sharing topology across layers.

  3. Optimized Weight Initialization on the Stiefel Manifold for Deep ReLU Neural Networks

    cs.LG 2025-08 conditional novelty 6.0 of 10

    An initialization that maximizes a semi-orthogonal weight matrix's alignment with the all-ones vector prevents dying ReLU and keeps 100-layer ReLU networks trainable.

  4. Pre-Strings Lectures on Artificial Intelligence

    hep-th 2026-07 accept novelty 5.5 of 10

    Lecture notes define neural-network field theory and survey how it recovers known QFT/string results plus applied AI techniques for string problems.

  5. Criticality analysis of nuclear binding energy neural networks

    nucl-th 2025-08 conditional novelty 5.0 of 10

    On a two-input nuclear binding energy network, the paper validates ANNFT predictions for variance, kurtosis, and an optimal depth-to-width ratio r*=0.034 under SGD, while adaptive optimizers obscure criticality.

  6. Machine Learning is Good for Physics - and Vice Versa

    hep-ph 2026-08 unverdicted novelty 3.0 of 10

    A perspective essay arguing that AI should be integrated into fundamental physics while preserving the field's statistical and theory-based standards, and that physics can enrich machine learning.

Pith tools