Pith. sign in

REVIEW 11 cited by

Deep Learning is Not So Mysterious or Different

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.02113 v2 pith:3R3VNOZV submitted 2025-03-03 cs.LG stat.ML

Deep Learning is Not So Mysterious or Different

classification cs.LG stat.ML
keywords deepgeneralizationlearningclassesdifferenthypothesismodelmysterious
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Deep neural networks are often seen as different from other model classes by defying conventional notions of generalization. Popular examples of anomalous generalization behaviour include benign overfitting, double descent, and the success of overparametrization. We argue that these phenomena are not distinct to neural networks, or particularly mysterious. Moreover, this generalization behaviour can be intuitively understood, and rigorously characterized, using long-standing generalization frameworks such as PAC-Bayes and countable hypothesis bounds. We present soft inductive biases as a key unifying principle in explaining these phenomena: rather than restricting the hypothesis space to avoid overfitting, embrace a flexible hypothesis space, with a soft preference for simpler solutions that are consistent with the data. This principle can be encoded in many model classes, and thus deep learning is not as mysterious or different from other model classes as it might seem. However, we also highlight how deep learning is relatively distinct in other ways, such as its ability for representation learning, phenomena such as mode connectivity, and its relative universality.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Pointwise Generalization in Deep Neural Networks

    cs.LG 2026-05 unverdicted novelty 7.0

    Proposes pointwise Riemannian Dimension from feature eigenvalues to derive tighter, representation-aware generalization bounds for deep networks in the nonlinear regime.

  2. Estimating Implicit Regularization in Deep Learning

    stat.ML 2026-05 unverdicted novelty 7.0

    Gradient matching empirically recovers implicit regularization effects such as l2 penalties from early stopping and dropout in neural networks.

  3. Spectral Born machines: classically trainable quantum generative models for discrete data

    quant-ph 2026-07 conditional novelty 6.0

    Spectral Born machines are Fourier-phase quantum generative models over Z_d^n that train classically via graph-spectral MMD and show reduced parameters plus apparent overfitting resistance on integer data.

  4. To select or not to select: predictively consistent priors instead of model selection

    stat.ME 2026-06 unverdicted novelty 6.0

    Predictively consistent priors let complex Bayesian models match or beat the out-of-sample performance of selected simpler models across linear, logistic, and nonlinear examples without explicit selection.

  5. The Chandra-Gaia Catalog of Counterparts: Resolving ambiguous Gaia matches to X-ray sources in the Chandra Source Catalog using Machine Learning

    astro-ph.IM 2026-06 unverdicted novelty 6.0

    A LightGBM classifier trained on NWAY Bayesian matches identifies true Chandra-Gaia counterparts for 113k X-ray sources, flags 7k ambiguous cases, and attributes half of 20k separation-only matches to chance coinciden...

  6. Overfitting has a limitation: a model-independent generalization gap bound based on R\'enyi entropy

    stat.ML 2025-05 unverdicted novelty 6.0

    A model-independent upper bound on generalization gap is established that depends solely on the Rényi entropy of the data-generating distribution for histogram-determined algorithms such as ERM.

  7. The Virtue of Sparsity in Complexity

    q-fin.GN 2026-04 unverdicted novelty 5.0

    Expanding feature capacity enables discovery of sparse priced-risk structures, so nonlinear expansions plus basis pursuit outperform ridgeless methods beyond a complexity threshold.

  8. Spectral methods: crucial for machine learning, natural for quantum computers?

    quant-ph 2026-03 unverdicted novelty 5.0

    Quantum computers may enable more natural manipulation of Fourier spectra in ML models via the Quantum Fourier Transform, potentially leading to resource-efficient spectral methods.

  9. Are We Ready for AI-Driven Discovery? AI Verification Before the Next Fundamental Physics Breakthrough

    physics.data-an 2026-07 accept novelty 4.0

    Verification of ML in fundamental physics is essential precisely when models enter statistical modeling, inference, or hypothesis testing, and is bounded by unavoidable inductive bias, sample complexity, and experimen...

  10. Statistical Properties of Training & Generalization

    stat.ML 2026-06 unverdicted novelty 2.0

    Neural scaling laws in deep learning interact with physics constraints and inductive biases beyond classical statistics.

  11. Statistical Properties of Training & Generalization

    stat.ML 2026-06 unverdicted novelty 1.0

    Review of neural scaling laws and their relation to constraints and inductive biases when applying machine learning to physics problems.