Pith. sign in

REVIEW 4 cited by

Double Descent Demystified: Identifying, Interpreting & Ablating the Sources of a Deep Learning Puzzle

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.14151 v1 pith:7ZW7PNAP submitted 2023-03-24 cs.LG stat.ML

classification cs.LGstat.ML
keywords descentdoubledatalearningnumberlinearmodelsregression
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Double descent is a surprising phenomenon in machine learning, in which as the number of model parameters grows relative to the number of data, test error drops as models grow ever larger into the highly overparameterized (data undersampled) regime. This drop in test error flies against classical learning theory on overfitting and has arguably underpinned the success of large models in machine learning. This non-monotonic behavior of test loss depends on the number of data, the dimensionality of the data and the number of model parameters. Here, we briefly describe double descent, then provide an explanation of why double descent occurs in an informal and approachable manner, requiring only familiarity with linear algebra and introductory probability. We provide visual intuition using polynomial regression, then mathematically analyze double descent with ordinary linear regression and identify three interpretable factors that, when simultaneously all present, together create double descent. We demonstrate that double descent occurs on real data when using ordinary linear regression, then demonstrate that double descent does not occur when any of the three factors are ablated. We use this understanding to shed light on recent observations in nonlinear models concerning superposition and double descent. Code is publicly available.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Deep learning inference with the Event Horizon Telescope II. The Zingularity framework for Bayesian artificial neural networks

    astro-ph.IM 2025-06 conditional novelty 5.0 of 10

    Bayesian neural networks trained on synthetic EHT observations of Sgr A* and M87* recover spin and magnetic state well in cross-code tests, but give overconfident wrong estimates for temperature ratio and inclination ...

  2. Black hole/quantum machine learning correspondence

    quant-ph 2025-06 conditional novelty 5.0 of 10

    The authors identify the black hole Page time with the interpolation threshold of quantum linear regression, linking information recovery to double descent through the Marchenko-Pastur law.

  3. Large Language Models and Emergence: A Complex Systems Perspective

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A perspective paper arguing that LLM emergence claims are incomplete without evidence of internal coarse-grained representations, and that LLMs have not shown emergent intelligence.

  4. Approach to Finding a Robust Deep Learning Model

    cs.LG 2025-05 conditional novelty 4.0 of 10

    A robustness measure based on the spread of test losses across independently trained instances, plus a pruning algorithm, selects stable small CNNs for calorimeter energy and position reconstruction.

Pith tools