REVIEW 10 cited by
Deep Double Descent: Where Bigger Models and More Data Hurt
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We show that a variety of modern deep learning tasks exhibit a "double-descent" phenomenon where, as we increase model size, performance first gets worse and then gets better. Moreover, we show that double descent occurs not just as a function of model size, but also as a function of the number of training epochs. We unify the above phenomena by defining a new complexity measure we call the effective model complexity and conjecture a generalized double descent with respect to this measure. Furthermore, our notion of model complexity allows us to identify certain regimes where increasing (even quadrupling) the number of train samples actually hurts test performance.
Forward citations
Cited by 10 Pith papers
-
Impact of Bottleneck Layers and Skip Connections on the Generalization of Linear Denoising Autoencoders
Two-layer linear denoising autoencoders show a bias-variance trade-off in bottleneck width, and skip connections reduce variance near the interpolation peak.
-
Breaking the Simplification Bottleneck in Amortized Neural Symbolic Regression
A fast hash-based simplification engine enables a transformer-based symbolic regression system to train on 512M simplified expressions and match PySR on the FastSRB benchmark with better parsimony scaling.
-
On Spectral Properties of Gradient-based Explanation Methods
Gradient-based explanations behave like frequency-band selectors: the gradient acts as a high-pass filter, perturbation as a low-pass filter, and their combination creates explanations that shift with the perturbation scale.
-
Detecting AI Assistance in Abstract Complex Tasks
Converting behavioral search traces into image channels plus an exploration/exploitation time series lets a small CNN-RNN detect AI assistance with about 86% accuracy on a balanced lab task.
-
How much do language models memorize?
A compression-based measurement puts GPT-style model memorization capacity at roughly 3.6 bits per parameter, with membership inference success following a sigmoid in the dataset-to-capacity ratio.
-
Double Descent and Overparameterization in Particle Physics Data
Double descent, a test-error peak followed by recovery at high capacity, appears in jet regression and event classification on ATLAS open data; with early stopping, overparameterized models can beat classical ones.
-
BlueGlass: A Framework for Composite AI Safety
BlueGlass provides composite AI safety infrastructure; its case studies on object-detection VLMs reveal dataset trade-offs, a decoder-layer phase transition in probe accuracy, and SAE-discovered concepts including spu...
-
Does Order Matter : Connecting The Law of Robustness to Robust Generalization
The paper proves R(ℓρ∘B_L∘S) ≤ 8R(B_L∘S) but does not derive the advertised Ω(n^{1/d}) recovery or the missing local-scale result.
-
Optimizers Qualitatively Alter Solutions And We Should Leverage This
Deep learning optimizers should be designed to induce desired solution properties, not just convergence speed; different optimizers demonstrably land in qualitatively different minima.
-
Statistical Machine Learning for Astronomy -- A Textbook
A systematic Bayesian-first textbook that derives classical and modern machine learning methods for astronomy from probability theory, explicitly without new research results.
Discussion (0). Continue with ORCID to comment.