Pith. sign in

REVIEW 8 cited by

A Constructive Prediction of the Generalization Error Across Scales

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.12673 v2 pith:G4WMOLK4 submitted 2019-09-27 cs.LG cs.CLcs.CVstat.ML

classification cs.LGcs.CLcs.CVstat.ML
keywords modelformscalesacrossdataerrorgeneralizationdependency
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The dependency of the generalization error of neural networks on model and dataset size is of critical importance both in practice and for understanding the theory of neural networks. Nevertheless, the functional form of this dependency remains elusive. In this work, we present a functional form which approximates well the generalization error in practice. Capitalizing on the successful concept of model scaling (e.g., width, depth), we are able to simultaneously construct such a form and specify the exact models which can attain it across model/data scales. Our construction follows insights obtained from observations conducted over a range of model/data scales, in various model types and datasets, in vision and language tasks. We show that the form both fits the observations well across scales, and provides accurate predictions from small- to large-scale models and data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. In situ fine-tuning of in silico trained Optical Neural Networks

    cs.NE 2025-06 reject novelty 6.0 of 10

    GIFT computes a gradient-informed direction based on how the training loss gradient changes with the assumed noise level and line-searches along it in situ, improving accuracy under noise misspecification.

  2. A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A skill-graph random-walk model gives closed-form accuracy-versus-compute formulas for four reasoning strategies and connects them to training scaling.

  3. Bayesian Neural Scaling Law Extrapolation with Prior-Data Fitted Networks

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A Prior-data Fitted Network with a scaling-law-specific prior gives better point and uncertainty predictions for neural scaling law extrapolation than MCMC, BNSL, and LC-PFN baselines.

  4. Scaling Pre-training to One Hundred Billion Data for Vision Language Models

    cs.CV 2025-02 conditional novelty 6.0 of 10

    Scaling VLM pretraining from 10B to 100B image-text pairs yields saturation on standard benchmarks but large gains on cultural diversity, low-resource language retrieval, and subgroup disparity.

  5. Lightweight Dataset Pruning without Full Training via Example Difficulty and Prediction Uncertainty

    cs.LG 2025-02 conditional novelty 6.0 of 10

    A lightweight score combining prediction mean and variance, plus ratio-adaptive Beta sampling, prunes datasets early in training and reaches 60% ImageNet accuracy at 90% pruning.

  6. Data-Efficient Deep Learning: Empirical Guidelines for Training Set Size Estimation in Inertial Sensor Classification

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Classification accuracy on inertial HAR and SLR tasks follows a consistent logarithmic growth with training-set size, enabling a MAPD-based stability-point metric that often saturates far below traditional heuristics.

  7. Unifying Learning Dynamics and Generalization in Transformers Scaling Law

    cs.LG 2025-12 reject novelty 4.0 of 10

    Claims a two-stage transformer scaling law (exponential then C^{-1/6}) with matching bounds, but the lower bounds are missing, the exponent is inconsistent (-1/7 vs -1/6), and the law is an artifact of hand-set M = Θ(...

  8. Beyond Scaling Curves: Internal Dynamics of Neural Networks Through the NTK Lens

    cs.LG 2025-07 conditional novelty 4.0 of 10

    Using NTK trace and effective rank, this paper shows that model and data scaling improve test loss at similar rates but drive internal dynamics in opposite directions, and estimates a feature-learning width limit well...

Pith tools