Pith. sign in

REVIEW 3 cited by

Scaling Laws for Deep Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2108.07686 v1 pith:XIFD35E2 submitted 2021-08-17 cs.LG

classification cs.LG
keywords lawsscalingerrorlearningsourcescasedeepfield
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Running faster will only get you so far -- it is generally advisable to first understand where the roads lead, then get a car ... The renaissance of machine learning (ML) and deep learning (DL) over the last decade is accompanied by an unscalable computational cost, limiting its advancement and weighing on the field in practice. In this thesis we take a systematic approach to address the algorithmic and methodological limitations at the root of these costs. We first demonstrate that DL training and pruning are predictable and governed by scaling laws -- for state of the art models and tasks, spanning image classification and language modeling, as well as for state of the art model compression via iterative pruning. Predictability, via the establishment of these scaling laws, provides the path for principled design and trade-off reasoning, currently largely lacking in the field. We then continue to analyze the sources of the scaling laws, offering an approximation-theoretic view and showing through the exploration of a noiseless realizable case that DL is in fact dominated by error sources very far from the lower error limit. We conclude by building on the gained theoretical understanding of the scaling laws' origins. We present a conjectural path to eliminate one of the current dominant error sources -- through a data bandwidth limiting hypothesis and the introduction of Nyquist learners -- which can, in principle, reach the generalization error lower limit (e.g. 0 in the noiseless case), at finite dataset size.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Bayesian Neural Scaling Law Extrapolation with Prior-Data Fitted Networks

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A Prior-data Fitted Network with a scaling-law-specific prior gives better point and uncertainty predictions for neural scaling law extrapolation than MCMC, BNSL, and LC-PFN baselines.

  2. Phase Transitions in Large Language Models and the $O(N)$ Model

    cs.LG 2025-01 reject novelty 5.0 of 10

    A proposed O(N)-model reformulation of Transformers yields fitted estimates of an internal dimension near 6 and a claimed capability-emergence threshold near 7B parameters, though both rest on fragile fits.

  3. Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI

    cs.CY 2024-12 conditional novelty 4.0 of 10

    The paper proposes the AI-45 degree law, a Causal Ladder framework, and five trustworthiness levels as a roadmap toward trustworthy AGI.

Pith tools