Pith. sign in

REVIEW 5 cited by

Compute Trends Across Three Eras of Machine Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.05924 v2 pith:SHRPS3F6 submitted 2022-02-11 cs.LG cs.AIcs.CY

classification cs.LGcs.AIcs.CY
keywords computelearningtrainingdeepthreedoublingerasevery
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Compute, data, and algorithmic advances are the three fundamental factors that guide the progress of modern Machine Learning (ML). In this paper we study trends in the most readily quantified factor - compute. We show that before 2010 training compute grew in line with Moore's law, doubling roughly every 20 months. Since the advent of Deep Learning in the early 2010s, the scaling of training compute has accelerated, doubling approximately every 6 months. In late 2015, a new trend emerged as firms developed large-scale ML models with 10 to 100-fold larger requirements in training compute. Based on these observations we split the history of compute in ML into three eras: the Pre Deep Learning Era, the Deep Learning Era and the Large-Scale Era. Overall, our work highlights the fast-growing compute requirements for training advanced ML systems.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Bridging Compute- and Data-Optimal Pretraining

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Pretraining loss obeys a single law in which repeated or paraphrased tokens count as η(N, data-per-parameter, expansion-ratio) fresh tokens, with total effective data saturating as derived tokens grow.

  2. Scaling Laws of Global Weather Models

    cs.LG 2026-02 conditional novelty 6.0 of 10

    Across five global weather models, validation loss follows power-law scaling, with wider architectures and larger training datasets outperforming deeper or smaller-data configurations.

  3. Polaritonic Machine Learning for Graph-based Data Analysis

    cond-mat.dis-nn 2025-07 conditional novelty 6.0 of 10

    Simulated polariton condensate lattices act as physics-based feature generators for CNNs and improve classification of cliques and asymmetries in point clouds over raw point images in three synthetic tasks.

  4. Subjective Experience in AI Systems: What Do AI Researchers and the Public Believe?

    cs.CY 2025-06 conditional novelty 6.0 of 10

    Both AI researchers and the US public see AI subjective experience as a likely reality by 2100, while disagreeing on how to treat and govern such systems.

  5. Jolting Technologies: Superexponential Acceleration in AI Capabilities and Implications for AGI

    cs.AI 2025-07 reject novelty 3.0 of 10

    The paper formalizes superexponential AI growth as a positive third derivative (a 'jolt') and claims a simulation-based detector can identify such jolts, though no empirical benchmark validation is provided.

Pith tools