REVIEW 6 cited by
Machine Learning Model Sizes and the Parameter Gap
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We study trends in model size of notable machine learning systems over time using a curated dataset. From 1950 to 2018, model size in language models increased steadily by seven orders of magnitude. The trend then accelerated, with model size increasing by another five orders of magnitude in just 4 years from 2018 to 2022. Vision models grew at a more constant pace, totaling 7 orders of magnitude of growth between 1950 and 2022. We also identify that, since 2020, there have been many language models below 20B parameters, many models above 70B parameters, but a scarcity of models in the 20-70B parameter range. We refer to that scarcity as the parameter gap. We provide some stylized facts about the parameter gap and propose a few hypotheses to explain it. The explanations we favor are: (a) increasing model size beyond 20B parameters requires adopting different parallelism techniques, which makes mid-sized models less cost-effective, (b) GPT-3 was one order of magnitude larger than previous language models, and researchers afterwards primarily experimented with bigger models to outperform it. While these dynamics likely exist, and we believe they play some role in generating the gap, we don't have high confidence that there are no other, more important dynamics at play.
Forward citations
Cited by 6 Pith papers
-
400-Gbps/$\lambda$ Ultrafast Silicon Microring Modulator for Scalable Optical Compute Interconnects
A heavily-doped narrow-trench silicon microring modulator demonstrates open-eye 400 Gbps PAM6, 360 Gbps PAM4, and 200 Gbps NRZ, plus a 0.97 fJ/bit bias-free 32 Gbps mode.
-
On the Performance of Concept Probing: The Influence of the Data (Extended Version)
A systematic empirical study shows concept probes need surprisingly little data for task-relevant concepts, tolerate data reuse and moderate label noise, and benefit slightly from larger probed models.
-
GeFL: Model-Agnostic Federated Learning with Generative Models
Generative model-aided federated learning (GeFL) enables model-heterogeneous FL by sharing a federated generator, and its feature-level version GeFL-F improves scalability and privacy.
-
Revisiting Weight Averaging for Model Merging
Centering task vectors around the weight average and keeping their top singular vectors yields a merged multi-task model that outperforms prior merging methods on vision and NLP benchmarks.
-
Foundational values for foundation models
A Socratic analysis of research values produces a network of reasons for using or abstaining from foundation models in medical imaging.
-
Time-multiplexed layer reuse for physical neural networks
ReLaX-Net cycles a small set of fixed weight matrices to deepen physical neural networks, but controlled experiments show a single repeated large layer is the best use of a fixed parameter budget.
Discussion (0). Continue with ORCID to comment.