REVIEW 4 major objections 5 minor 1 cited by
Compute Requirements for Algorithmic Innovation in Frontier AI Models
T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Even strict compute caps would still leave room for half of AI's algorithmic innovations, a catalog of 36 techniques suggests.
desk verdict First useful dataset on compute costs for algorithmic innovations, but the 'half of innovations' count does not establish the policy conclusion about progress. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is a hand-built catalog of 36 pre-training algorithmic innovations used in Llama 3 and DeepSeek-V3, each with an estimated total FLOP and hardware capacity in TFLOP/s drawn from the experiments in its original paper. The counterfactual analysis plots the cumulative fraction of innovations whose estimated requirements fall below a given FLOP cap or hardware cap, using GPT-2 training compute and H100 counts as reference points. The distinction between total operations and hardware capacity matters because innovations like parallelization techniques may use negligible FLOP but require large clusters, so the two cap types are analyzed separately. Roughly 25 percent of the cataloged innovations used negligible FLOP.
What would settle it
A credible test would take a sample of the 36 innovations and compare the reported paper compute against audited internal development records, including failed runs; if the true development compute is consistently several times higher, the caps' allowed fraction would drop below half.
Extended reading notes
Core claim
The central claim is that compute caps alone are unlikely to dramatically slow AI algorithmic progress. The paper builds a dataset of 36 pre-training innovations from two open frontier model families, assigns each an estimated development FLOP and TFLOP/s based on the experiments in its original paper, and shows that non-negligible innovations grow at roughly 2.5 times per year in FLOP and 2.1 times per year in hardware capacity. Under counterfactual caps at GPT-2-level FLOP or 8 H100s, about half of the cataloged innovations still fall below the cap. The author reads this as evidence that algorithmic progress has a low near-term compute threshold, so restricting compute alone is unlikely to stop or dramatically slow the development of new pretraining algorithms.
Load-bearing premise
The analysis assumes the compute reported in each original paper is close to the compute actually needed to develop that innovation, even though failed experiments and validation at scale are omitted from the papers.
Editorial extensions
If this is right
- If the estimates hold, regulators cannot rely on compute caps as a standalone brake on algorithmic progress; half of recent innovations would survive even very tight caps.
- Export controls that limit rival states to modest hardware would probably not stop algorithmic improvement, since existing hardware is already sufficient for many innovations.
- The doubling trend implies that caps set at fixed 2025-style levels may become increasingly binding over time, but until then most innovations remain below them.
- Innovations that save compute at training time remain discoverable under caps precisely because their development does not require large training runs.
- Compute governance appears more credible when paired with legal or institutional measures, as the paper itself concludes.
Reading between the lines
- The paper's estimates are lower bounds on real R&D compute because failed experiments and scale-up validation are excluded; a systematic calibration of reported versus true compute could shift the cap curve downward.
- Under enforced scarcity, researchers would likely shift toward low-FLOP innovations, so the realized innovation rate under caps might be higher than the static catalog suggests.
- The catalog covers only open pre-training innovations; post-training and closed-lab innovations could have different compute profiles, so the policy conclusion may not transfer to those domains.
- A natural testable extension is to apply the same catalog method to post-training and reasoning innovations, where compute requirements are currently lower but rising.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper catalogs 36 pre-training algorithmic innovations used in Llama 3 and DeepSeek-V3, estimates for each the total FLOP and hardware capacity (TFLOP/s) used in the introducing paper, and analyzes how these estimates evolve over time. The authors report that non-negligible-FLOP innovations have grown at roughly 2.53x per year in total operations and 2.14x per year in hardware capacity. They then use the catalog to estimate what fraction of innovations would remain available under compute caps, concluding that even stringent caps (GPT-2-scale FLOP or 8 H100s) would still allow about half of the cataloged innovations, and hence that compute caps alone are unlikely to dramatically slow AI algorithmic progress. The paper is transparent about limitations, including omitted failed experiments, exclusion of proprietary labs, and the absence of impact weighting.
Significance. If the central conclusion were established, this would be a valuable contribution to compute-governance debates: it provides the first systematic catalog of compute requirements for algorithmic innovations in open frontier models, along with concrete trend estimates and a falsifiable 2028 projection. The Appendix A dataset is transparent and will be useful for future work on algorithmic progress. The paper is also commendably candid in Section 5.1 about the many ways its estimates could be biased. However, the headline policy inference moves from an unweighted count of innovations to a claim about the rate of algorithmic progress, and the manuscript's own Section 5.2 defers impact weighting to future work. That gap is load-bearing, so the central claim is not yet supported as stated.
major comments (4)
- [§4, Abstract] The headline conclusion equates the number of cataloged innovations below a cap with the rate of algorithmic progress. Figure 3 reports cumulative counts, but no impact weighting is provided; Section 5.2 explicitly lists 'Quantifying Innovation Impact (CEG)' as future work. The blocked set can plausibly contain the highest-impact innovations: under a GPT-2-scale FLOP cap, the blocked set includes Chinchilla scaling laws, DeepSeekMoE, MLA, and FP8-LM, and under an 8-H100 hardware cap it includes ZeRO, tensor parallelism, and FSDP. If these carry most of the compute-equivalent gain, progress could slow dramatically even though half the count is available. The count statistic is therefore insufficient for the progress-based conclusion; the text should either provide an impact-weighted analysis or restrict the claim to 'half of the cataloged innovations.'
- [§5.1, §4] The counterfactual analysis treats the compute reported in the introducing paper as the compute required to develop the innovation. Section 5.1 acknowledges that reported costs omit failed experiments and preliminary explorations, and that validation at scale may require much more compute, making the estimates lower bounds on actual R&D compute. The Section 4 caveat that researchers could be more efficient under a cap is a different bias and does not cancel this one. If actual development compute is systematically higher than reported, caps would block more than half of the innovations. The paper should bound or quantify this bias, for example through the researcher interviews suggested in Section 5.2 or through sensitivity analysis that inflates reported compute by plausible factors.
- [Table 2, Figure 3] The cap fractions are deterministic functions of point estimates in Table 2, yet those estimates are reported to several significant figures with no uncertainty and come from heterogeneous sources, including personal correspondence. Several entries have missing TFLOP/s values (marked '—'), and it is unclear how Figure 3 (bottom) treats these missing values when computing the fraction below a hardware cap. A small number of mis-estimated entries could move the 'half' result. The paper should clarify the handling of missing values and add sensitivity analysis around the point estimates, for example by showing how the cap fractions change under factor-of-2 or factor-of-10 perturbations.
- [§1, §5.1] The sample is restricted to 36 innovations in two open model families, and Section 5.1 acknowledges the exclusion of proprietary labs. Yet the Abstract and Section 4 draw conclusions about 'AI algorithmic progress' in general. Since the most compute-intensive innovations may occur in closed labs, the representativeness of the open-model catalog is load-bearing for the general policy claim. The manuscript should either justify that the open-model sample is representative of frontier algorithmic progress or explicitly restrict the conclusion to the population of open-model pre-training innovations.
minor comments (5)
- [§2] The phrase 'terraFLOP/s' should be 'teraFLOP/s'.
- [§4] The sentence 'These caps are would still be fairly low' contains a grammatical error and should read 'These caps would still be fairly low.'
- [Figure 1 caption] The caption should state explicitly that the trend line is fit only to innovations that do not use negligible FLOP, and should clarify what the shaded area represents (the text says 95% CI but does not say whether it is a confidence band for the mean or prediction interval).
- [Table 2] The 'Math equiv' column uses a star symbol that is never defined in the table caption; please define it and explain how 'Negligible' FLOP and missing TFLOP/s entries are treated in Figure 3.
- [§4] The sentence beginning 'A further implication is that US-led export controls...' goes beyond the evidence presented, since export controls target specific actors and involve different enforcement and verification dynamics than the domestic caps modeled here; consider softening this claim or adding supporting reasoning.
Circularity Check
No circularity: the cap analysis is a transparent cumulative count of the paper's own catalog, not a hidden fit or self-citation chain.
full rationale
The paper's derivation chain is self-contained and non-circular. The central empirical result — that half of the cataloged innovations fall below the GPT-2 FLOP cap and the 8-H100 hardware cap — is a direct cumulative distribution of the paper's own dataset, explicitly presented as a preliminary measure in Section 4 ('As a preliminary measure of the impact of compute caps on algorithmic progress, we calculate what fraction of the cataloged innovations fall below a given FLOP or hardware cap'). No parameter is fitted to a target and then relabeled as a prediction; the trend regressions in Section 3 are descriptive summaries and are not used to construct the cap conclusion. The paper's own limitations section candidly notes that the reported compute is a lower bound on actual R&D compute and that impact weighting (CEG) is left to future work, but those are evidentiary and interpretive gaps, not circularity. There are no load-bearing self-citations, no imported uniqueness theorems, and no ansatz smuggled in via citation. The inference from 'half of the cataloged innovations available' to 'compute caps alone are unlikely to dramatically slow AI algorithmic progress' is an extrapolation that may be debated on representativeness or impact-weighting grounds, but it is not equivalent to its inputs by construction.
Assumptions & free parameters
free parameters (2)
- FLOP growth rate (log-linear trend slope) =
2.53x per year (95% CI 1.86-3.38)
- Hardware capacity growth rate =
2.14x per year (95% CI 1.44-2.76)
assumptions (4)
- domain assumption The compute reported in the paper that introduced an innovation measures the compute required to develop it.
- domain assumption The 36 innovations used in Llama 3 and DeepSeek-V3 are a representative sample of pre-training algorithmic innovation.
- domain assumption An innovation is 'available' under a cap if its estimated development compute is below the cap, independent of other innovations and prior lineage.
- domain assumption Log-linear extrapolation of the observed trend to 2028 is valid for forecasting future median compute requirements.
Cite this review
Pith. "Pith review of Compute Requirements for Algorithmic Innovation in Frontier AI Models." pith.science (2026). https://pith.science/paper/4I2VG32M
@misc{pith2026250710618,
author = {Pith},
title = {Pith review of: Compute Requirements for Algorithmic Innovation in Frontier AI Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/4I2VG32M}},
note = {Machine review of arXiv:2507.10618}
}
read the original abstract
Algorithmic innovation in the pretraining of large language models has driven a massive reduction in the total compute required to reach a given level of capability. In this paper we empirically investigate the compute requirements for developing algorithmic innovations. We catalog 36 pre-training algorithmic innovations used in Llama 3 and DeepSeek-V3. For each innovation we estimate both the total FLOP used in development and the FLOP/s of the hardware utilized. Innovations using significant resources double in their requirements each year. We then use this dataset to investigate the effect of compute caps on innovation. Our analysis suggests that compute caps alone are unlikely to dramatically slow AI algorithmic progress. Even stringent compute caps -- such as capping total operations to the compute used to train GPT-2 or capping hardware capacity to 8 H100 GPUs -- could still have allowed for half of the cataloged innovations.
Figures
Forward citations
Cited by 1 Pith paper
-
How to Catch a GPU: A Taxonomy of Verification and Enforcement Mechanisms for International AI Agreements
Verification of international AI agreements will fail first at detecting hidden compute facilities, around the 10,000-H100-equivalent scale, before other enforcement mechanisms break.
Reference graph
Works this paper leans on
-
[1]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...
-
[2]
Aguirre, A. Keep the Future Human: Why and How We Should Close the Gates to AGI and Superintelligence, and What We Should Build Instead . arXiv preprint arXiv:2311.09452, 2025. URL https://arxiv.org/abs/2311.09452v4
arXiv 2025
-
[3]
d., Zemlyanskiy, Y., Lebron, F., and Sanghai, S
Ainslie, J., Lee-Thorp, J., Jong, M. d., Zemlyanskiy, Y., Lebron, F., and Sanghai, S. GQA : Training Generalized Multi - Query Transformer Models from Multi - Head Checkpoints . December 2023. URL https://openreview.net/forum?id=hmOwOZWzYE
work page 2023
-
[4]
Singe: leveraging warp specialization for high performance on GPUs
Bauer, M., Treichler, S., and Aiken, A. Singe: leveraging warp specialization for high performance on GPUs . SIGPLAN Not., 49 0 (8): 0 119--130, February 2014. ISSN 0362-1340. doi:10.1145/2692916.2555258. URL https://doi.org/10.1145/2692916.2555258
-
[5]
Efficient Training of Language Models to Fill in the Middle , July 2022
Bavarian, M., Jun, H., Tezak, N., Schulman, J., McLeavey, C., Tworek, J., and Chen, M. Efficient Training of Language Models to Fill in the Middle , July 2022. URL http://arxiv.org/abs/2207.14255. arXiv:2207.14255 [cs]
arXiv 2022
-
[6]
Brass, A. and Aarne, O. Location verification for ai chips. Issue brief, Institute for AI Policy and Strategy (IAPS), April 2024. URL https://www.iaps.ai/research/location-verification-for-ai-chips
work page 2024
-
[7]
Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models
Dai, D., Deng, C., Zhao, C., Xu, R., Gao, H., Chen, D., Li, J., Zeng, W., Yu, X., Wu, Y., et al. Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models. arXiv preprint arXiv:2401.06066, 2024
arXiv 2024
-
[8]
Y., Ermon, S., Rudra, A., and Ré, C
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C. FLASHATTENTION : fast and memory-efficient exact attention with IO -awareness. In Proceedings of the 36th International Conference on Neural Information Processing Systems , NIPS '22, pp.\ 16344--16359, Red Hook, NY, USA, November 2022. Curran Associates Inc. ISBN 978-1-71387-108-8
work page 2022
Show all 61 references
-
[9]
Ai capabilities can be significantly improved without expensive retraining
Davidson, T., Denain, J.-S., Villalobos, P., and Bas, G. Ai capabilities can be significantly improved without expensive retraining. arXiv preprint arXiv:2312.07413, 2023
2023 arXiv
-
[10]
DeepSeek - Coder - V2 : Breaking the Barrier of Closed - Source Models in Code Intelligence , June 2024
DeepSeek-AI, Zhu, Q., Guo, D., Shao, Z., Yang, D., Wang, P., Xu, R., Wu, Y., Li, Y., Gao, H., Ma, S., Zeng, W., Bi, X., Gu, Z., Xu, H., Dai, D., Dong, K., Zhang, L., Piao, Y., Gou, Z., Xie, Z., Hao, Z., Wang, B., Song, J., Chen, D., Xie, X., Guan, K., You, Y., Liu, A., Du, Q.,...
2024 arXiv
-
[11]
LLM .int8(): 8-bit matrix multiplication for transformers at scale
Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L. LLM .int8(): 8-bit matrix multiplication for transformers at scale. In Proceedings of the 36th International Conference on Neural Information Processing Systems , NIPS '22, pp.\ 30318--30332, Red Hook, NY, USA, November...
2022
-
[12]
A., Chhaparia, R., Donchev, Y., Kuncoro, A., Ranzato, M., Szlam, A., and Shen, J
Douillard, A., Feng, Q., Rusu, A. A., Chhaparia, R., Donchev, Y., Kuncoro, A., Ranzato, M., Szlam, A., and Shen, J. Diloco: Distributed low-communication training of language models. arXiv preprint arXiv:2311.08105, 2023
2023 arXiv
-
[13]
Data on machine learning hardware”, 10 2024 a
Epoch AI . Data on machine learning hardware”, 10 2024 a . URL https://epoch.ai/data/machine-learning-hardware. Accessed: 2025-04-28
2024
-
[14]
Data on notable ai models, 6 2024 b
Epoch AI . Data on notable ai models, 6 2024 b . URL https://epoch.ai/data/notable-ai-models. Accessed: 2025-05-11
2024
-
[15]
Y., Rozière, B., Lopez-Paz, D., and Synnaeve, G
Gloeckle, F., Idrissi, B. Y., Rozière, B., Lopez-Paz, D., and Synnaeve, G. Better & faster large language models via multi-token prediction. In Proceedings of the 41st International Conference on Machine Learning , volume 235 of ICML '24 , pp.\ 15706--15734, Vienna, Austria, J...
2024
-
[16]
The llama 3 herd of models, 2024
Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Vaughan, A., et al. The llama 3 herd of models, 2024
2024
-
[17]
H., Ivison, H., Magnusson, I., Wang, Y., et al
Groeneveld, D., Beltagy, I., Walsh, P., Bhagia, A., Kinney, R., Tafjord, O., Jha, A. H., Ivison, H., Magnusson, I., Wang, Y., et al. Olmo: Accelerating the science of language models. arXiv preprint arXiv:2402.00838, 2024
2024 arXiv
-
[18]
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025
2025 arXiv
-
[19]
and Koessler, L
Heim, L. and Koessler, L. Training compute thresholds: Features and functions in ai regulation. arXiv preprint arXiv:2405.10799, 2024
2024 arXiv
-
[20]
A., and Zilberman, N
Heim, L., Fist, T., Egan, J., Huang, S., Zekany, S., Trager, R., Osborne, M. A., and Zilberman, N. Governing through the cloud: The intermediary role of compute providers in ai regulation. arXiv preprint arXiv:2403.08501, 2024
2024 arXiv
-
[21]
C., Atkinson, D., Thompson, N., and Sevilla, J
Ho, A., Besiroglu, T., Erdil, E., Owen, D., Rahman, R., Guo, Z. C., Atkinson, D., Thompson, N., and Sevilla, J. Algorithmic progress in language models, 2024
2024
-
[22]
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., Driessche, G. v. d., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Rae, J. W., Vinyal...
2022 arXiv
-
[23]
X., Chen, D., Lee, H., Ngiam, J., Le, Q
Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, M. X., Chen, D., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., and Chen, Z. GPipe : efficient training of giant neural networks using pipeline parallelism. In Proceedings of the 33rd International Conference on Neural Information Proc...
2019
-
[24]
Openai o1 system card
Jaech, A., Kalai, A., Lerer, A., Richardson, A., El-Kishky, A., Low, A., Helyar, A., Madry, A., Beutel, A., Carney, A., et al. Openai o1 system card. arXiv preprint arXiv:2412.16720, 2024
2024 arXiv
-
[25]
M., Basra, M., Obeid, F., Straube, J., Keiblinger, M., Bakouch, E., Atkins, L., Panahi, M., Goddard, C., et al
Jaghouar, S., Ong, J. M., Basra, M., Obeid, F., Straube, J., Keiblinger, M., Bakouch, E., Atkins, L., Panahi, M., Goddard, C., et al. Intellect-1 technical report. arXiv preprint arXiv:2412.01152, 2024
2024 arXiv
-
[26]
A., Casper, J., Lym, S., McAfee, L., Andersch, M., Shoeybi, M., and Catanzaro, B
Korthikanti, V. A., Casper, J., Lym, S., McAfee, L., Andersch, M., Shoeybi, M., and Catanzaro, B. Reducing activation recomputation in large transformer models. Proceedings of Machine Learning and Systems, 5: 0 341--353, 2023
2023
-
[27]
and Richardson, J
Kudo, T. and Richardson, J. SentencePiece : A simple and language independent subword tokenizer and detokenizer for Neural Text Processing . In Blanco, E. and Lu, W. (eds.), Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing : System Demonst...
2018 doi
-
[28]
Kulp, G., Gonzales, D., Smith, E., Heim, L., Puri, P., Vermeer, M. J. D., and Winkelman, Z. Hardware-Enabled Governance Mechanisms: Developing Technical Solutions to Exempt Items Otherwise Classified Under Export Control Classification Numbers 3A090 and 4A090. RAND Corporation...
2024 doi
-
[29]
Lambert, N., Morrison, J., Pyatkin, V., Huang, S., Ivison, H., Brahman, F., Miranda, L. J. V., Liu, A., Dziri, N., Lyu, S., et al. T\"ulu 3: Pushing frontiers in open language model post-training. arXiv preprint arXiv:2411.15124, 2024
2024 arXiv
-
[30]
Breadth- First Pipeline Parallelism
Lamy-Poirier, J. Breadth- First Pipeline Parallelism . Proceedings of Machine Learning and Systems, 5: 0 48--67, March 2023. URL https://proceedings.mlsys.org/paper_files/paper/2023/hash/24e845415c1486dd2d582a9d639237f9-Abstract-mlsys2023.html
2023
-
[31]
Y., Bansal, H., Guha, E., Keh, S
Li, J., Fang, A., Smyrnis, G., Ivgi, M., Jordan, M., Gadre, S. Y., Bansal, H., Guha, E., Keh, S. S., Arora, K., et al. Datacomp-lm: In search of the next generation of training sets for language models. Advances in Neural Information Processing Systems, 37: 0 14200--14282, 2024
2024
-
[32]
DeepSeek - V2 : A Strong , Economical , and Efficient Mixture -of- Experts Language Model , 2024 a
Liu, A., Feng, B., Wang, B., Wang, B., Liu, B., Zhao, C., Dengr, C., Ruan, C., Dai, D., Guo, D., et al. DeepSeek - V2 : A Strong , Economical , and Efficient Mixture -of- Experts Language Model , 2024 a
2024
-
[33]
DeepSeek-V3 Technical Report , 2024 b
Liu, A., Feng, B., Xue, B., Wang, B., Wu, B., Lu, C., Zhao, C., Deng, C., Zhang, C., Ruan, C., et al. DeepSeek-V3 Technical Report , 2024 b
2024
-
[34]
RingAttention with Blockwise Transformers for Near - Infinite Context
Liu, H., Zaharia, M., and Abbeel, P. RingAttention with Blockwise Transformers for Near - Infinite Context . October 2023. URL https://openreview.net/forum?id=WsRHpHH4s0
2023
-
[35]
and Hutter, F
Loshchilov, I. and Hutter, F. Decoupled Weight Decay Regularization . September 2018. URL https://openreview.net/forum?id=Bkg6RiCqY7
2018
-
[36]
A Narrow Path , December 2024
Miotti, A., Bilge, T., Kasten, D., and Newport, J. A Narrow Path , December 2024. URL https://www.narrowpath.co/
2024
-
[37]
Efficient large-scale language model training on GPU clusters using megatron- LM
Narayanan, D., Shoeybi, M., Casper, J., LeGresley, P., Patwary, M., Korthikanti, V., Vainbrand, D., Kashinkunti, P., Bernauer, J., Catanzaro, B., Phanishayee, A., and Zaharia, M. Efficient large-scale language model training on GPU clusters using megatron- LM . In Proceedings ...
2021
-
[38]
8-bit Numerical Formats for Deep Neural Networks , June 2022
Noune, B., Jones, P., Justus, D., Masters, D., and Luschi, C. 8-bit Numerical Formats for Deep Neural Networks , June 2022. URL http://arxiv.org/abs/2206.02915. arXiv:2206.02915 [cs]
2022 arXiv
-
[39]
2 olmo 2 furious
OLMo Team , Walsh, P., Soldaini, L., Groeneveld, D., Lo, K., Arora, S., Bhagia, A., Gu, Y., Huang, S., Jordan, M., et al. 2 olmo 2 furious. arXiv preprint arXiv:2501.00656, 2024
2024 arXiv
-
[40]
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35: 0 27730--27744, 2022
2022
-
[41]
FP8 - LM : Training FP8 Large Language Models , December 2023
Peng, H., Wu, K., Wei, Y., Zhao, G., Yang, Y., Liu, Z., Xiong, Y., Yang, Z., Ni, B., Hu, J., Li, R., Zhang, M., Li, C., Ning, J., Wang, R., Zhang, Z., Liu, S., Chau, J., Hu, H., and Cheng, P. FP8 - LM : Training FP8 Large Language Models , December 2023. URL http://arxiv.org/a...
2023 arXiv
-
[42]
Interim report: Mechanisms for flexible hardware-enabled guarantees
Petrie, J., Aarne, O., Ammann, N., and Dalrymple, D. Interim report: Mechanisms for flexible hardware-enabled guarantees. Technical report, 8 2024
2024
-
[43]
Zero Bubble ( Almost ) Pipeline Parallelism
Qi, P., Wan, X., Huang, G., and Lin, M. Zero Bubble ( Almost ) Pipeline Parallelism . October 2023. URL https://openreview.net/forum?id=tuzTN0eIO5
2023
-
[44]
Rabe, M. N. and Staats, C. Self-attention does not need o (n^2) memory, 2021
2021
-
[45]
ZeRO : memory optimizations toward training trillion parameter models
Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y. ZeRO : memory optimizations toward training trillion parameter models. In Proceedings of the International Conference for High Performance Computing , Networking , Storage and Analysis , SC '20, pp.\ 1--16, Atlanta, Georgia, ...
2020
-
[46]
Y., Ruwase, O., Yang, S., Zhang, M., Li, D., and He, Y
Ren, J., Rajbhandari, S., Aminabadi, R. Y., Ruwase, O., Yang, S., Zhang, M., Li, D., and He, Y. ZeRO - Offload : Democratizing Billion - Scale Model Training . January 2021. URL https://openreview.net/forum?id=qXFQtGMHRa
2021
-
[47]
K., Ngo, R., Pilz, K., et al
Sastry, G., Heim, L., Belfield, H., Anderljung, M., Brundage, M., Hazell, J., O'Keefe, C., Hadfield, G. K., Ngo, R., Pilz, K., et al. Computing power and the governance of artificial intelligence. arXiv preprint arXiv:2402.08797, 2024
2024 arXiv
-
[48]
and Thiergart, L
Scher, A. and Thiergart, L. Mechanisms to Verify International Agreements About AI Development , November 2024. URL https://techgov.intelligence.org/research/mechanisms-to-verify-international-agreements-about-ai-development
2024
-
[49]
Neural Machine Translation of Rare Words with Subword Units
Sennrich, R., Haddow, B., and Birch, A. Neural Machine Translation of Rare Words with Subword Units . In Erk, K. and Smith, N. A. (eds.), Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics ( Volume 1: Long Papers ) , pp.\ 1715--1725, Berlin...
2016 doi
-
[50]
GLU Variants Improve Transformer , February 2020
Shazeer, N. GLU Variants Improve Transformer , February 2020. URL http://arxiv.org/abs/2002.05202. arXiv:2002.05202 [cs]
2020 arXiv
-
[51]
Megatron- LM : Training Multi - Billion Parameter Language Models Using Model Parallelism , March 2020
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B. Megatron- LM : Training Multi - Billion Parameter Language Models Using Model Parallelism , March 2020. URL http://arxiv.org/abs/1909.08053. arXiv:1909.08053 [cs]
2020 arXiv
-
[52]
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F. Learning to summarize with human feedback. Advances in neural information processing systems, 33: 0 3008--3021, 2020
2020
-
[53]
RoFormer : Enhanced transformer with Rotary Position Embedding
Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y. RoFormer : Enhanced transformer with Rotary Position Embedding . Neurocomputing, 568: 0 127063, February 2024. ISSN 0925-2312. doi:10.1016/j.neucom.2023.127063. URL https://www.sciencedirect.com/science/article/pii/S09252...
2024
-
[54]
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozi \`e re, B., Goyal, N., Hambro, E., Azhar, F., et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023 a
2023 arXiv
-
[55]
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023 b
2023 arXiv
-
[56]
N., Kaiser, ., and Polosukhin, I
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, ., and Polosukhin, I. Attention is all you need. volume 30, 2017
2017
-
[57]
Auxiliary- Loss - Free Load Balancing Strategy for Mixture -of- Experts
Wang, L., Gao, H., Zhao, C., Sun, X., and Dai, D. Auxiliary- Loss - Free Load Balancing Strategy for Mixture -of- Experts . October 2024. URL https://openreview.net/forum?id=y1iU5czYpE
2024
-
[58]
Ccnet: Extracting high quality monolingual datasets from web crawl data
Wenzek, G., Lachaux, M.-A., Conneau, A., Chaudhary, V., Guzm \'a n, F., Joulin, A., and Grave, E. Ccnet: Extracting high quality monolingual datasets from web crawl data. arXiv preprint arXiv:1911.00359, 2019
1911 arXiv
-
[59]
A., Oguz, B., Khabsa, M., Fang, H., Mehdad, Y., Narang, S., Malik, K., Fan, A., Bhosale, S., Edunov, S., Lewis, M., Wang, S., and Ma, H
Xiong, W., Liu, J., Molybog, I., Zhang, H., Bhargava, P., Hou, R., Martin, L., Rungta, R., Sankararaman, K. A., Oguz, B., Khabsa, M., Fang, H., Mehdad, Y., Narang, S., Malik, K., Fan, A., Bhosale, S., Edunov, S., Lewis, M., Wang, S., and Ma, H. Effective Long - Context Scaling...
2024
-
[60]
and Sennrich, R
Zhang, B. and Sennrich, R. Root mean square layer normalization. In Proceedings of the 33rd International Conference on Neural Information Processing Systems , number 1110, pp.\ 12381--12392. Curran Associates Inc., Red Hook, NY, USA, December 2019
2019
-
[61]
PyTorch FSDP : Experiences on Scaling Fully Sharded Data Parallel , September 2023
Zhao, Y., Gu, A., Varma, R., Luo, L., Huang, C.-C., Xu, M., Wright, L., Shojanazeri, H., Ott, M., Shleifer, S., Desmaison, A., Balioglu, C., Damania, P., Nguyen, B., Chauhan, G., Hao, Y., Mathews, A., and Li, S. PyTorch FSDP : Experiences on Scaling Fully Sharded Data Parallel...
2023 arXiv
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.