REVIEW 4 major objections 4 minor 88 references
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism
T0 review · 4 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read This paper claims that variance-based norm tests, extended to data and model parallel training, let language model pretraining grow batch sizes on demand, matching small-batch quality at large-batch speed, with Adam still provably…
desk verdict Real distributed variance estimator, but the empirical claims don't survive the paper's own tables on the larger models. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the norm test statistic: the ratio of an estimate of the per-sample gradient variance to the squared norm of the current batch gradient, compared against a threshold $\eta^2$. In the distributed setting this statistic is computed without per-sample gradients by measuring how much the per-worker minibatch gradients differ from the global batch gradient, a version the paper calls DDP-Norm (and FSDP-Norm under model parallelism). The proof rests on a coordinate-wise version of the exact-variance norm test, which the authors show implies the coordinate-wise expected strong growth condition that Adam's iteration-complexity analysis requires.
What would settle it
Train one small model twice with identical hyperparameters, once using the approximate worker-level variance statistic to choose batches and once using exact per-sample gradient variance, and compare the batch-size trajectories and validation losses; if the approximate schedule consistently selects notably different batch sizes or yields worse validation loss at matched sample counts, the claim that the practical test supports the convergence guarantee is refuted.
Extended reading notes
Core claim
In the paper's own terms, the discovery is that the norm test—the rule that increases the next batch size to $\lceil \| \operatorname{Var}_{i \in B_k}(\nabla \ell_i(w_k))\|_1 / (\eta^2 \|\nabla \mathcal{L}_{B_k}(w_k)\|^2) \rceil$ when the current batch gradient is too noisy—can be made practical for distributed, model-parallel training and can carry Adam to convergence. The theoretical result, Theorem 1, proves a convergence bound of order $\sqrt{K}$ up to logarithmic factors on the cumulative expected full-gradient norm under the coordinate-wise exact-variance norm test, which implies the coordinate-wise expected strong growth condition. The empirical result is that, at equal numbers of training samples or steps, the adaptive schedules match or improve validation loss relative to constant large batches and to heuristic stagewise warmups, while using fewer steps when the batch can grow. The authors state this as a general-purpose schedule applicable beyond language models, with a particular demonstration on up-to-3-billion-parameter Llama 2 family models.
Load-bearing premise
The proof assumes the exact coordinate-wise variance norm test is satisfied at every iteration, but the deployed algorithm checks a cheaper approximate statistic built from differences between workers' minibatch gradients and never verifies that the condition actually holds for the next batch.
Editorial extensions
If this is right
- Practitioners can start training with a small batch and let the method increase it only when gradient noise is small, obtaining small-batch validation quality at large-batch throughput.
- The same schedule is compatible with data and model parallelism, so memory-constrained multi-GPU setups can pretrain billion-parameter models without hand-tuning a batch warmup recipe.
- Adam's convergence under the schedule is guaranteed for smooth nonconvex objectives when the coordinate-wise norm test holds, removing the need for a global variance assumption.
- With comparable wall-clock time, the adaptive schedules produce validation losses closer to those of the best small constant batch while using far fewer gradient steps.
- The tuning effort for pretraining shifts to the single threshold $\eta$, which controls how aggressively the batch grows.
Reading between the lines
- The approximation that replaces per-sample gradient variance with between-worker minibatch variance has only $J$ samples (the number of workers); when $J$ is small the batch-size schedule will inherit high variance, so smoothing the statistic over a few iterations is a natural extension the paper does not test.
- A falsifiable check of the theory-practice link is to run a small model with exact per-sample gradient variance as the test and compare the resulting batch-size trajectory with the approximate worker-level test; close trajectories would certify the approximation that the proof needs.
- The same norm-test machinery could be coupled with sequence-length warmup or learning-rate schedules, since both affect the gradient-noise estimate and the descent direction; the paper mentions sequence length warmup but does not combine them.
- The claimed scaling-law connection between $\eta$ and the critical batch size is left open; a direct experiment varying $\eta$ while measuring the final batch size plateau would test whether the plateau tracks the critical batch size.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes adaptive batch size schedules for distributed training of language models, based on a norm test that grows the global batch when the approximate gradient variance is too large relative to the batch gradient norm. The authors present DDP-Norm and FSDP-Norm implementations and a convergence theorem for Adam under a coordinate-wise exact variance norm test. Experiments are reported on MicroLlama 300M, TinyLlama 1.1B, and OpenLlama 3B pretrained on C4, with the claim that the adaptive schedules outperform constant batch sizes and heuristic batch-size warmup schedules for Llama-family models.
Significance. If the empirical and theoretical claims were correct, this would be a useful contribution: a theoretically motivated, practically implemented method for adjusting batch sizes during large-scale distributed LLM pretraining, with an open-source FSDP implementation and reproducible experimental configuration. The paper also gives credit for shipping code and for a seriousness of purpose in attempting to connect adaptive sampling theory with distributed training practice. However, the central empirical claim is contradicted by the paper's own tables on the two larger testbeds, and the convergence theorem does not apply to the algorithm actually run. The strengths are real but are outweighed by the failure of the two headline claims.
major comments (4)
- [§5.2 Table 2; §5.3 Table 3; Abstract] The abstract's claim that the proposed approaches "outperform constant batch sizes" in the pretraining of Llama-family models is directly contradicted by the paper's own results. On TinyLlama 1.1B (Table 2), the best adaptive run (eta=0.085) reaches validation loss 4.256, while the constant batch size 4096 baseline reaches 3.817. On OpenLlama 3B (Table 3), the best adaptive run (eta=0.15) reaches 4.554, while constant 4096 reaches 3.956. On both larger models, every adaptive schedule is worse than the constant 4096 baseline, which is an internal inconsistency with the stated contribution.
- [§5.3] The wall-clock argument does not rescue the empirical claim. The authors concede in §5.3 that "using a constant batch size 4096 achieves an even lower validation loss," but the reported time savings are small: about 5% on TinyLlama (34.48h vs 32.83h) and about 6% on OpenLlama (20.75h vs 19.59h), while the validation-loss differences are large (roughly 0.44 and 0.60). Because all runs consume the same 2,000,000 training sequences, the constant batch size 4096 baseline dominates on a per-token basis, so the efficiency framing does not compensate for the worse final validation loss.
- [§4, Theorem 1; Algorithm 1; Appendix B, Remark B.1] There is a load-bearing mismatch between the convergence theorem and the implemented algorithm. Theorem 1 assumes that the coordinate-wise exact variance norm test, Ek[(∂iℒBk(wk)−∂iℒ(wk))2] ≤ η2(∂iℒ(wk))2, is satisfied at every iteration. Algorithm 1, however, implements the aggregate, non-coordinate-wise approximate test in Eq. (5) (DDP-Norm/FSDP-Norm) and does not verify the exact condition for the next batch. Appendix B Remark B.1 explicitly concedes that "the exact variance test is not implemented in practice but its approximate version instead." Consequently, the proof does not establish convergence for the algorithm whose empirical behavior is reported.
- [§4, Theorem 1; Appendix B, Theorem B.1 and definition of c2] The stated convergence guarantee is vacuous as written. Theorem 1 bounds ∑k=1K E[‖∇ℒ(wk)‖] by ~O(K). A bound of O(K) on a sum of K terms holds trivially for any algorithm with bounded per-iteration gradient norms and does not imply convergence to a stationary point. The explicit bound in Theorem B.1 is also at least linear in K: the constant c2 contains the term 2c1∑i (log(1/√β2 v0,i) − K log β2), which grows linearly in K since log β2 < 0. Thus the average gradient norm need not decay, so the theorem does not deliver the advertised convergence guarantee.
minor comments (4)
- [§3.1, Eq. (2)] The notation Vari∈B(∇ℓi(w)) is used as a vector while the displayed expression mixes norms and scalar quantities; please define the per-coordinate variance vector explicitly and distinguish it from its L1 norm.
- [§5.4] The text contains the typo "prupose" in the paragraph on the effect of η; it should read "purpose."
- [Figures 2 and 3] The batch-size panels label the horizontal axis as "sample ×105" or "sample ×106"; the unit should be "samples" and the exponent formatting made consistent across panels.
- [§4] The experiments use AdamW with decoupled weight decay, whereas Theorem 1 concerns Adam without weight decay; the paper acknowledges this in the text, but the main-text discussion would be clearer if the limitation were stated immediately after the theorem rather than in the later discussion.
Circularity Check
No significant circularity: the convergence proof is conditional on an exact norm-test assumption and does not reduce to the implemented schedule or to a fitted value.
full rationale
The paper's derivation chain is not circular. Theorem 1 (formalized as Theorem B.1) assumes the coordinate-wise exact variance norm test, Eq. (4), holds at every iteration and derives an O~(√K) bound on the sum of expected gradient norms. This assumption is a strong-growth-type noise bound; the conclusion is a different statement, so no equation reduces to an input by construction. The practical DDP-Norm/FSDP-Norm statistic, Eq. (5), is explicitly acknowledged in Remark B.1 to be an approximate version of the exact test ('the exact variance test is not implemented in practice but its approximate version instead'), so the theory-implementation gap is a soundness problem, not circularity. Self-citations to [42] for the E-SG nomenclature and [41] for related local-gradient extensions are not load-bearing: the norm test originates from Byrd et al. [12] and the proof technique is attributed to the external work [76]. Finally, the empirical claim in the abstract is weakened by the paper's own Tables 2 and 3, where constant batch size 4096 reaches lower validation loss than every adaptive schedule, and Section 5.3 concedes 'While using a constant batch size 4096 achieves an even lower validation loss'; this is an evidence/correctness concern, not a circularity. No fitted parameter is fed back into the theorem, and no uniqueness claim is imported from the authors' prior work.
Assumptions & free parameters
free parameters (4)
- norm test threshold eta =
0.15/0.2/0.25/0.275 (MicroLlama); 0.05/0.075/0.08/0.085 (TinyLlama); 0.05/0.1/0.15 (OpenLlama)
- initial global batch size =
256 (MicroLlama); 128 (TinyLlama and OpenLlama)
- maximum global batch size cap =
8192
- micro-batch configuration =
base micro batch 4, maximum micro batch 8, gradient accumulation 16
assumptions (5)
- domain assumption The loss is L-Lipschitz smooth (Assumption 1).
- ad hoc to paper The exact coordinate-wise variance norm test holds at every iteration.
- ad hoc to paper The approximate variance estimator in Eq. (5) is a valid proxy for the per-sample gradient variance.
- standard math The technical lemmas of Wang et al. [76] are correct and applicable.
- domain assumption The theoretical Adam update omits bias correction, weight decay, learning rate schedules, and gradient clipping.
Cite this review
Pith. "Pith review of Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism." pith.science (2026). https://pith.science/paper/25X6ZO74
@misc{pith2026241221124,
author = {Pith},
title = {Pith review of: Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism},
year = {2026},
howpublished = {\url{https://pith.science/paper/25X6ZO74}},
note = {Machine review of arXiv:2412.21124}
}
read the original abstract
An appropriate choice of batch sizes in large-scale model training is crucial, yet it involves an intrinsic yet inevitable dilemma: large-batch training improves training efficiency in terms of memory utilization, while generalization performance often deteriorates due to small amounts of gradient noise. Despite this dilemma, the common practice of choosing batch sizes in language model training often prioritizes training efficiency -- employing either constant large sizes with data parallelism or implementing batch size warmup schedules. However, such batch size schedule designs remain heuristic and often fail to adapt to training dynamics, presenting the challenge of designing adaptive batch size schedules. Given the abundance of available datasets and the data-hungry nature of language models, data parallelism has become an indispensable distributed training paradigm, enabling the use of larger batch sizes for gradient computation. However, vanilla data parallelism requires replicas of model parameters, gradients, and optimizer states at each worker, which prohibits training larger models with billions of parameters. To optimize memory usage, more advanced parallelism strategies must be employed. In this work, we propose general-purpose and theoretically principled adaptive batch size schedules compatible with data parallelism and model parallelism. We develop a practical implementation with PyTorch Fully Sharded Data Parallel, facilitating the pretraining of language models of different sizes. We empirically demonstrate that our proposed approaches outperform constant batch sizes and heuristic batch size warmup schedules in the pretraining of models in the Llama 2 family, with particular focus on smaller models with up to 3 billion parameters. We also establish theoretical convergence guarantees for such adaptive batch size schedules with Adam for general smooth nonconvex objectives.
Figures
Reference graph
Works this paper leans on
-
[1]
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dandelion Mané, Rajat Monga, Sherry Moore, Derek Murra...
2015
-
[2]
Takuya Akiba, Shuji Suzuki, and Keisuke Fukuda. Extremely large minibatch SGD: Training ResNet-50 on ImageNet in 15 minutes.arXiv preprint arXiv:1711.04325, 2017
arXiv 2017
-
[3]
Qwen technical report.arXiv preprint arXiv:2309.16609, 2023
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Benfeng X...
arXiv 2023
-
[4]
Coupling adaptive batch sizes with learning rates
Lukas Balles, Javier Romero, and Philipp Hennig. Coupling adaptive batch sizes with learning rates. In Proceedings of the Conference on Uncertainty in Artificial Intelligence (UAI), 2017
2017
-
[5]
Stable LM 2 1.6B technical report.arXiv preprint arXiv:2402.17834, 2024
Marco Bellagente, Jonathan Tow, Dakota Mahan, Duy Phung, Maksym Zhuravinskyi, Reshinth Adithyan, James Baicoianu, Ben Brooks, Nathan Cooper, Ashish Datta, Meng Lee, Emad Mostaque, Michael Pieler, Nikhil Pinnaparju, Paulo Rocha, Harry Saini, Hannah Teufel, Niccolo Zanichelli, and Carlos Riquelme. Stable LM 2 1.6B technical report.arXiv preprint arXiv:2402....
arXiv 2024
-
[6]
BigScience Workshop, Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Benoît Sagot, Niklas Muennighoff, Albert Villanova del Moral, Olatu...
arXiv 2022
-
[7]
Raghu Bollapragada and Stefan M. Wild. Adaptive sampling quasi-Newton methods for zeroth-order stochastic optimization. Mathematical Programming Computation, 15(2):327–364, 2023
2023
-
[8]
Adaptive sampling strategies for stochastic optimization
Raghu Bollapragada, Richard Byrd, and Jorge Nocedal. Adaptive sampling strategies for stochastic optimization. SIAM Journal on Optimization, 28(4):3312–3343, 2018
2018
Show all 88 references
-
[9]
Curtis, and Jorge Nocedal
Léon Bottou, Frank E. Curtis, and Jorge Nocedal. Optimization methods for large-scale machine learning. SIAM Review, 60(2):223–311, 2018
2018
-
[10]
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang. JAX: composable transformations of Python+NumPy programs, 2018. URLhttp://github.com/google/jax
2018
-
[11]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffr...
2020
-
[12]
Byrd, Gillian M
Richard H. Byrd, Gillian M. Chin, Jorge Nocedal, and Yuchen Wu. Sample size selection in optimization methods for machine learning.Mathematical Programming, 134(1):127–155, 2012
2012
-
[13]
Big batch SGD: Automated inference using adaptive batch sizes.arXiv preprint arXiv:1610.05792, 2016
Soham De, Abhay Yadav, David Jacobs, and Tom Goldstein. Big batch SGD: Automated inference using adaptive batch sizes.arXiv preprint arXiv:1610.05792, 2016
2016 arXiv
-
[14]
Automated inference with adaptive batches
Soham De, Abhay Yadav, David Jacobs, and Tom Goldstein. Automated inference with adaptive batches. In Proceedings of the International Conference on Artificial Intelligence and Statistics (AISTATS), 2017
2017
-
[15]
Zhang, Hanwei Xu, Hao Yang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J
DeepSeek-AI, Aixin Liu, Bei Feng, Bin Wang, Bingxuan Wang, Bo Liu, Chenggang Zhao, Chengqi Dengr, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei Li, H. Zhang, Hanwei Xu, Hao Yang, Haowei...
2024 arXiv
-
[16]
Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J
DeepSeek-AI, Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei L...
2024 arXiv
-
[17]
Adabatch: Adaptive batch sizes for training deep neural networks.arXiv preprint arXiv:1712.02029, 2017
Aditya Devarakonda, Maxim Naumov, and Michael Garland. Adabatch: Adaptive batch sizes for training deep neural networks.arXiv preprint arXiv:1712.02029, 2017
2017 arXiv
-
[18]
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at...
2021
-
[19]
Susskind, and Armand Joulin
Alaaeldin El-Nouby, Michal Klein, Shuangfei Zhai, Miguel Angel Bautista, Alexander Toshev, Vaishaal Shankar, Joshua M. Susskind, and Armand Joulin. Scalable pre-training of large autoregressive image models. In Proceedings of the International Conference on Machine Learning (I...
2024
-
[20]
PyTorch Lightning, 2019
William Falcon and The PyTorch Lightning team. PyTorch Lightning, 2019. URLhttps://github.com/ Lightning-AI/lightning. Version 2.0.8
2019
-
[21]
Friedlander and Mark Schmidt
Michael P. Friedlander and Mark Schmidt. Hybrid deterministic-stochastic methods for data fitting. SIAM Journal on Scientific Computing, 34(3):A1380–A1405, 2012
2012
-
[22]
Compiling machine learning programs via high-level tracing
Roy Frostig, Matthew James Johnson, and Chris Leary. Compiling machine learning programs via high-level tracing. InProceedings of Machine Learning and Systems (MLSys), 2018
2018
-
[23]
Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, Johan Ferret, Peter Liu, Pouya Tafti, Abe Friesen, Michelle Casbon, Sabela Ramos, Ravin Kumar, Charline Le Lan...
2024 arXiv
-
[24]
Gemma: Open models based on Gemini research and technology
Google DeepMind Gemma Team. Gemma: Open models based on Gemini research and technology. arXiv preprint arXiv:2403.08295, 2024
2024 arXiv
-
[25]
OpenLLaMA: An open reproduction of LLaMA, May 2023
Xinyang Geng and Hao Liu. OpenLLaMA: An open reproduction of LLaMA, May 2023. URL https://github.com/openlm-research/open_llama
2023
-
[26]
Accurate, large minibatch SGD: Training ImageNet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. Accurate, large minibatch SGD: Training ImageNet in 1 hour. arXiv preprint arXiv:1706.02677, 2017
2017 arXiv
-
[27]
Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Raghavi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack...
2024 arXiv
-
[28]
Scaling laws for single-agent reinforcement learning.arXiv preprint arXiv:2301.13442, 2023
Jacob Hilton, Jie Tang, and John Schulman. Scaling laws for single-agent reinforcement learning.arXiv preprint arXiv:2301.13442, 2023
2023 arXiv
-
[29]
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry. Train longer, generalize better: closing the generalization gap in large batch training of neural networks. InAdvances in Neural Information Processing Systems (NeurIPS), 2017
2017
-
[30]
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas ...
-
[31]
GrowLength: Accelerating LLMs pretraining by progressively growing training length.arXiv preprint arXiv:2310.00576, 2023
Hongye Jin, Xiaotian Han, Jingfeng Yang, Zhimeng Jiang, Chia-Yuan Chang, and Xia Hu. GrowLength: Accelerating LLMs pretraining by progressively growing training length.arXiv preprint arXiv:2310.00576, 2023
2023 arXiv
-
[32]
AdaScale SGD: A user-friendly algorithm for distributed training
Tyler Johnson, Pulkit Agrawal, Haijie Gu, and Carlos Guestrin. AdaScale SGD: A user-friendly algorithm for distributed training. InProceedings of the International Conference on Machine Learning (ICML), 2020. 15
2020
-
[33]
Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models.arXiv preprint arXiv:2001.08361, 2020
2001 arXiv
-
[34]
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang. On large-batch training for deep learning: Generalization gap and sharp minima. InInternational Conference on Learning Representations (ICLR), 2017
2017
-
[35]
Better theory for SGD in the nonconvex world.Transactions on Machine Learning Research, 2023
Ahmed Khaled and Peter Richtárik. Better theory for SGD in the nonconvex world.Transactions on Machine Learning Research, 2023
2023
-
[36]
Kingma and Jimmy Lei Ba
Diederik P. Kingma and Jimmy Lei Ba. Adam: a method for stochastic optimization. InInternational Conference on Learning Representations (ICLR), 2015
2015
-
[37]
Reducing activation recomputation in large transformer models
Vijay Anand Korthikanti, Jared Casper, Sangkug Lym, Lawrence McAfee, Michael Andersch, Mohammad Shoeybi, and Bryan Catanzaro. Reducing activation recomputation in large transformer models. In Proceedings of Machine Learning and Systems (MLSys), 2023
2023
-
[38]
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. ImageNet classification with deep convolutional neural networks. InAdvances in Neural Information Processing Systems (NeurIPS), 2012
2012
-
[39]
Noise is not the main factor behind the gap between SGD and Adam on transformers, but sign descent might be
Frederik Kunstner, Jacques Chen, Jonathan Wilder Lavington, and Mark Schmidt. Noise is not the main factor behind the gap between SGD and Adam on transformers, but sign descent might be. In International Conference on Learning Representations (ICLR), 2023
2023
-
[40]
Heavy-tailed class imbalance and why Adam outperforms gradient descent on language models.arXiv preprint arXiv:2402.19449, 2024
Frederik Kunstner, Robin Yadav, Alan Milligan, Mark Schmidt, and Alberto Bietti. Heavy-tailed class imbalance and why Adam outperforms gradient descent on language models.arXiv preprint arXiv:2402.19449, 2024
2024 arXiv
-
[41]
Communication-efficient adaptive batch size strategies for distributed local gradient methods.arXiv preprint arXiv:2406.13936, 2024
Tim Tsz-Kit Lau, Weijian Li, Chenwei Xu, Han Liu, and Mladen Kolar. Communication-efficient adaptive batch size strategies for distributed local gradient methods.arXiv preprint arXiv:2406.13936, 2024
2024 arXiv
-
[42]
arXiv preprint arXiv:2402.11215, 2024
Tim Tsz-Kit Lau, Han Liu, and Mladen Kolar.AdAdaGrad: Adaptive batch size schemes for adaptive gradient methods. arXiv preprint arXiv:2402.11215, 2024
2024 arXiv
-
[43]
Orr, and Klaus Robert Müller
Yann LeCun, Leon Bottou, Genevieve B. Orr, and Klaus Robert Müller. Efficient BackProp. In Genevieve B. Orr and Klaus-Robert Müller, editors,Neural Networks: Tricks of the Trade, pages 9–50. Springer Berlin Heidelberg, 2002
2002
-
[44]
GShard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. GShard: Scaling giant models with conditional computation and automatic sharding. InInternational Conference on Learning Representations (ICLR), 2021
2021
-
[45]
The stability-efficiency dilemma: Investigating sequence length warmup for training GPT models.Advances in Neural Information Processing Systems (NeurIPS), 2022
Conglong Li, Minjia Zhang, and Yuxiong He. The stability-efficiency dilemma: Investigating sequence length warmup for training GPT models.Advances in Neural Information Processing Systems (NeurIPS), 2022
2022
-
[46]
PyTorch distributed: Experiences on accelerating data parallel training
Shen Li, Yanli Zhao, Rohan Varma, Omkar Salpekar, Pieter Noordhuis, Teng Li, Adam Paszke, Jeff Smith, Brian Vaughan, Pritam Damania, and Soumith Chintala. PyTorch distributed: Experiences on accelerating data parallel training. InProceedings of the VLDB Endowment, 2020
2020
-
[47]
TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training
Wanchao Liang, Tianyu Liu, Less Wright, Will Constable, Andrew Gu, Chien-Chin Huang, Iris Zhang, Wei Feng, Howard Huang, Junjie Wang, Sanket Purandare, Gokul Nadathur, and Stratos Idreos. TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training. InInt...
2025
-
[48]
LitGPT, 2023
Lightning AI. LitGPT, 2023. URLhttps://github.com/Lightning-AI/litgpt. 16
2023
-
[49]
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. RoBERTa: A robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692, 2019
1907 arXiv
-
[50]
The Llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024
Llama Team, AI @ Meta. The Llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024
2024 arXiv
-
[51]
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. InInternational Conference on Learning Representations (ICLR), 2019
2019
-
[52]
An empirical model of large-batch training
Sam McCandlish, Jared Kaplan, Dario Amodei, and OpenAI Dota Team. An empirical model of large-batch training. arXiv preprint arXiv:1812.06162, 2018
2018 arXiv
-
[53]
Efficient large-scale language model training on GPU clusters using Megatron-LM
Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley, Mostofa Patwary, Vi- jay Anand Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, Amar Phanishayee, and Matei Zaharia. Efficient large-scale language model training on GPU cl...
2021
-
[54]
AdaBatchGrad: Combining adaptive batch size and adaptive step size.arXiv preprint arXiv:2402.05264, 2024
Petr Ostroukhov, Aigerim Zhumabayeva, Chulu Xiang, Alexander Gasnikov, Martin Takáč, and Dmitry Kamzolov. AdaBatchGrad: Combining adaptive batch size and adaptive step size.arXiv preprint arXiv:2402.05264, 2024
2024 arXiv
-
[55]
Toward understanding why Adam converges faster than SGD for transformers
Yan Pan and Yuanzhi Li. Toward understanding why Adam converges faster than SGD for transformers. arXiv preprint arXiv:2306.00204, 2023
2023 arXiv
-
[56]
Nemotron-4 15B technical report.arXiv preprint arXiv:2402.16819, 2024
Jupinder Parmar, Shrimai Prabhumoye, Joseph Jennings, Mostofa Patwary, Sandeep Subramanian, Dan Su, Chen Zhu, Deepak Narayanan, Aastha Jhunjhunwala, Ayush Dattagupta, Vibhu Jawa, Jiwei Liu, Ameya Mahabaleshwarkar, Osvald Nitski, Annika Brundyn, James Maki, Miguel Martinez, Jia...
2024 arXiv
-
[57]
PyTorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fan...
2019
-
[58]
Large scale language modeling: Converging on 40GB of text in four hours
Raul Puri, Robert Kirby, Nikolai Yakovenko, and Bryan Catanzaro. Large scale language modeling: Converging on 40GB of text in four hours. InProceedings of the International Symposium on Computer Architecture and High Performance Computing (SBAC-PAD), 2018
2018
-
[59]
SimiGrad: Fine-grained adaptive batching for large scale training using gradient similarity measurement
Heyang Qin, Samyam Rajbhandari, Olatunji Ruwase, Feng Yan, Lei Yang, and Yuxiong He. SimiGrad: Fine-grained adaptive batching for large scale training using gradient similarity measurement. InAdvances in Neural Information Processing Systems (NeurIPS), 2021
2021
-
[60]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67, 2020
2020
-
[61]
ZeRO: Memory optimizations toward training trillion parameter models
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. ZeRO: Memory optimizations toward training trillion parameter models. InSC20: International Conference for High Performance Computing, Networking, Storage and Analysis, pages 1–16. IEEE, 2020
2020
-
[62]
DeepSpeed: System optimizations enable training deep learning models with over 100 billion parameters
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. DeepSpeed: System optimizations enable training deep learning models with over 100 billion parameters. InProceedings of the ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2020. 17
2020
-
[63]
ZeRO-Offload: Democratizing billion-scale model training
Jie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, and Yuxiong He. ZeRO-Offload: Democratizing billion-scale model training. InUSENIX Annual Technical Conference (USENIX ATC), 2021
2021
-
[64]
On the different regimes of stochastic gradient descent.Proceedings of the National Academy of Sciences, 121(9):e2316301121, 2024
Antonio Sclocchi and Matthieu Wyart. On the different regimes of stochastic gradient descent.Proceedings of the National Academy of Sciences, 121(9):e2316301121, 2024
2024
-
[65]
Shallue, Jaehoon Lee, Joseph Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E
Christopher J. Shallue, Jaehoon Lee, Joseph Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E. Dahl. Measuring the effects of data parallelism on neural network training.Journal of Machine Learning Research, 20(112):1–49, 2019
2019
-
[66]
Megatron-LM: Training multi-billion parameter language models using model parallelism.arXiv preprint arXiv:1909.08053, 2019
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. Megatron-LM: Training multi-billion parameter language models using model parallelism.arXiv preprint arXiv:1909.08053, 2019
1909 arXiv
-
[67]
Smith and Quoc V
Samuel L. Smith and Quoc V. Le. A Bayesian perspective on generalization and stochastic gradient descent. InInternational Conference on Learning Representations (ICLR), 2018
2018
-
[68]
Smith, Pieter-Jan Kindermans, and Quoc V
Samuel L. Smith, Pieter-Jan Kindermans, and Quoc V. Le. Don’t decay the learning rate, increase the batch size. InInternational Conference on Learning Representations (ICLR), 2018
2018
-
[69]
Using DeepSpeed and Megatron to train Megatron-Turing NLG 530B, a large-scale generative language model.arXiv preprint arXiv:2201.11990, 2022
Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, Elton Zhang, Rewon Child, Reza Yazdani Aminabadi, Julie Bernauer, Xia Song, Mohammad Shoeybi, Yuxiong He, Michae...
2022 arXiv
-
[70]
Unraveling the mystery of scaling laws: Part I.arXiv preprint arXiv:2403.06563, 2024
Hui Su, Zhi Tian, Xiaoyu Shen, and Xunliang Cai. Unraveling the mystery of scaling laws: Part I.arXiv preprint arXiv:2403.06563, 2024
2024 arXiv
-
[71]
Team OLMo, Pete Walsh, Luca Soldaini, Dirk Groeneveld, Kyle Lo, Shane Arora, Akshita Bhagia, Yuling Gu, Shengyi Huang, Matt Jordan, Nathan Lambert, Dustin Schwenk, Oyvind Tafjord, Taira Anderson, David Atkinson, Faeze Brahman, Christopher Clark, Pradeep Dasigi, Nouha Dziri, Mi...
2025 arXiv
-
[72]
Introducing DBRX: A new state-of-the-art open LLM
The Mosaic Research Team. Introducing DBRX: A new state-of-the-art open LLM. https://www. databricks.com/blog/introducing-dbrx-new-state-art-open-llm , 2024
2024
-
[73]
Llama 2: Open foundation and fine-tuned chat models.arXiv preprint arXiv:2307.09288, 2023
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, W...
2023 arXiv
-
[74]
Gomez, Łukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. InAdvances in Neural Information Processing Systems (NeurIPS), 2017. 18
2017
-
[75]
Meta Lingua: A minimal PyTorch LLM training library, 2024
Mathurin Videau, Badr Youbi Idrissi, Daniel Haziza, Luca Wehrstedt, Jade Copet, Olivier Teytaud, and David Lopez-Paz. Meta Lingua: A minimal PyTorch LLM training library, 2024. URLhttps: //github.com/facebookresearch/lingua
2024
-
[76]
Closing the gap between the upper bound and lower bound of Adam’s iteration complexity
Bohan Wang, Jingwen Fu, Huishuai Zhang, Nanning Zheng, and Wei Chen. Closing the gap between the upper bound and lower bound of Adam’s iteration complexity. InAdvances in Neural Information Processing Systems (NeurIPS), 2023
2023
-
[77]
MicroLlama-300M
Ken Wang. MicroLlama-300M. https://github.com/keeeeenw/MicroLlama, 2024
2024
-
[78]
Liu, Lechao Xiao, Katie Everett, Alex Alemi, Ben Adlam, John D
Mitchell Wortsman, Peter J. Liu, Lechao Xiao, Katie Everett, Alex Alemi, Ben Adlam, John D. Co-Reyes, Izzeddin Gur, Abhishek Kumar, Roman Novak, Jeffrey Pennington, Jascha Sohl-dickstein, Kelvin Xu, Jaehoon Lee, Justin Gilmer, and Simon Kornblith. Small-scale proxies for large...
2024
-
[79]
Baichuan 2: Open large-scale language models.arXiv preprint arXiv:2309.10305, 2023
Aiyuan Yang, Bin Xiao, Bingning Wang, Borong Zhang, Ce Bian, Chao Yin, Chenxu Lv, Da Pan, Dian Wang, Dong Yan, Fan Yang, Fei Deng, Feng Wang, Feng Liu, Guangwei Ai, Guosheng Dong, Haizhou Zhao, Hang Xu, Haoze Sun, Hongda Zhang, Hui Liu, Jiaming Ji, Jian Xie, JunTao Dai, Kun Fa...
2023 arXiv
-
[80]
Qwen2 technical report
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jianxin Yang, Jin Xu, Jingren Zhou, Jinze...
2024 arXiv
-
[81]
Large batch optimization for deep learning: Training BERT in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh. Large batch optimization for deep learning: Training BERT in 76 minutes. InInternational Conference on Learning Representations (IC...
2020
-
[82]
Susskind
Shuangfei Zhai, Tatiana Likhomanenko, Etai Littwin, Dan Busbridge, Jason Ramapuram, Yizhe Zhang, Jiatao Gu, and Joshua M. Susskind. Stabilizing transformer training by preventing attention entropy collapse. In Proceedings of the International Conference on Machine Learning (IC...
2023
-
[83]
How does critical batch size scale in pre-training? InInternational Conference on Learning Representations (ICLR), 2025
Hanlin Zhang, Depen Morwani, Nikhil Vyas, Jingfeng Wu, Difan Zou, Udaya Ghai, Dean Foster, and Sham Kakade. How does critical batch size scale in pre-training? InInternational Conference on Learning Representations (ICLR), 2025
2025
-
[84]
TinyLlama: An open-source small language model
Peiyuan Zhang, Guangtao Zeng, Tianduo Wang, and Wei Lu. TinyLlama: An open-source small language model. arXiv preprint arXiv:2401.02385, 2024
2024 arXiv
-
[85]
OPT: Open pre-trained transformer language models.arXiv preprint arXiv:2205.01068, 2022
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. O...
2022 arXiv
-
[86]
Why transformers need Adam: A Hessian perspective.arXiv preprint arXiv:2402.16788, 2024
Yushun Zhang, Congliang Chen, Tian Ding, Ziniu Li, Ruoyu Sun, and Zhi-Quan Luo. Why transformers need Adam: A Hessian perspective.arXiv preprint arXiv:2402.16788, 2024
2024 arXiv
-
[87]
Picotron: Distributed training framework for education and research experimentation, 2025
Haojun Zhao and Ferdinand Mom. Picotron: Distributed training framework for education and research experimentation, 2025. URL https://github.com/huggingface/picotron. 19
2025
-
[88]
first-order term
Yanli Zhao, Andrew Gu, Rohan Varma, Liang Luo, Chien-Chin Huang, Min Xu, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, Alban Desmaison, Can Balioglu, Pritam Damania, Bernard Nguyen, Geeta Chauhan, Yuchen Hao, Ajit Mathews, and Shen Li. PyTorch FSDP: Experiences on sc...
2023
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.