Pith. sign in

REVIEW 4 major objections 4 minor 88 references

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism

T0 review · 4 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This paper claims that variance-based norm tests, extended to data and model parallel training, let language model pretraining grow batch sizes on demand, matching small-batch quality at large-batch speed, with Adam still provably…

desk verdict Real distributed variance estimator, but the empirical claims don't survive the paper's own tables on the larger models. read the letter →

arxiv 2412.21124 v2 pith:25X6ZO74 submitted 2024-12-30 cs.LG math.OCstat.ML

classification cs.LGmath.OCstat.ML MSC 68T0790C2690C1568W10
keywords adaptivebatchsizenormtestlanguagemodelpretrainingdistributeddataparallelismAdamconvergencegeneralizationgapfullyshardedparallel
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper's central claim is that the long-standing trade-off in language model training—large batches give throughput but worse generalization, small batches generalize better but train slowly—can be resolved by an adaptive batch size schedule that grows the batch only when the current gradient estimate is provably too noisy. The authors extend the norm test of adaptive sampling to distributed settings that combine data parallelism with model parallelism, implement it in a fully sharded data parallel framework, and report that it beats constant batch sizes and hand-designed warmup schedules on decoder-only models up to three billion parameters. On the theory side, they prove that Adam, the standard optimizer for pretraining, converges under these schedules for smooth nonconvex objectives if a coordinate-wise variance norm test is satisfied. If correct, this gives practitioners a way to adapt batch sizes to training dynamics rather than following fixed recipes.

What carries the argument

The load-bearing object is the norm test statistic: the ratio of an estimate of the per-sample gradient variance to the squared norm of the current batch gradient, compared against a threshold $\eta^2$. In the distributed setting this statistic is computed without per-sample gradients by measuring how much the per-worker minibatch gradients differ from the global batch gradient, a version the paper calls DDP-Norm (and FSDP-Norm under model parallelism). The proof rests on a coordinate-wise version of the exact-variance norm test, which the authors show implies the coordinate-wise expected strong growth condition that Adam's iteration-complexity analysis requires.

What would settle it

Train one small model twice with identical hyperparameters, once using the approximate worker-level variance statistic to choose batches and once using exact per-sample gradient variance, and compare the batch-size trajectories and validation losses; if the approximate schedule consistently selects notably different batch sizes or yields worse validation loss at matched sample counts, the claim that the practical test supports the convergence guarantee is refuted.

Watch

Extended reading notes

Core claim

In the paper's own terms, the discovery is that the norm test—the rule that increases the next batch size to $\lceil \| \operatorname{Var}_{i \in B_k}(\nabla \ell_i(w_k))\|_1 / (\eta^2 \|\nabla \mathcal{L}_{B_k}(w_k)\|^2) \rceil$ when the current batch gradient is too noisy—can be made practical for distributed, model-parallel training and can carry Adam to convergence. The theoretical result, Theorem 1, proves a convergence bound of order $\sqrt{K}$ up to logarithmic factors on the cumulative expected full-gradient norm under the coordinate-wise exact-variance norm test, which implies the coordinate-wise expected strong growth condition. The empirical result is that, at equal numbers of training samples or steps, the adaptive schedules match or improve validation loss relative to constant large batches and to heuristic stagewise warmups, while using fewer steps when the batch can grow. The authors state this as a general-purpose schedule applicable beyond language models, with a particular demonstration on up-to-3-billion-parameter Llama 2 family models.

Load-bearing premise

The proof assumes the exact coordinate-wise variance norm test is satisfied at every iteration, but the deployed algorithm checks a cheaper approximate statistic built from differences between workers' minibatch gradients and never verifies that the condition actually holds for the next batch.

Editorial extensions

If this is right

  • Practitioners can start training with a small batch and let the method increase it only when gradient noise is small, obtaining small-batch validation quality at large-batch throughput.
  • The same schedule is compatible with data and model parallelism, so memory-constrained multi-GPU setups can pretrain billion-parameter models without hand-tuning a batch warmup recipe.
  • Adam's convergence under the schedule is guaranteed for smooth nonconvex objectives when the coordinate-wise norm test holds, removing the need for a global variance assumption.
  • With comparable wall-clock time, the adaptive schedules produce validation losses closer to those of the best small constant batch while using far fewer gradient steps.
  • The tuning effort for pretraining shifts to the single threshold $\eta$, which controls how aggressively the batch grows.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The approximation that replaces per-sample gradient variance with between-worker minibatch variance has only $J$ samples (the number of workers); when $J$ is small the batch-size schedule will inherit high variance, so smoothing the statistic over a few iterations is a natural extension the paper does not test.
  • A falsifiable check of the theory-practice link is to run a small model with exact per-sample gradient variance as the test and compare the resulting batch-size trajectory with the approximate worker-level test; close trajectories would certify the approximation that the proof needs.
  • The same norm-test machinery could be coupled with sequence-length warmup or learning-rate schedules, since both affect the gradient-noise estimate and the descent direction; the paper mentions sequence length warmup but does not combine them.
  • The claimed scaling-law connection between $\eta$ and the critical batch size is left open; a direct experiment varying $\eta$ while measuring the final batch size plateau would test whether the plateau tracks the critical batch size.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper proposes adaptive batch size schedules for distributed training of language models, based on a norm test that grows the global batch when the approximate gradient variance is too large relative to the batch gradient norm. The authors present DDP-Norm and FSDP-Norm implementations and a convergence theorem for Adam under a coordinate-wise exact variance norm test. Experiments are reported on MicroLlama 300M, TinyLlama 1.1B, and OpenLlama 3B pretrained on C4, with the claim that the adaptive schedules outperform constant batch sizes and heuristic batch-size warmup schedules for Llama-family models.

Significance. If the empirical and theoretical claims were correct, this would be a useful contribution: a theoretically motivated, practically implemented method for adjusting batch sizes during large-scale distributed LLM pretraining, with an open-source FSDP implementation and reproducible experimental configuration. The paper also gives credit for shipping code and for a seriousness of purpose in attempting to connect adaptive sampling theory with distributed training practice. However, the central empirical claim is contradicted by the paper's own tables on the two larger testbeds, and the convergence theorem does not apply to the algorithm actually run. The strengths are real but are outweighed by the failure of the two headline claims.

major comments (4)
  1. [§5.2 Table 2; §5.3 Table 3; Abstract] The abstract's claim that the proposed approaches "outperform constant batch sizes" in the pretraining of Llama-family models is directly contradicted by the paper's own results. On TinyLlama 1.1B (Table 2), the best adaptive run (eta=0.085) reaches validation loss 4.256, while the constant batch size 4096 baseline reaches 3.817. On OpenLlama 3B (Table 3), the best adaptive run (eta=0.15) reaches 4.554, while constant 4096 reaches 3.956. On both larger models, every adaptive schedule is worse than the constant 4096 baseline, which is an internal inconsistency with the stated contribution.
  2. [§5.3] The wall-clock argument does not rescue the empirical claim. The authors concede in §5.3 that "using a constant batch size 4096 achieves an even lower validation loss," but the reported time savings are small: about 5% on TinyLlama (34.48h vs 32.83h) and about 6% on OpenLlama (20.75h vs 19.59h), while the validation-loss differences are large (roughly 0.44 and 0.60). Because all runs consume the same 2,000,000 training sequences, the constant batch size 4096 baseline dominates on a per-token basis, so the efficiency framing does not compensate for the worse final validation loss.
  3. [§4, Theorem 1; Algorithm 1; Appendix B, Remark B.1] There is a load-bearing mismatch between the convergence theorem and the implemented algorithm. Theorem 1 assumes that the coordinate-wise exact variance norm test, Ek[(∂iℒBk(wk)−∂iℒ(wk))2] ≤ η2(∂iℒ(wk))2, is satisfied at every iteration. Algorithm 1, however, implements the aggregate, non-coordinate-wise approximate test in Eq. (5) (DDP-Norm/FSDP-Norm) and does not verify the exact condition for the next batch. Appendix B Remark B.1 explicitly concedes that "the exact variance test is not implemented in practice but its approximate version instead." Consequently, the proof does not establish convergence for the algorithm whose empirical behavior is reported.
  4. [§4, Theorem 1; Appendix B, Theorem B.1 and definition of c2] The stated convergence guarantee is vacuous as written. Theorem 1 bounds ∑k=1K E[‖∇ℒ(wk)‖] by ~O(K). A bound of O(K) on a sum of K terms holds trivially for any algorithm with bounded per-iteration gradient norms and does not imply convergence to a stationary point. The explicit bound in Theorem B.1 is also at least linear in K: the constant c2 contains the term 2c1∑i (log(1/√β2 v0,i) − K log β2), which grows linearly in K since log β2 < 0. Thus the average gradient norm need not decay, so the theorem does not deliver the advertised convergence guarantee.
minor comments (4)
  1. [§3.1, Eq. (2)] The notation Vari∈B(∇ℓi(w)) is used as a vector while the displayed expression mixes norms and scalar quantities; please define the per-coordinate variance vector explicitly and distinguish it from its L1 norm.
  2. [§5.4] The text contains the typo "prupose" in the paragraph on the effect of η; it should read "purpose."
  3. [Figures 2 and 3] The batch-size panels label the horizontal axis as "sample ×105" or "sample ×106"; the unit should be "samples" and the exponent formatting made consistent across panels.
  4. [§4] The experiments use AdamW with decoupled weight decay, whereas Theorem 1 concerns Adam without weight decay; the paper acknowledges this in the text, but the main-text discussion would be clearer if the limitation were stated immediately after the theorem rather than in the later discussion.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the convergence proof is conditional on an exact norm-test assumption and does not reduce to the implemented schedule or to a fitted value.

full rationale

The paper's derivation chain is not circular. Theorem 1 (formalized as Theorem B.1) assumes the coordinate-wise exact variance norm test, Eq. (4), holds at every iteration and derives an O~(√K) bound on the sum of expected gradient norms. This assumption is a strong-growth-type noise bound; the conclusion is a different statement, so no equation reduces to an input by construction. The practical DDP-Norm/FSDP-Norm statistic, Eq. (5), is explicitly acknowledged in Remark B.1 to be an approximate version of the exact test ('the exact variance test is not implemented in practice but its approximate version instead'), so the theory-implementation gap is a soundness problem, not circularity. Self-citations to [42] for the E-SG nomenclature and [41] for related local-gradient extensions are not load-bearing: the norm test originates from Byrd et al. [12] and the proof technique is attributed to the external work [76]. Finally, the empirical claim in the abstract is weakened by the paper's own Tables 2 and 3, where constant batch size 4096 reaches lower validation loss than every adaptive schedule, and Section 5.3 concedes 'While using a constant batch size 4096 achieves an even lower validation loss'; this is an evidence/correctness concern, not a circularity. No fitted parameter is fed back into the theorem, and no uniqueness claim is imported from the authors' prior work.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The adaptive schedule is driven by a manually chosen threshold eta and by initial and maximum batch size caps, all selected per experiment. The convergence proof relies on an exact test that the implementation does not run, and it inherits unproved lemmas from Wang et al. [76]. No new physical or algorithmic entities beyond the DDP-Norm and FSDP-Norm variants are introduced.

free parameters (4)
  • norm test threshold eta = 0.15/0.2/0.25/0.275 (MicroLlama); 0.05/0.075/0.08/0.085 (TinyLlama); 0.05/0.1/0.15 (OpenLlama)
    The threshold controls batch-size growth and is manually selected per experiment. Section 5.4 says choosing the right eta is vital, and the reported results are sensitive to it.
  • initial global batch size = 256 (MicroLlama); 128 (TinyLlama and OpenLlama)
    The starting point of the adaptive schedule is chosen by hand rather than derived; it affects early training dynamics and the number of steps.
  • maximum global batch size cap = 8192
    The schedule stops growing at this manually chosen cap. In several runs the batch size quickly reaches the cap, so part of the adaptivity is simply an early ramp to a fixed large value.
  • micro-batch configuration = base micro batch 4, maximum micro batch 8, gradient accumulation 16
    These memory-driven choices define how the global batch size is mapped to workers and gradient accumulation steps; they are chosen by hardware constraints rather than by the theory.
assumptions (5)
  • domain assumption The loss is L-Lipschitz smooth (Assumption 1).
    Standard nonconvex smoothness assumption used in the convergence theorem; it may not hold in practice for transformer losses with attention and normalization.
  • ad hoc to paper The exact coordinate-wise variance norm test holds at every iteration.
    Theorem 1 and Proposition 1 require this exact condition. The implemented DDP-Norm and FSDP-Norm use the approximate estimator in Eq. (5) and do not check the condition on the next batch.
  • ad hoc to paper The approximate variance estimator in Eq. (5) is a valid proxy for the per-sample gradient variance.
    The implementation uses between-worker minibatch gradient differences scaled by 1/bk, but the paper gives no bound relating this estimator to the exact variance norm test used in the theory.
  • standard math The technical lemmas of Wang et al. [76] are correct and applicable.
    Appendix B states these lemmas without proof and builds the Adam convergence argument directly on them.
  • domain assumption The theoretical Adam update omits bias correction, weight decay, learning rate schedules, and gradient clipping.
    Section 4 explicitly excludes these components, while the experiments use AdamW with decoupled weight decay, linear warmup, cosine decay, and gradient clipping. The theorem therefore does not cover the experimental setting.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism." pith.science (2026). https://pith.science/paper/25X6ZO74

@misc{pith2026241221124,
  author       = {Pith},
  title        = {Pith review of: Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/25X6ZO74}},
  note         = {Machine review of arXiv:2412.21124}
}
read the original abstract

An appropriate choice of batch sizes in large-scale model training is crucial, yet it involves an intrinsic yet inevitable dilemma: large-batch training improves training efficiency in terms of memory utilization, while generalization performance often deteriorates due to small amounts of gradient noise. Despite this dilemma, the common practice of choosing batch sizes in language model training often prioritizes training efficiency -- employing either constant large sizes with data parallelism or implementing batch size warmup schedules. However, such batch size schedule designs remain heuristic and often fail to adapt to training dynamics, presenting the challenge of designing adaptive batch size schedules. Given the abundance of available datasets and the data-hungry nature of language models, data parallelism has become an indispensable distributed training paradigm, enabling the use of larger batch sizes for gradient computation. However, vanilla data parallelism requires replicas of model parameters, gradients, and optimizer states at each worker, which prohibits training larger models with billions of parameters. To optimize memory usage, more advanced parallelism strategies must be employed. In this work, we propose general-purpose and theoretically principled adaptive batch size schedules compatible with data parallelism and model parallelism. We develop a practical implementation with PyTorch Fully Sharded Data Parallel, facilitating the pretraining of language models of different sizes. We empirically demonstrate that our proposed approaches outperform constant batch sizes and heuristic batch size warmup schedules in the pretraining of models in the Llama 2 family, with particular focus on smaller models with up to 3 billion parameters. We also establish theoretical convergence guarantees for such adaptive batch size schedules with Adam for general smooth nonconvex objectives.

Figures

Figures reproduced from arXiv: 2412.21124 by the authors.

Figure 1
Figure 1. Generalization gap in transformer pretraining. Various curves represent distinct batch sizes. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Training loss, validation loss and batch size schedule for MicroLlama 300M [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Training loss, validation loss and batch size schedule for TinyLlama 1.1B [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Training loss, validation loss and batch size schedule for OpenLlama 3B [PITH_FULL_IMAGE:figures/full_fig_p010_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

88 extracted references · 42 canonical work pages

  1. [1]

    Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dandelion Mané, Rajat Monga, Sherry Moore, Derek Murra...

  2. [2]

    Extremely large minibatch SGD: Training ResNet-50 on ImageNet in 15 minutes.arXiv preprint arXiv:1711.04325, 2017

    Takuya Akiba, Shuji Suzuki, and Keisuke Fukuda. Extremely large minibatch SGD: Training ResNet-50 on ImageNet in 15 minutes.arXiv preprint arXiv:1711.04325, 2017

  3. [3]

    Qwen technical report.arXiv preprint arXiv:2309.16609, 2023

    Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Benfeng X...

  4. [4]

    Coupling adaptive batch sizes with learning rates

    Lukas Balles, Javier Romero, and Philipp Hennig. Coupling adaptive batch sizes with learning rates. In Proceedings of the Conference on Uncertainty in Artificial Intelligence (UAI), 2017

  5. [5]

    Stable LM 2 1.6B technical report.arXiv preprint arXiv:2402.17834, 2024

    Marco Bellagente, Jonathan Tow, Dakota Mahan, Duy Phung, Maksym Zhuravinskyi, Reshinth Adithyan, James Baicoianu, Ben Brooks, Nathan Cooper, Ashish Datta, Meng Lee, Emad Mostaque, Michael Pieler, Nikhil Pinnaparju, Paulo Rocha, Harry Saini, Hannah Teufel, Niccolo Zanichelli, and Carlos Riquelme. Stable LM 2 1.6B technical report.arXiv preprint arXiv:2402....

  6. [6]

    BigScience Workshop, Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Benoît Sagot, Niklas Muennighoff, Albert Villanova del Moral, Olatu...

  7. [7]

    Raghu Bollapragada and Stefan M. Wild. Adaptive sampling quasi-Newton methods for zeroth-order stochastic optimization. Mathematical Programming Computation, 15(2):327–364, 2023

  8. [8]

    Adaptive sampling strategies for stochastic optimization

    Raghu Bollapragada, Richard Byrd, and Jorge Nocedal. Adaptive sampling strategies for stochastic optimization. SIAM Journal on Optimization, 28(4):3312–3343, 2018

Show all 88 references
  1. [9]

    Curtis, and Jorge Nocedal

    Léon Bottou, Frank E. Curtis, and Jorge Nocedal. Optimization methods for large-scale machine learning. SIAM Review, 60(2):223–311, 2018

  2. [10]

    JAX: composable transformations of Python+NumPy programs, 2018

    James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang. JAX: composable transformations of Python+NumPy programs, 2018. URLhttp://github.com/google/jax

  3. [11]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffr...

  4. [12]

    Byrd, Gillian M

    Richard H. Byrd, Gillian M. Chin, Jorge Nocedal, and Yuchen Wu. Sample size selection in optimization methods for machine learning.Mathematical Programming, 134(1):127–155, 2012

  5. [13]

    Big batch SGD: Automated inference using adaptive batch sizes.arXiv preprint arXiv:1610.05792, 2016

    Soham De, Abhay Yadav, David Jacobs, and Tom Goldstein. Big batch SGD: Automated inference using adaptive batch sizes.arXiv preprint arXiv:1610.05792, 2016

  6. [14]

    Automated inference with adaptive batches

    Soham De, Abhay Yadav, David Jacobs, and Tom Goldstein. Automated inference with adaptive batches. In Proceedings of the International Conference on Artificial Intelligence and Statistics (AISTATS), 2017

  7. [15]

    Zhang, Hanwei Xu, Hao Yang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J

    DeepSeek-AI, Aixin Liu, Bei Feng, Bin Wang, Bingxuan Wang, Bo Liu, Chenggang Zhao, Chengqi Dengr, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei Li, H. Zhang, Hanwei Xu, Hao Yang, Haowei...

  8. [16]

    Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J

    DeepSeek-AI, Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei L...

  9. [17]

    Adabatch: Adaptive batch sizes for training deep neural networks.arXiv preprint arXiv:1712.02029, 2017

    Aditya Devarakonda, Maxim Naumov, and Michael Garland. Adabatch: Adaptive batch sizes for training deep neural networks.arXiv preprint arXiv:1712.02029, 2017

  10. [18]

    An image is worth 16x16 words: Transformers for image recognition at scale

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at...

  11. [19]

    Susskind, and Armand Joulin

    Alaaeldin El-Nouby, Michal Klein, Shuangfei Zhai, Miguel Angel Bautista, Alexander Toshev, Vaishaal Shankar, Joshua M. Susskind, and Armand Joulin. Scalable pre-training of large autoregressive image models. In Proceedings of the International Conference on Machine Learning (I...

  12. [20]

    PyTorch Lightning, 2019

    William Falcon and The PyTorch Lightning team. PyTorch Lightning, 2019. URLhttps://github.com/ Lightning-AI/lightning. Version 2.0.8

  13. [21]

    Friedlander and Mark Schmidt

    Michael P. Friedlander and Mark Schmidt. Hybrid deterministic-stochastic methods for data fitting. SIAM Journal on Scientific Computing, 34(3):A1380–A1405, 2012

  14. [22]

    Compiling machine learning programs via high-level tracing

    Roy Frostig, Matthew James Johnson, and Chris Leary. Compiling machine learning programs via high-level tracing. InProceedings of Machine Learning and Systems (MLSys), 2018

  15. [23]

    Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, Johan Ferret, Peter Liu, Pouya Tafti, Abe Friesen, Michelle Casbon, Sabela Ramos, Ravin Kumar, Charline Le Lan...

  16. [24]

    Gemma: Open models based on Gemini research and technology

    Google DeepMind Gemma Team. Gemma: Open models based on Gemini research and technology. arXiv preprint arXiv:2403.08295, 2024

  17. [25]

    OpenLLaMA: An open reproduction of LLaMA, May 2023

    Xinyang Geng and Hao Liu. OpenLLaMA: An open reproduction of LLaMA, May 2023. URL https://github.com/openlm-research/open_llama

  18. [26]

    Accurate, large minibatch SGD: Training ImageNet in 1 hour

    Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. Accurate, large minibatch SGD: Training ImageNet in 1 hour. arXiv preprint arXiv:1706.02677, 2017

  19. [27]

    Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Raghavi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack...

  20. [28]

    Scaling laws for single-agent reinforcement learning.arXiv preprint arXiv:2301.13442, 2023

    Jacob Hilton, Jie Tang, and John Schulman. Scaling laws for single-agent reinforcement learning.arXiv preprint arXiv:2301.13442, 2023

  21. [29]

    Train longer, generalize better: closing the generalization gap in large batch training of neural networks

    Elad Hoffer, Itay Hubara, and Daniel Soudry. Train longer, generalize better: closing the generalization gap in large batch training of neural networks. InAdvances in Neural Information Processing Systems (NeurIPS), 2017

  22. [30]

    Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas ...

  23. [31]

    GrowLength: Accelerating LLMs pretraining by progressively growing training length.arXiv preprint arXiv:2310.00576, 2023

    Hongye Jin, Xiaotian Han, Jingfeng Yang, Zhimeng Jiang, Chia-Yuan Chang, and Xia Hu. GrowLength: Accelerating LLMs pretraining by progressively growing training length.arXiv preprint arXiv:2310.00576, 2023

  24. [32]

    AdaScale SGD: A user-friendly algorithm for distributed training

    Tyler Johnson, Pulkit Agrawal, Haijie Gu, and Carlos Guestrin. AdaScale SGD: A user-friendly algorithm for distributed training. InProceedings of the International Conference on Machine Learning (ICML), 2020. 15

  25. [33]

    Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei

    Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models.arXiv preprint arXiv:2001.08361, 2020

  26. [34]

    On large-batch training for deep learning: Generalization gap and sharp minima

    Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang. On large-batch training for deep learning: Generalization gap and sharp minima. InInternational Conference on Learning Representations (ICLR), 2017

  27. [35]

    Better theory for SGD in the nonconvex world.Transactions on Machine Learning Research, 2023

    Ahmed Khaled and Peter Richtárik. Better theory for SGD in the nonconvex world.Transactions on Machine Learning Research, 2023

  28. [36]

    Kingma and Jimmy Lei Ba

    Diederik P. Kingma and Jimmy Lei Ba. Adam: a method for stochastic optimization. InInternational Conference on Learning Representations (ICLR), 2015

  29. [37]

    Reducing activation recomputation in large transformer models

    Vijay Anand Korthikanti, Jared Casper, Sangkug Lym, Lawrence McAfee, Michael Andersch, Mohammad Shoeybi, and Bryan Catanzaro. Reducing activation recomputation in large transformer models. In Proceedings of Machine Learning and Systems (MLSys), 2023

  30. [38]

    Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. ImageNet classification with deep convolutional neural networks. InAdvances in Neural Information Processing Systems (NeurIPS), 2012

  31. [39]

    Noise is not the main factor behind the gap between SGD and Adam on transformers, but sign descent might be

    Frederik Kunstner, Jacques Chen, Jonathan Wilder Lavington, and Mark Schmidt. Noise is not the main factor behind the gap between SGD and Adam on transformers, but sign descent might be. In International Conference on Learning Representations (ICLR), 2023

  32. [40]

    Heavy-tailed class imbalance and why Adam outperforms gradient descent on language models.arXiv preprint arXiv:2402.19449, 2024

    Frederik Kunstner, Robin Yadav, Alan Milligan, Mark Schmidt, and Alberto Bietti. Heavy-tailed class imbalance and why Adam outperforms gradient descent on language models.arXiv preprint arXiv:2402.19449, 2024

  33. [41]

    Communication-efficient adaptive batch size strategies for distributed local gradient methods.arXiv preprint arXiv:2406.13936, 2024

    Tim Tsz-Kit Lau, Weijian Li, Chenwei Xu, Han Liu, and Mladen Kolar. Communication-efficient adaptive batch size strategies for distributed local gradient methods.arXiv preprint arXiv:2406.13936, 2024

  34. [42]

    arXiv preprint arXiv:2402.11215, 2024

    Tim Tsz-Kit Lau, Han Liu, and Mladen Kolar.AdAdaGrad: Adaptive batch size schemes for adaptive gradient methods. arXiv preprint arXiv:2402.11215, 2024

  35. [43]

    Orr, and Klaus Robert Müller

    Yann LeCun, Leon Bottou, Genevieve B. Orr, and Klaus Robert Müller. Efficient BackProp. In Genevieve B. Orr and Klaus-Robert Müller, editors,Neural Networks: Tricks of the Trade, pages 9–50. Springer Berlin Heidelberg, 2002

  36. [44]

    GShard: Scaling giant models with conditional computation and automatic sharding

    Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. GShard: Scaling giant models with conditional computation and automatic sharding. InInternational Conference on Learning Representations (ICLR), 2021

  37. [45]

    The stability-efficiency dilemma: Investigating sequence length warmup for training GPT models.Advances in Neural Information Processing Systems (NeurIPS), 2022

    Conglong Li, Minjia Zhang, and Yuxiong He. The stability-efficiency dilemma: Investigating sequence length warmup for training GPT models.Advances in Neural Information Processing Systems (NeurIPS), 2022

  38. [46]

    PyTorch distributed: Experiences on accelerating data parallel training

    Shen Li, Yanli Zhao, Rohan Varma, Omkar Salpekar, Pieter Noordhuis, Teng Li, Adam Paszke, Jeff Smith, Brian Vaughan, Pritam Damania, and Soumith Chintala. PyTorch distributed: Experiences on accelerating data parallel training. InProceedings of the VLDB Endowment, 2020

  39. [47]

    TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training

    Wanchao Liang, Tianyu Liu, Less Wright, Will Constable, Andrew Gu, Chien-Chin Huang, Iris Zhang, Wei Feng, Howard Huang, Junjie Wang, Sanket Purandare, Gokul Nadathur, and Stratos Idreos. TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training. InInt...

  40. [48]

    LitGPT, 2023

    Lightning AI. LitGPT, 2023. URLhttps://github.com/Lightning-AI/litgpt. 16

  41. [49]

    RoBERTa: A robustly optimized BERT pretraining approach

    Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. RoBERTa: A robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692, 2019

  42. [50]

    The Llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024

    Llama Team, AI @ Meta. The Llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024

  43. [51]

    Decoupled weight decay regularization

    Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. InInternational Conference on Learning Representations (ICLR), 2019

  44. [52]

    An empirical model of large-batch training

    Sam McCandlish, Jared Kaplan, Dario Amodei, and OpenAI Dota Team. An empirical model of large-batch training. arXiv preprint arXiv:1812.06162, 2018

  45. [53]

    Efficient large-scale language model training on GPU clusters using Megatron-LM

    Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley, Mostofa Patwary, Vi- jay Anand Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, Amar Phanishayee, and Matei Zaharia. Efficient large-scale language model training on GPU cl...

  46. [54]

    AdaBatchGrad: Combining adaptive batch size and adaptive step size.arXiv preprint arXiv:2402.05264, 2024

    Petr Ostroukhov, Aigerim Zhumabayeva, Chulu Xiang, Alexander Gasnikov, Martin Takáč, and Dmitry Kamzolov. AdaBatchGrad: Combining adaptive batch size and adaptive step size.arXiv preprint arXiv:2402.05264, 2024

  47. [55]

    Toward understanding why Adam converges faster than SGD for transformers

    Yan Pan and Yuanzhi Li. Toward understanding why Adam converges faster than SGD for transformers. arXiv preprint arXiv:2306.00204, 2023

  48. [56]

    Nemotron-4 15B technical report.arXiv preprint arXiv:2402.16819, 2024

    Jupinder Parmar, Shrimai Prabhumoye, Joseph Jennings, Mostofa Patwary, Sandeep Subramanian, Dan Su, Chen Zhu, Deepak Narayanan, Aastha Jhunjhunwala, Ayush Dattagupta, Vibhu Jawa, Jiwei Liu, Ameya Mahabaleshwarkar, Osvald Nitski, Annika Brundyn, James Maki, Miguel Martinez, Jia...

  49. [57]

    PyTorch: An imperative style, high-performance deep learning library

    Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fan...

  50. [58]

    Large scale language modeling: Converging on 40GB of text in four hours

    Raul Puri, Robert Kirby, Nikolai Yakovenko, and Bryan Catanzaro. Large scale language modeling: Converging on 40GB of text in four hours. InProceedings of the International Symposium on Computer Architecture and High Performance Computing (SBAC-PAD), 2018

  51. [59]

    SimiGrad: Fine-grained adaptive batching for large scale training using gradient similarity measurement

    Heyang Qin, Samyam Rajbhandari, Olatunji Ruwase, Feng Yan, Lei Yang, and Yuxiong He. SimiGrad: Fine-grained adaptive batching for large scale training using gradient similarity measurement. InAdvances in Neural Information Processing Systems (NeurIPS), 2021

  52. [60]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67, 2020

  53. [61]

    ZeRO: Memory optimizations toward training trillion parameter models

    Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. ZeRO: Memory optimizations toward training trillion parameter models. InSC20: International Conference for High Performance Computing, Networking, Storage and Analysis, pages 1–16. IEEE, 2020

  54. [62]

    DeepSpeed: System optimizations enable training deep learning models with over 100 billion parameters

    Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. DeepSpeed: System optimizations enable training deep learning models with over 100 billion parameters. InProceedings of the ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2020. 17

  55. [63]

    ZeRO-Offload: Democratizing billion-scale model training

    Jie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, and Yuxiong He. ZeRO-Offload: Democratizing billion-scale model training. InUSENIX Annual Technical Conference (USENIX ATC), 2021

  56. [64]

    On the different regimes of stochastic gradient descent.Proceedings of the National Academy of Sciences, 121(9):e2316301121, 2024

    Antonio Sclocchi and Matthieu Wyart. On the different regimes of stochastic gradient descent.Proceedings of the National Academy of Sciences, 121(9):e2316301121, 2024

  57. [65]

    Shallue, Jaehoon Lee, Joseph Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E

    Christopher J. Shallue, Jaehoon Lee, Joseph Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E. Dahl. Measuring the effects of data parallelism on neural network training.Journal of Machine Learning Research, 20(112):1–49, 2019

  58. [66]

    Megatron-LM: Training multi-billion parameter language models using model parallelism.arXiv preprint arXiv:1909.08053, 2019

    Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. Megatron-LM: Training multi-billion parameter language models using model parallelism.arXiv preprint arXiv:1909.08053, 2019

  59. [67]

    Smith and Quoc V

    Samuel L. Smith and Quoc V. Le. A Bayesian perspective on generalization and stochastic gradient descent. InInternational Conference on Learning Representations (ICLR), 2018

  60. [68]

    Smith, Pieter-Jan Kindermans, and Quoc V

    Samuel L. Smith, Pieter-Jan Kindermans, and Quoc V. Le. Don’t decay the learning rate, increase the batch size. InInternational Conference on Learning Representations (ICLR), 2018

  61. [69]

    Using DeepSpeed and Megatron to train Megatron-Turing NLG 530B, a large-scale generative language model.arXiv preprint arXiv:2201.11990, 2022

    Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, Elton Zhang, Rewon Child, Reza Yazdani Aminabadi, Julie Bernauer, Xia Song, Mohammad Shoeybi, Yuxiong He, Michae...

  62. [70]

    Unraveling the mystery of scaling laws: Part I.arXiv preprint arXiv:2403.06563, 2024

    Hui Su, Zhi Tian, Xiaoyu Shen, and Xunliang Cai. Unraveling the mystery of scaling laws: Part I.arXiv preprint arXiv:2403.06563, 2024

  63. [71]

    Team OLMo, Pete Walsh, Luca Soldaini, Dirk Groeneveld, Kyle Lo, Shane Arora, Akshita Bhagia, Yuling Gu, Shengyi Huang, Matt Jordan, Nathan Lambert, Dustin Schwenk, Oyvind Tafjord, Taira Anderson, David Atkinson, Faeze Brahman, Christopher Clark, Pradeep Dasigi, Nouha Dziri, Mi...

  64. [72]

    Introducing DBRX: A new state-of-the-art open LLM

    The Mosaic Research Team. Introducing DBRX: A new state-of-the-art open LLM. https://www. databricks.com/blog/introducing-dbrx-new-state-art-open-llm , 2024

  65. [73]

    Llama 2: Open foundation and fine-tuned chat models.arXiv preprint arXiv:2307.09288, 2023

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, W...

  66. [74]

    Gomez, Łukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. InAdvances in Neural Information Processing Systems (NeurIPS), 2017. 18

  67. [75]

    Meta Lingua: A minimal PyTorch LLM training library, 2024

    Mathurin Videau, Badr Youbi Idrissi, Daniel Haziza, Luca Wehrstedt, Jade Copet, Olivier Teytaud, and David Lopez-Paz. Meta Lingua: A minimal PyTorch LLM training library, 2024. URLhttps: //github.com/facebookresearch/lingua

  68. [76]

    Closing the gap between the upper bound and lower bound of Adam’s iteration complexity

    Bohan Wang, Jingwen Fu, Huishuai Zhang, Nanning Zheng, and Wei Chen. Closing the gap between the upper bound and lower bound of Adam’s iteration complexity. InAdvances in Neural Information Processing Systems (NeurIPS), 2023

  69. [77]

    MicroLlama-300M

    Ken Wang. MicroLlama-300M. https://github.com/keeeeenw/MicroLlama, 2024

  70. [78]

    Liu, Lechao Xiao, Katie Everett, Alex Alemi, Ben Adlam, John D

    Mitchell Wortsman, Peter J. Liu, Lechao Xiao, Katie Everett, Alex Alemi, Ben Adlam, John D. Co-Reyes, Izzeddin Gur, Abhishek Kumar, Roman Novak, Jeffrey Pennington, Jascha Sohl-dickstein, Kelvin Xu, Jaehoon Lee, Justin Gilmer, and Simon Kornblith. Small-scale proxies for large...

  71. [79]

    Baichuan 2: Open large-scale language models.arXiv preprint arXiv:2309.10305, 2023

    Aiyuan Yang, Bin Xiao, Bingning Wang, Borong Zhang, Ce Bian, Chao Yin, Chenxu Lv, Da Pan, Dian Wang, Dong Yan, Fan Yang, Fei Deng, Feng Wang, Feng Liu, Guangwei Ai, Guosheng Dong, Haizhou Zhao, Hang Xu, Haoze Sun, Hongda Zhang, Hui Liu, Jiaming Ji, Jian Xie, JunTao Dai, Kun Fa...

  72. [80]

    Qwen2 technical report

    An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jianxin Yang, Jin Xu, Jingren Zhou, Jinze...

  73. [81]

    Large batch optimization for deep learning: Training BERT in 76 minutes

    Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh. Large batch optimization for deep learning: Training BERT in 76 minutes. InInternational Conference on Learning Representations (IC...

  74. [82]

    Susskind

    Shuangfei Zhai, Tatiana Likhomanenko, Etai Littwin, Dan Busbridge, Jason Ramapuram, Yizhe Zhang, Jiatao Gu, and Joshua M. Susskind. Stabilizing transformer training by preventing attention entropy collapse. In Proceedings of the International Conference on Machine Learning (IC...

  75. [83]

    How does critical batch size scale in pre-training? InInternational Conference on Learning Representations (ICLR), 2025

    Hanlin Zhang, Depen Morwani, Nikhil Vyas, Jingfeng Wu, Difan Zou, Udaya Ghai, Dean Foster, and Sham Kakade. How does critical batch size scale in pre-training? InInternational Conference on Learning Representations (ICLR), 2025

  76. [84]

    TinyLlama: An open-source small language model

    Peiyuan Zhang, Guangtao Zeng, Tianduo Wang, and Wei Lu. TinyLlama: An open-source small language model. arXiv preprint arXiv:2401.02385, 2024

  77. [85]

    OPT: Open pre-trained transformer language models.arXiv preprint arXiv:2205.01068, 2022

    Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. O...

  78. [86]

    Why transformers need Adam: A Hessian perspective.arXiv preprint arXiv:2402.16788, 2024

    Yushun Zhang, Congliang Chen, Tian Ding, Ziniu Li, Ruoyu Sun, and Zhi-Quan Luo. Why transformers need Adam: A Hessian perspective.arXiv preprint arXiv:2402.16788, 2024

  79. [87]

    Picotron: Distributed training framework for education and research experimentation, 2025

    Haojun Zhao and Ferdinand Mom. Picotron: Distributed training framework for education and research experimentation, 2025. URL https://github.com/huggingface/picotron. 19

  80. [88]

    first-order term

    Yanli Zhao, Andrew Gu, Rohan Varma, Liang Luo, Chien-Chin Huang, Min Xu, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, Alban Desmaison, Can Balioglu, Pritam Damania, Bernard Nguyen, Geeta Chauhan, Yuchen Hao, Ajit Mathews, and Shen Li. PyTorch FSDP: Experiences on sc...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.