Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T23:08:28.873206Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 88 of 88 outbound references and 1 inbound Pith citation observation for arXiv:2412.21124.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T23:08:28.873206Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-10T18:38:58.883757Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T00:10:53.383914Z
88 of 88 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b6810cf5-1d1c-4834-9b3b-ea411c930456 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4d1e638-3900-4a90-86bc-ff63b3adc3cc · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Extremely Large Minibatch SGD: Training ResNet-50 on ImageNet in 15 Minutes
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84a20f5c-bdf2-4da6-9d0d-6056e3d8527e · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Qwen Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e363a1bf-ad3a-4431-a9ee-ebf276a9553a · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Coupling adaptive batch sizes with learning rates
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7be17e8-052a-492e-8dfd-c091eb85b34f · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Stable LM 2 1.6B Technical Report
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de41598d-de41-4249-9e38-d6153b205e8a · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48f87ad8-556f-456e-8af3-b79a5c0de339 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0154dee3-e45b-4bc6-9507-a01fcdd8e549 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Adaptive sampling strategies for stochastic optimization
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e54d7e66-2702-49f4-a396-a9a7e2f8d840 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Curtis, and Jorge Nocedal
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23065d4c-fbea-43bd-b372-852a508c27d4 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism JAX: composable transformations of Python+NumPy programs, 2018
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0caed567-023f-4a1d-8a3b-1d2683f1dd07 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0eeab79-a6bc-47b0-b2fc-8445b74c20f9 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Byrd, Gillian M
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6be660f-e3e5-4157-b5c8-1d6ac47cf887 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Big Batch SGD: Automated Inference using Adaptive Batch Sizes
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35e837c6-2948-45ac-95f5-cd22a7fe7355 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Automated inference with adaptive batches
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4b30ab26-4207-4482-920c-93f59be5f71e · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15bc84be-5671-4d28-8a77-668b9b38bd04 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism DeepSeek-V3 Technical Report
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 987fb9ad-f1b0-40b2-8a3d-0e4f7f93e3ab · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism AdaBatch: Adaptive Batch Sizes for Training Deep Neural Networks
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13543c64-9340-4ab2-b218-39a27f7781cf · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism An image is worth 16x16 words: Transformers for image recognition at scale
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b961533f-5b32-49cb-a584-d36ecb1e2447 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Susskind, and Armand Joulin
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 29e69ced-3080-4adc-8562-8aea5b81ec8c · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism PyTorch Lightning, 2019
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 38fb4cec-205f-4ba5-b869-2103668853ad · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Friedlander and Mark Schmidt
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 13baffce-448c-4f9c-9c6c-e347e157cb1a · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Compiling machine learning programs via high-level tracing
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation cf72c474-a5cb-42c7-9018-3fb9dcfd2392 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Gemma 2: Improving Open Language Models at a Practical Size
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ac85a57-6330-4013-bb12-7df6911a50a6 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Gemma: Open Models Based on Gemini Research and Technology
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3ed2a93-05a1-4f3d-aa6c-667d7e12fdf4 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism OpenLLaMA: An open reproduction of LLaMA, May 2023
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c6ef5449-5106-4836-bb71-1b19318437a1 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac423a62-c31f-41a4-906f-d01ac3142681 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism OLMo: Accelerating the Science of Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5c3fcb0-70c9-4acf-9a10-3e43e0d12e93 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Scaling laws for single-agent reinforcement learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51935083-cbe8-4bab-932a-60ee23cdffa6 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 78be173c-05ad-4d19-b32c-bd48f41bc3c7 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Mistral 7B
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17c167d4-30b6-4bdd-819a-27934a167b44 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism GrowLength: Accelerating LLMs Pretraining by Progressively Growing Training Length
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1760b5be-98d9-4526-b657-3081a272bab5 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism AdaScale SGD: A user-friendly algorithm for distributed training
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f48c3594-b81e-4b68-8d15-4c9f33f392e6 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Scaling Laws for Neural Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17647574-453a-49ec-af1d-8bfb0abe4257 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism On large-batch training for deep learning: Generalization gap and sharp minima
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a94984a0-e9fb-43e4-b2d8-10affe2a0b38 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Better theory for SGD in the nonconvex world.Transactions on Machine Learning Research, 2023
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5e987de-2f6f-4743-89b1-93ad0cfb3414 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Kingma and Jimmy Lei Ba
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 347cc37f-44ae-4a79-b544-a41c08c30de7 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Reducing activation recomputation in large transformer models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6c6e722c-e118-4edf-b158-1dec81a7baf7 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c3a741b7-853d-4ea0-a2b0-fe8418af1843 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Noise is not the main factor behind the gap between SGD and Adam on transformers, but sign descent might be
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2f3cd7b1-45a6-4f41-be3f-4401f52edc0a · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6560157b-eab2-4e15-89df-67a2ec763523 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Communication-Efficient Adaptive Batch Size Strategies for Distributed Local Gradient Methods
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2928c122-0501-4787-9a67-f490766b7b7a · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism AdAdaGrad: Adaptive Batch Size Schemes for Adaptive Gradient Methods
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0085f04e-753d-40ad-9c68-87cb0bf298a9 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Orr, and Klaus Robert Müller
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 38d7e523-8b65-408a-bd78-b257912d9c85 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism GShard: Scaling giant models with conditional computation and automatic sharding
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8fd0fe8-57ef-4aa5-92c7-0ac612b66797 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism The stability-efficiency dilemma: Investigating sequence length warmup for training GPT models.Advances in Neural Information Processing Systems (NeurIPS), 2022
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation fd20966e-19e9-431f-9293-2f6afc57c197 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism PyTorch distributed: Experiences on accelerating data parallel training
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 80ff2617-719a-41af-9d88-a06457a4f1c6 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 58b284de-9ed3-4d16-8f14-8f780d34b680 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism LitGPT, 2023
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9a41a052-d67d-4158-85f3-5931beda7774 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism RoBERTa: A Robustly Optimized BERT Pretraining Approach
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c644634f-0da3-4edd-82d2-d513fe3408c7 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism The Llama 3 Herd of Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34f011ae-18a7-4607-a505-725a14048efd · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Decoupled weight decay regularization
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d34ffa08-e6a3-40e1-b3e0-5cd1ba38d35a · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism An Empirical Model of Large-Batch Training
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 074aab96-b4b6-4f21-aac4-0bb83b2515f9 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Efficient large-scale language model training on GPU clusters using Megatron-LM
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 22200e34-27dc-4aa2-874f-5bee70eadc5f · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation de7bce79-ad55-4a69-8474-051b04a9546c · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Toward Understanding Why Adam Converges Faster Than SGD for Transformers
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7f506ca-a3ad-4760-862d-0eda79403039 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Nemotron-4 15B Technical Report
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a8dd1b2-4148-4ef9-b20f-d797a19836dc · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism PyTorch: An imperative style, high-performance deep learning library
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4137c201-4efe-4573-8233-8ca7b0e2ad11 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Large scale language modeling: Converging on 40GB of text in four hours
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation bf88acb2-dc51-432b-a251-2a59cc317473 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism SimiGrad: Fine-grained adaptive batching for large scale training using gradient similarity measurement
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9dc0ab40-ecdb-4c15-adbb-fb68b7f9e0b9 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Unresolved cited work
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f68f0b6-4b00-4022-8e2e-3133f6e05054 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism ZeRO: Memory optimizations toward training trillion parameter models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 435361fa-002d-4e5c-8082-e025b9ff84c9 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism DeepSpeed: System optimizations enable training deep learning models with over 100 billion parameters
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 413252af-5483-45ae-801c-a90863ecd40b · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism ZeRO-Offload: Democratizing billion-scale model training
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 284f4916-6eaa-4116-a4f4-5957e00e21ae · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism On the different regimes of stochastic gradient descent.Proceedings of the National Academy of Sciences, 121(9):e2316301121, 2024
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b5b9e4a3-0569-436b-837d-6e4b093c86d9 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Shallue, Jaehoon Lee, Joseph Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2c58e921-f588-416c-9b0f-8721c32559fe · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63a15759-513c-44eb-a4a2-e140610e06c2 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Smith and Quoc V
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 234a99ec-997c-4c5c-81e8-00d68670c9f9 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Smith, Pieter-Jan Kindermans, and Quoc V
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 366ce334-1037-486c-a24f-3c327afabc8c · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6140cea0-75ee-42cf-ad62-feac0f1759e2 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Unraveling the Mystery of Scaling Laws: Part I
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40cd2d9e-5b48-473f-8c59-f296d3c16d4b · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism 2 OLMo 2 Furious
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a15bb9d-3ca6-4f94-9491-526cf29f2130 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Introducing DBRX: A new state-of-the-art open LLM
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 68c1c185-1b47-43b3-a13b-42f30562146f · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff6ea462-df90-4351-9f61-f16f054ce4e9 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Gomez, Łukasz Kaiser, and Illia Polosukhin
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 38068dd4-107f-4687-83b4-8cf6ff6ba778 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Meta Lingua: A minimal PyTorch LLM training library, 2024
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1f15343f-3aa1-4764-a255-45564d336b2c · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Closing the gap between the upper bound and lower bound of Adam’s iteration complexity
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 86ddbe66-ccb9-4fec-8b2f-07de7a059f34 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism MicroLlama-300M
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ed4e73fc-e733-4218-9e62-787e408a693e · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Liu, Lechao Xiao, Katie Everett, Alex Alemi, Ben Adlam, John D
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b954a9c7-ac0e-4775-8057-031783abea46 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Baichuan 2: Open Large-scale Language Models
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52d360f0-e63b-4c57-abf8-3f126710905b · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Qwen2 Technical Report
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cdb2a43-e351-4882-944a-b155be40774b · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Large batch optimization for deep learning: Training BERT in 76 minutes
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 041b3c68-5256-4840-819c-0e3d2633bbbc · outbound
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation aca9d4ba-94cd-475e-9e9b-83fb0df939aa · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism How does critical batch size scale in pre-training? InInternational Conference on Learning Representations (ICLR), 2025
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c1b70b2a-0fa1-4414-9fd0-5e12aa5c7f3a · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism TinyLlama: An Open-Source Small Language Model
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a79759a-0bd9-41a0-986e-8732853ebe77 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism OPT: Open Pre-trained Transformer Language Models
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a1a938c-4694-44ef-8f10-c7c54ec86094 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Why Transformers Need Adam: A Hessian Perspective
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26383e5e-bdd6-4643-b43f-11e84065244a · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Picotron: Distributed training framework for education and research experimentation, 2025
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3103d643-bf42-404a-9da8-2952792fec22 · outbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism first-order term
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3ef300f3-0353-4936-a83f-cac4e59f93c4 · inbound
Convergence of Riemannian Stochastic Gradient Descents: Varying Batch Sizes And Nonstandard Batch Forming Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.