Pith. sign in

Paper Citation Record · LEDGER

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism

As of 16 August 2026, this Paper Citation Record lists 88 of 88 outbound references and 1 inbound Pith citation observation for arXiv:2412.21124.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.21124 v2

Coverage vector

measured 88 of 88 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:08:28.873206Z

measured 89 of 89 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T18:38:58.883757Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T00:10:53.383914Z

Reference resolution

88 of 88 outbound references displayed

  • verified exact2
  • verified fuzzy39
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b6810cf5-1d1c-4834-9b3b-ea411c930456 · outbound

This paper cites an unresolved cited work.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:26.118184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:26.118184Z digest=sha256:09201d05f7867b027324850a24ec8333a32e2bce22a4fb0cd64b46636a5b28f1

Observation b4d1e638-3900-4a90-86bc-ff63b3adc3cc · outbound

This paper cites Extremely Large Minibatch SGD: Training ResNet-50 on ImageNet in 15 Minutes.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Extremely Large Minibatch SGD: Training ResNet-50 on ImageNet in 15 Minutes

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:26.212271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:26.212271Z digest=sha256:1ea7f0f87bd8ac41a4efeb6db2941c74b7f506bd25a5f89a75c705aa4a209afc

Observation 84a20f5c-bdf2-4da6-9d0d-6056e3d8527e · outbound

This paper cites Qwen Technical Report.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:26.262650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:26.262650Z digest=sha256:2831f1173eb73387c9de37c4a7826123742826e5e048f0aafd85bc0252c866bc

Observation e363a1bf-ad3a-4431-a9ee-ebf276a9553a · outbound

This paper cites Coupling adaptive batch sizes with learning rates.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Coupling adaptive batch sizes with learning rates

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:26.267815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:26.267815Z digest=sha256:ac25c89b3de1d19b10d3248d0391df98580e1551eafd8880eaa99ccc30e21533

Observation e7be17e8-052a-492e-8dfd-c091eb85b34f · outbound

This paper cites Stable LM 2 1.6B Technical Report.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Stable LM 2 1.6B Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:26.271570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:26.271570Z digest=sha256:5039f9077ea2d085e56724c3cf7bd343f3a2b78a94fc6e4030255f377b935595

Observation de41598d-de41-4249-9e38-d6153b205e8a · outbound

This paper cites BLOOM: A 176B-Parameter Open-Access Multilingual Language Model.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:26.276815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:26.276815Z digest=sha256:d8da463b4c2d6f4009f163612247b1446ce24085705fcfd80676d184560072ff

Observation 48f87ad8-556f-456e-8af3-b79a5c0de339 · outbound

This paper cites an unresolved cited work.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:26.312733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:26.312733Z digest=sha256:a0dd1753d66bcce344f3ec6fc48bdb80014e29aa039f5f3d70ae46c2557fb8d5

Observation 0154dee3-e45b-4bc6-9507-a01fcdd8e549 · outbound

This paper cites Adaptive sampling strategies for stochastic optimization.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Adaptive sampling strategies for stochastic optimization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:26.440847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:26.440847Z digest=sha256:dc58758859f6155340fad49a5ce792fe492644d539f6e5a9479e1c0737863a15

Observation e54d7e66-2702-49f4-a396-a9a7e2f8d840 · outbound

This paper cites Curtis, and Jorge Nocedal.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Curtis, and Jorge Nocedal

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:26.486876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:26.486876Z digest=sha256:dd7cde4979a3ff52e410072da130b21a20700f60e5b1ff608202ca4fb5defaa2

Observation 23065d4c-fbea-43bd-b372-852a508c27d4 · outbound

This paper cites JAX: composable transformations of Python+NumPy programs, 2018.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism JAX: composable transformations of Python+NumPy programs, 2018

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:26.491193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:26.491193Z digest=sha256:7e58b999b28567ca346bc740aeaa0c7644085a4bcbfc8101b1277e9d5add1165

Observation 0caed567-023f-4a1d-8a3b-1d2683f1dd07 · outbound

This paper cites an unresolved cited work.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:26.495360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:26.495360Z digest=sha256:38ddba2636e27d22b02324cdf53d4d5b47088321c7b9b84a1653f1488fdb8739

Observation d0eeab79-a6bc-47b0-b2fc-8445b74c20f9 · outbound

This paper cites Byrd, Gillian M.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Byrd, Gillian M

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:26.501738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:26.501738Z digest=sha256:dab23fba80a7904890b50fabcc4f2b2ec703394d0e0324ed96b4e1f057e9f221

Observation c6be660f-e3e5-4157-b5c8-1d6ac47cf887 · outbound

This paper cites Big Batch SGD: Automated Inference using Adaptive Batch Sizes.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Big Batch SGD: Automated Inference using Adaptive Batch Sizes

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:26.505657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:26.505657Z digest=sha256:4815211d4adddf7ece1f59c13e1c70759bf2714e4a029ad5a788624b9d50be3c

Observation 35e837c6-2948-45ac-95f5-cd22a7fe7355 · outbound

This paper cites Automated inference with adaptive batches.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Automated inference with adaptive batches

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:31.508603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:26.510052Z digest=sha256:ec04427ab0fa5b9db06d39e7efd40f10d669808c921a602746226f28f919da02

Observation 4b30ab26-4207-4482-920c-93f59be5f71e · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:26.514375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:26.514375Z digest=sha256:2f4d80ff4ad2a4d113bfbafbd12a39a791dc7d972ba42d95fe7d132a103ed36c

Observation 15bc84be-5671-4d28-8a77-668b9b38bd04 · outbound

This paper cites DeepSeek-V3 Technical Report.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism DeepSeek-V3 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:26.574840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:26.574840Z digest=sha256:14bb97bbd84aad99835243ed9b3324e0fb31e9313cf5c254ec4b66f1f20d2696

Observation 987fb9ad-f1b0-40b2-8a3d-0e4f7f93e3ab · outbound

This paper cites AdaBatch: Adaptive Batch Sizes for Training Deep Neural Networks.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism AdaBatch: Adaptive Batch Sizes for Training Deep Neural Networks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:26.676251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:26.676251Z digest=sha256:68027d0bc8935a4da5f58a72e254b7d6a102e5124ee70a7b42bba8eb07262825

Observation 13543c64-9340-4ab2-b218-39a27f7781cf · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism An image is worth 16x16 words: Transformers for image recognition at scale

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:27.007580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:27.007580Z digest=sha256:3c9117e58bc5fd041fd59cba0142663d05c350dd244de99a00db71163bb712af

Observation b961533f-5b32-49cb-a584-d36ecb1e2447 · outbound

This paper cites Susskind, and Armand Joulin.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Susskind, and Armand Joulin

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:31.486551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:27.787032Z digest=sha256:513a8825e16bcd9b551485c6e45eb19db016ec7e89a8c63e22fe114f5e47568d

Observation 29e69ced-3080-4adc-8562-8aea5b81ec8c · outbound

This paper cites PyTorch Lightning, 2019.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism PyTorch Lightning, 2019

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:31.416593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:27.838868Z digest=sha256:354a1a0052751f0da4d69db0b17f85a060404b0bcad7bd091ad48fabebd572b9

Observation 38fb4cec-205f-4ba5-b869-2103668853ad · outbound

This paper cites Friedlander and Mark Schmidt.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Friedlander and Mark Schmidt

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:31.402273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:27.843948Z digest=sha256:be3b96455a3556fb59cb2ddda9db3af6c42e61c73b3e7e2e5b5337350ca6cb09

Observation 13baffce-448c-4f9c-9c6c-e347e157cb1a · outbound

This paper cites Compiling machine learning programs via high-level tracing.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Compiling machine learning programs via high-level tracing

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:31.303387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:27.849740Z digest=sha256:ba04c663a51ac942ff1e2d5e0b23aae8993aef34afdc7254c5afec99d4b17ac5

Observation cf72c474-a5cb-42c7-9018-3fb9dcfd2392 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Gemma 2: Improving Open Language Models at a Practical Size

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:27.854625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:27.854625Z digest=sha256:7fdf09d5c9c188a239dd5799a2675741e4411a02a5fab1cf309d79d3326a23a7

Observation 9ac85a57-6330-4013-bb12-7df6911a50a6 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Gemma: Open Models Based on Gemini Research and Technology

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:27.860027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:27.860027Z digest=sha256:be929333d7b8596c82c9ca84b06203b78bd37e35e9bf80a4a91ee760d201431b

Observation a3ed2a93-05a1-4f3d-aa6c-667d7e12fdf4 · outbound

This paper cites OpenLLaMA: An open reproduction of LLaMA, May 2023.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism OpenLLaMA: An open reproduction of LLaMA, May 2023

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:31.288817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:27.863843Z digest=sha256:f4e9f53dffe80bbd00d1b94271f927507ed5e75bb24ba80d78da11b6c535c8c9

Observation c6ef5449-5106-4836-bb71-1b19318437a1 · outbound

This paper cites Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:27.940336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:27.940336Z digest=sha256:c674ecececfb2abce779b73296a3d57a30873809b0a040a5321279680699a079

Observation ac423a62-c31f-41a4-906f-d01ac3142681 · outbound

This paper cites OLMo: Accelerating the Science of Language Models.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism OLMo: Accelerating the Science of Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:27.945364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:27.945364Z digest=sha256:c475115492d61bd2a36df8aaeb0ca1aa9d3cb5f1706f80e9d755f81849781698

Observation b5c3fcb0-70c9-4acf-9a10-3e43e0d12e93 · outbound

This paper cites Scaling laws for single-agent reinforcement learning.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Scaling laws for single-agent reinforcement learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:27.949748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:27.949748Z digest=sha256:ef7448464e018c630091646d88e458bdd36a257d5fbde8f700da92d8ead6e5f9

Observation 51935083-cbe8-4bab-932a-60ee23cdffa6 · outbound

This paper cites Train longer, generalize better: closing the generalization gap in large batch training of neural networks.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Train longer, generalize better: closing the generalization gap in large batch training of neural networks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:31.275809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:27.954773Z digest=sha256:baecf5090be79151295fb957623e986f5bd8a6b5a0562568acf2be7699e88566

Observation 78be173c-05ad-4d19-b32c-bd48f41bc3c7 · outbound

This paper cites Mistral 7B.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Mistral 7B

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:27.958688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:27.958688Z digest=sha256:d4eb267ee7123cce00fc45e2c6acbba064d87c9ddecd96f7176cb2ea5c3502ed

Observation 17c167d4-30b6-4bdd-819a-27934a167b44 · outbound

This paper cites GrowLength: Accelerating LLMs Pretraining by Progressively Growing Training Length.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism GrowLength: Accelerating LLMs Pretraining by Progressively Growing Training Length

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:08:29.256142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:27.964430Z digest=sha256:ba940029571bcab1f2f885351a9833a43351e0cc720fc13847c431f31c2e8293

Observation 1760b5be-98d9-4526-b657-3081a272bab5 · outbound

This paper cites AdaScale SGD: A user-friendly algorithm for distributed training.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism AdaScale SGD: A user-friendly algorithm for distributed training

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:31.172939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.018228Z digest=sha256:9a288405bd415630c55d71bdb2d6fe95b6190906c3dafdcf1043cbd3df4222f9

Observation f48c3594-b81e-4b68-8d15-4c9f33f392e6 · outbound

This paper cites Scaling Laws for Neural Language Models.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Scaling Laws for Neural Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.022357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.022357Z digest=sha256:d4e005c5ba5bc7ab2d5755596800416297a61b9c4608b7f967af3920f1685607

Observation 17647574-453a-49ec-af1d-8bfb0abe4257 · outbound

This paper cites On large-batch training for deep learning: Generalization gap and sharp minima.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism On large-batch training for deep learning: Generalization gap and sharp minima

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:31.146639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.026043Z digest=sha256:98901a5a357243206d70c0679e3b05a3f209bd4991c92dc99bbd8e76e42192b9

Observation a94984a0-e9fb-43e4-b2d8-10affe2a0b38 · outbound

This paper cites Better theory for SGD in the nonconvex world.Transactions on Machine Learning Research, 2023.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Better theory for SGD in the nonconvex world.Transactions on Machine Learning Research, 2023

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.030194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.030194Z digest=sha256:ef940a05d241093c2c277efc32b7f576c24173760b88e67f45b14e05a6683867

Observation a5e987de-2f6f-4743-89b1-93ad0cfb3414 · outbound

This paper cites Kingma and Jimmy Lei Ba.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Kingma and Jimmy Lei Ba

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:31.060307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.033499Z digest=sha256:3ac8441da9ba297709ef04ee8bd9ac2b11d524ed85973be56d09d4abacc8ec06

Observation 347cc37f-44ae-4a79-b544-a41c08c30de7 · outbound

This paper cites Reducing activation recomputation in large transformer models.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Reducing activation recomputation in large transformer models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:31.045296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.037092Z digest=sha256:09e6242fb381f6b0665b911f2a01bab7e230c9a018ebb5101b0ace1087749d00

Observation 6c6e722c-e118-4edf-b158-1dec81a7baf7 · outbound

This paper cites an unresolved cited work.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:08:30.949976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.041594Z digest=sha256:f5bf204a44d4f9c80273ec1b42d2a6ed6faebc32c2a474a5012c4c5f8869e9ab

Observation c3a741b7-853d-4ea0-a2b0-fe8418af1843 · outbound

This paper cites Noise is not the main factor behind the gap between SGD and Adam on transformers, but sign descent might be.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Noise is not the main factor behind the gap between SGD and Adam on transformers, but sign descent might be

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:30.936710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.086417Z digest=sha256:fa39705693c972b5bee5a48379ce9e8dadeff35002f3f33252e17b9f0c562ab2

Observation 2f3cd7b1-45a6-4f41-be3f-4401f52edc0a · outbound

This paper cites Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.116499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.116499Z digest=sha256:4adf24a8cb36e3754b7765eb4f8fab499c156d6684ff79c148bafaaa5fd44a3e

Observation 6560157b-eab2-4e15-89df-67a2ec763523 · outbound

This paper cites Communication-Efficient Adaptive Batch Size Strategies for Distributed Local Gradient Methods.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Communication-Efficient Adaptive Batch Size Strategies for Distributed Local Gradient Methods

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.119718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.119718Z digest=sha256:49686676fb1975e7c574c263ec2b70ba94a486276a88aff75ed7908cd931f686

Observation 2928c122-0501-4787-9a67-f490766b7b7a · outbound

This paper cites AdAdaGrad: Adaptive Batch Size Schemes for Adaptive Gradient Methods.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism AdAdaGrad: Adaptive Batch Size Schemes for Adaptive Gradient Methods

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.123888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.123888Z digest=sha256:07b558e237bc4c3af49346e1fcdb6f9ae71e340340b89d0c1ebdc5c162ea7ec0

Observation 0085f04e-753d-40ad-9c68-87cb0bf298a9 · outbound

This paper cites Orr, and Klaus Robert Müller.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Orr, and Klaus Robert Müller

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:30.901128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.128430Z digest=sha256:a2a7fbf72174a2726ca5da48d5b2d75d549d6a7011457b52f2e16e44c0570b77

Observation 38d7e523-8b65-408a-bd78-b257912d9c85 · outbound

This paper cites GShard: Scaling giant models with conditional computation and automatic sharding.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism GShard: Scaling giant models with conditional computation and automatic sharding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.132217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.132217Z digest=sha256:f737ff24339b88113f56f821d39fc241785347102c0a5076a927008880bd08b8

Observation e8fd0fe8-57ef-4aa5-92c7-0ac612b66797 · outbound

This paper cites The stability-efficiency dilemma: Investigating sequence length warmup for training GPT models.Advances in Neural Information Processing Systems (NeurIPS), 2022.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism The stability-efficiency dilemma: Investigating sequence length warmup for training GPT models.Advances in Neural Information Processing Systems (NeurIPS), 2022

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:30.788644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.135256Z digest=sha256:3d6e68091b0f9f091cd978969a01f130aba4311b1de1a85ff5f7e6a59afb18d4

Observation fd20966e-19e9-431f-9293-2f6afc57c197 · outbound

This paper cites PyTorch distributed: Experiences on accelerating data parallel training.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism PyTorch distributed: Experiences on accelerating data parallel training

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:30.776699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.139259Z digest=sha256:86054592c74eccdeb7881a8d547d830c8b82a294c3761ebbbbffd5a9e9a0f1f7

Observation 80ff2617-719a-41af-9d88-a06457a4f1c6 · outbound

This paper cites TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:30.762450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.142617Z digest=sha256:b2e0fb28852724662f4418acc835efd074c55baab01e6a1be12e690124768681

Observation 58b284de-9ed3-4d16-8f14-8f780d34b680 · outbound

This paper cites LitGPT, 2023.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism LitGPT, 2023

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:30.685452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.191665Z digest=sha256:009dfd6af0bf3818348faa990d7b917b8d48011633f8728b0b5af4ac77469c88

Observation 9a41a052-d67d-4158-85f3-5931beda7774 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.197709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.197709Z digest=sha256:f82c0efa762ff1e6918e4f8464818998c16d521a97df30e704192c0828c0af52

Observation c644634f-0da3-4edd-82d2-d513fe3408c7 · outbound

This paper cites The Llama 3 Herd of Models.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism The Llama 3 Herd of Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.203669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.203669Z digest=sha256:a3f58008f73f495376d7270c012523565bb293a6626e41341a6ec2f2c7369ca9

Observation 34f011ae-18a7-4607-a505-725a14048efd · outbound

This paper cites Decoupled weight decay regularization.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Decoupled weight decay regularization

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.208234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.208234Z digest=sha256:2c23a59895d93a4e20def89d5a86ba97cc0ece7d1400e6a4d6377c4262d9f813

Observation d34ffa08-e6a3-40e1-b3e0-5cd1ba38d35a · outbound

This paper cites An Empirical Model of Large-Batch Training.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism An Empirical Model of Large-Batch Training

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.211454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.211454Z digest=sha256:e1c8e00f743e1bf8d47872975c55efc29cb7752ce01401da6b3369203b0cda5b

Observation 074aab96-b4b6-4f21-aac4-0bb83b2515f9 · outbound

This paper cites Efficient large-scale language model training on GPU clusters using Megatron-LM.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Efficient large-scale language model training on GPU clusters using Megatron-LM

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:30.563989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.215521Z digest=sha256:998ffc7711207e9449681ca96091401520497490187238ff100c0e0977a809f4

Observation 22200e34-27dc-4aa2-874f-5bee70eadc5f · outbound

This paper cites AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:08:29.105925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.219044Z digest=sha256:82bc42616c25b1fb00e874b7fadf8c3f26890b8d5982d6b06478873c8e3fac16

Observation de7bce79-ad55-4a69-8474-051b04a9546c · outbound

This paper cites Toward Understanding Why Adam Converges Faster Than SGD for Transformers.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Toward Understanding Why Adam Converges Faster Than SGD for Transformers

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.222610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.222610Z digest=sha256:55de91f5a580c7a62cd2535904467732221ece7f4a418b5bdf5f4b247fbb6b93

Observation a7f506ca-a3ad-4760-862d-0eda79403039 · outbound

This paper cites Nemotron-4 15B Technical Report.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Nemotron-4 15B Technical Report

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.313606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.313606Z digest=sha256:7d363f2834de71b21bab2d54c4ba9c55d2d742301954a2f21a0f756bd9c8ec77

Observation 8a8dd1b2-4148-4ef9-b20f-d797a19836dc · outbound

This paper cites PyTorch: An imperative style, high-performance deep learning library.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism PyTorch: An imperative style, high-performance deep learning library

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:30.479112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.354771Z digest=sha256:cebee3620fe067b4ee30ef1a729a4d857d95ae7f2630eccfe548cd122dbfc396

Observation 4137c201-4efe-4573-8233-8ca7b0e2ad11 · outbound

This paper cites Large scale language modeling: Converging on 40GB of text in four hours.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Large scale language modeling: Converging on 40GB of text in four hours

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:30.441402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.359098Z digest=sha256:d998fa22d0bb5056e0dcf0c0a5f212820033b66cb194c743914d971f1fdd2d27

Observation bf88acb2-dc51-432b-a251-2a59cc317473 · outbound

This paper cites SimiGrad: Fine-grained adaptive batching for large scale training using gradient similarity measurement.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism SimiGrad: Fine-grained adaptive batching for large scale training using gradient similarity measurement

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:30.356127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.362697Z digest=sha256:6dc58372d650834d2495bee4faed6e3b6efef128623fecfbc70ba46055f4f1c4

Observation 9dc0ab40-ecdb-4c15-adbb-fb68b7f9e0b9 · outbound

This paper cites an unresolved cited work.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.367120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.367120Z digest=sha256:ea989754ecb9c7d095fba69d201918e054f102bdd48697f5e3f5ce7c223385a6

Observation 6f68f0b6-4b00-4022-8e2e-3133f6e05054 · outbound

This paper cites ZeRO: Memory optimizations toward training trillion parameter models.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism ZeRO: Memory optimizations toward training trillion parameter models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:30.334289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.370025Z digest=sha256:5619981bed25d9d72004f8dae42c9d15b3e1c3cd4bd537839737c29bc3eb7a45

Observation 435361fa-002d-4e5c-8082-e025b9ff84c9 · outbound

This paper cites DeepSpeed: System optimizations enable training deep learning models with over 100 billion parameters.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism DeepSpeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:30.275184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.373751Z digest=sha256:3a45ca06390443f3ef73a6d4288cf64fc97f51aa6c78ae51a503c51b2e1d0301

Observation 413252af-5483-45ae-801c-a90863ecd40b · outbound

This paper cites ZeRO-Offload: Democratizing billion-scale model training.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism ZeRO-Offload: Democratizing billion-scale model training

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:30.232150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.378205Z digest=sha256:e022f4f5b1fb941b41902ae52459331762fc70ad91c799f5a4f7416100b16ec4

Observation 284f4916-6eaa-4116-a4f4-5957e00e21ae · outbound

This paper cites On the different regimes of stochastic gradient descent.Proceedings of the National Academy of Sciences, 121(9):e2316301121, 2024.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism On the different regimes of stochastic gradient descent.Proceedings of the National Academy of Sciences, 121(9):e2316301121, 2024

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:30.219609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.382219Z digest=sha256:ff6c13b71ec76a835c029ab255131225f9ed14f0cf064c77764b9562fc863ff7

Observation b5b9e4a3-0569-436b-837d-6e4b093c86d9 · outbound

This paper cites Shallue, Jaehoon Lee, Joseph Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Shallue, Jaehoon Lee, Joseph Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:30.125126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.401956Z digest=sha256:383f46fc5491930a2590b45e42a7bdab11921f7d577c376cf15c51c6e955f585

Observation 2c58e921-f588-416c-9b0f-8721c32559fe · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.459144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.459144Z digest=sha256:98f1a6f0fe719dcefc3737aa974d266e2ac8b2250807ddc76f4dfb3f75a4e8b5

Observation 63a15759-513c-44eb-a4a2-e140610e06c2 · outbound

This paper cites Smith and Quoc V.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Smith and Quoc V

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:30.031318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.513312Z digest=sha256:40362e4ba1353d2069bbc44b7142b3ba7d6b2074c48f25284018c1f84119fe5b

Observation 234a99ec-997c-4c5c-81e8-00d68670c9f9 · outbound

This paper cites Smith, Pieter-Jan Kindermans, and Quoc V.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Smith, Pieter-Jan Kindermans, and Quoc V

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:29.994206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.518013Z digest=sha256:b2b4aeb7536d9b03fe349f563d9672c2798d589105f5bca72dce42f20f173303

Observation 366ce334-1037-486c-a24f-3c327afabc8c · outbound

This paper cites Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.522233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.522233Z digest=sha256:397781c0e1c8f312c1299b9931fbe0f7b38d19aba58515c63da3c9daac1b0914

Observation 6140cea0-75ee-42cf-ad62-feac0f1759e2 · outbound

This paper cites Unraveling the Mystery of Scaling Laws: Part I.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Unraveling the Mystery of Scaling Laws: Part I

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.527070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.527070Z digest=sha256:9b67740a67c876839cb6893f7ab3ca585e6538d08b1a3698e67f866b957dc097

Observation 40cd2d9e-5b48-473f-8c59-f296d3c16d4b · outbound

This paper cites 2 OLMo 2 Furious.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism 2 OLMo 2 Furious

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.532201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.532201Z digest=sha256:cf87ef39c637e607c94eae243bc6d73368e962cd07860ad19172c9e712fd45ab

Observation 3a15bb9d-3ca6-4f94-9491-526cf29f2130 · outbound

This paper cites Introducing DBRX: A new state-of-the-art open LLM.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Introducing DBRX: A new state-of-the-art open LLM

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:29.973315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.536308Z digest=sha256:85cbf3ac3a72eef40e971c062582a9c5206bb93cbbeb1506223a13d046a6566b

Observation 68c1c185-1b47-43b3-a13b-42f30562146f · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.626060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.626060Z digest=sha256:a34b5dd397772a62453736854c9d6625963f2051dde5222b3a6a161c90d35bf3

Observation ff6ea462-df90-4351-9f61-f16f054ce4e9 · outbound

This paper cites Gomez, Łukasz Kaiser, and Illia Polosukhin.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Gomez, Łukasz Kaiser, and Illia Polosukhin

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:29.839689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.663765Z digest=sha256:ba8f0debfd9c8b5b861d60027e3e401f675b92c948130765598337822e8e3f7d

Observation 38068dd4-107f-4687-83b4-8cf6ff6ba778 · outbound

This paper cites Meta Lingua: A minimal PyTorch LLM training library, 2024.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Meta Lingua: A minimal PyTorch LLM training library, 2024

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:29.827044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.668757Z digest=sha256:adbe11cff7874f6ac641cfc9a3d2bd73b8e7cf8e88108ddbace6ea4df7da4f9f

Observation 1f15343f-3aa1-4764-a255-45564d336b2c · outbound

This paper cites Closing the gap between the upper bound and lower bound of Adam’s iteration complexity.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Closing the gap between the upper bound and lower bound of Adam’s iteration complexity

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:29.812872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.672935Z digest=sha256:062dbd2d5f0ddebf4b6413d652e8abbeac8ca178d8fc14202ff2b8fd53cce2ec

Observation 86ddbe66-ccb9-4fec-8b2f-07de7a059f34 · outbound

This paper cites MicroLlama-300M.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism MicroLlama-300M

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:29.799649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.675883Z digest=sha256:c7ec89dfb645ba08038493708040de109a35aa6fd93dd02cd2253290fb75f7c8

Observation ed4e73fc-e733-4218-9e62-787e408a693e · outbound

This paper cites Liu, Lechao Xiao, Katie Everett, Alex Alemi, Ben Adlam, John D.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Liu, Lechao Xiao, Katie Everett, Alex Alemi, Ben Adlam, John D

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:29.663870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.679793Z digest=sha256:7e361ece3da1af7fc46a14043258a2807dfa0953234160cbfd80932224b74574

Observation b954a9c7-ac0e-4775-8057-031783abea46 · outbound

This paper cites Baichuan 2: Open Large-scale Language Models.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Baichuan 2: Open Large-scale Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.684510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.684510Z digest=sha256:06c2594a328b703e452d883348f85fdf2e0910716a92da4d77e0d4072c63c35c

Observation 52d360f0-e63b-4c57-abf8-3f126710905b · outbound

This paper cites Qwen2 Technical Report.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Qwen2 Technical Report

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.733863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.733863Z digest=sha256:c07792927d63d7a17e144e18d5340afecb39672ef85e3c06da37c0a80f845312

Observation 6cdb2a43-e351-4882-944a-b155be40774b · outbound

This paper cites Large batch optimization for deep learning: Training BERT in 76 minutes.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Large batch optimization for deep learning: Training BERT in 76 minutes

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:29.620830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.778076Z digest=sha256:2e800c5a25e9e7bc01f2e800884ecafe46dbe509957b91462d4351c252eb4b06

Observation 041b3c68-5256-4840-819c-0e3d2633bbbc · outbound

This paper cites Susskind.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Susskind

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:29.606769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.809521Z digest=sha256:0b1e19bf4064c3852e82d770f6079af769c85f940d244df4aa5df313a1222bec

Observation aca9d4ba-94cd-475e-9e9b-83fb0df939aa · outbound

This paper cites How does critical batch size scale in pre-training? InInternational Conference on Learning Representations (ICLR), 2025.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism How does critical batch size scale in pre-training? InInternational Conference on Learning Representations (ICLR), 2025

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:29.593274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.813683Z digest=sha256:97fc6441e1d6be9f74551566b37211f368c50e8f4238c6b3a37979b916c1cf4a

Observation c1b70b2a-0fa1-4414-9fd0-5e12aa5c7f3a · outbound

This paper cites TinyLlama: An Open-Source Small Language Model.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism TinyLlama: An Open-Source Small Language Model

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.819083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.819083Z digest=sha256:4c27500a674baba62e29a24c3129126ac4838f064edf5bb5fe8fb48db72eab6f

Observation 8a79759a-0bd9-41a0-986e-8732853ebe77 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism OPT: Open Pre-trained Transformer Language Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.823254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.823254Z digest=sha256:149a9a15b99a5f9e3d3fea8c8e63a7b03884d8e20747f203d35bcc3118137a0e

Observation 2a1a938c-4694-44ef-8f10-c7c54ec86094 · outbound

This paper cites Why Transformers Need Adam: A Hessian Perspective.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Why Transformers Need Adam: A Hessian Perspective

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.827620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.827620Z digest=sha256:84a812e87973beef01c3b1d875b9c69a5dd955e3b3bfd154c7d96f6d7aa2ce8c

Observation 26383e5e-bdd6-4643-b43f-11e84065244a · outbound

This paper cites Picotron: Distributed training framework for education and research experimentation, 2025.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Picotron: Distributed training framework for education and research experimentation, 2025

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:29.500772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.831666Z digest=sha256:1bd5bcdcde563285c2ac2fd152d99e3d70059cc0664ad24cfe05dac3f41e3710

Observation 3103d643-bf42-404a-9da8-2952792fec22 · outbound

This paper cites first-order term.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism first-order term

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:29.489027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:08:28.873206Z digest=sha256:c5cfb26dc69cbaa55f77b0af9201a4471bcafa0f23575f722e4b0e0770572d8e

Pith citing papers

Observation 3ef300f3-0353-4936-a83f-cac4e59f93c4 · inbound

Convergence of Riemannian Stochastic Gradient Descents: Varying Batch Sizes And Nonstandard Batch Forming cites this paper.

Convergence of Riemannian Stochastic Gradient Descents: Varying Batch Sizes And Nonstandard Batch Forming Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:10:53.387410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T18:38:58.883757Z digest=sha256:4d45dfd93c6fc2bc652eec986ac5a0bf8cf26d5be516769e2b1b160e65c38178