Pith. sign in

Paper Citation Record · LEDGER

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism

As of 16 August 2026, this Paper Citation Record lists 88 of 88 outbound references and 1 inbound Pith citation observation for arXiv:2412.21124.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.21124 v2

Coverage vector

measured 88 of 88 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:08:28.873206Z

measured 89 of 89 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T18:38:58.883757Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T00:10:53.383914Z

Reference resolution

88 of 88 outbound references displayed

  • verified exact2
  • verified fuzzy39
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b6810cf5-1d1c-4834-9b3b-ea411c930456 · outbound

This paper cites an unresolved cited work.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:26.118184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:26.118184Z digest=sha256:09201d05f7867b027324850a24ec8333a32e2bce22a4fb0cd64b46636a5b28f1

Observation b4d1e638-3900-4a90-86bc-ff63b3adc3cc · outbound

This paper cites Extremely Large Minibatch SGD: Training ResNet-50 on ImageNet in 15 Minutes.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Extremely Large Minibatch SGD: Training ResNet-50 on ImageNet in 15 Minutes

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:26.212271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:26.212271Z digest=sha256:1ea7f0f87bd8ac41a4efeb6db2941c74b7f506bd25a5f89a75c705aa4a209afc

Observation 84a20f5c-bdf2-4da6-9d0d-6056e3d8527e · outbound

This paper cites Qwen Technical Report.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:26.262650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:26.262650Z digest=sha256:2831f1173eb73387c9de37c4a7826123742826e5e048f0aafd85bc0252c866bc

Observation e363a1bf-ad3a-4431-a9ee-ebf276a9553a · outbound

This paper cites Coupling adaptive batch sizes with learning rates.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Coupling adaptive batch sizes with learning rates

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:26.267815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:26.267815Z digest=sha256:ac25c89b3de1d19b10d3248d0391df98580e1551eafd8880eaa99ccc30e21533

Observation e7be17e8-052a-492e-8dfd-c091eb85b34f · outbound

This paper cites Stable LM 2 1.6B Technical Report.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Stable LM 2 1.6B Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:26.271570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:26.271570Z digest=sha256:5039f9077ea2d085e56724c3cf7bd343f3a2b78a94fc6e4030255f377b935595

Observation de41598d-de41-4249-9e38-d6153b205e8a · outbound

This paper cites BLOOM: A 176B-Parameter Open-Access Multilingual Language Model.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:26.276815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:26.276815Z digest=sha256:d8da463b4c2d6f4009f163612247b1446ce24085705fcfd80676d184560072ff

Observation 48f87ad8-556f-456e-8af3-b79a5c0de339 · outbound

This paper cites an unresolved cited work.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:26.312733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:26.312733Z digest=sha256:a0dd1753d66bcce344f3ec6fc48bdb80014e29aa039f5f3d70ae46c2557fb8d5

Observation 0154dee3-e45b-4bc6-9507-a01fcdd8e549 · outbound

This paper cites Adaptive sampling strategies for stochastic optimization.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Adaptive sampling strategies for stochastic optimization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:26.440847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:26.440847Z digest=sha256:dc58758859f6155340fad49a5ce792fe492644d539f6e5a9479e1c0737863a15

Observation e54d7e66-2702-49f4-a396-a9a7e2f8d840 · outbound

This paper cites Curtis, and Jorge Nocedal.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Curtis, and Jorge Nocedal

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:26.486876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:26.486876Z digest=sha256:dd7cde4979a3ff52e410072da130b21a20700f60e5b1ff608202ca4fb5defaa2

Observation 23065d4c-fbea-43bd-b372-852a508c27d4 · outbound

This paper cites JAX: composable transformations of Python+NumPy programs, 2018.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism JAX: composable transformations of Python+NumPy programs, 2018

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:26.491193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:26.491193Z digest=sha256:7e58b999b28567ca346bc740aeaa0c7644085a4bcbfc8101b1277e9d5add1165

Observation 0caed567-023f-4a1d-8a3b-1d2683f1dd07 · outbound

This paper cites an unresolved cited work.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:26.495360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:26.495360Z digest=sha256:38ddba2636e27d22b02324cdf53d4d5b47088321c7b9b84a1653f1488fdb8739

Observation d0eeab79-a6bc-47b0-b2fc-8445b74c20f9 · outbound

This paper cites Byrd, Gillian M.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Byrd, Gillian M

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:26.501738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:26.501738Z digest=sha256:dab23fba80a7904890b50fabcc4f2b2ec703394d0e0324ed96b4e1f057e9f221

Observation c6be660f-e3e5-4157-b5c8-1d6ac47cf887 · outbound

This paper cites Big Batch SGD: Automated Inference using Adaptive Batch Sizes.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Big Batch SGD: Automated Inference using Adaptive Batch Sizes

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:26.505657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:26.505657Z digest=sha256:4815211d4adddf7ece1f59c13e1c70759bf2714e4a029ad5a788624b9d50be3c

Observation 35e837c6-2948-45ac-95f5-cd22a7fe7355 · outbound

This paper cites Automated inference with adaptive batches.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Automated inference with adaptive batches

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:31.508603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:26.510052Z digest=sha256:9054f454f4311de2e96b345ec470677734845663fd3dfdfd7ce80c1843fd82bb

Observation 4b30ab26-4207-4482-920c-93f59be5f71e · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:26.514375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:26.514375Z digest=sha256:2f4d80ff4ad2a4d113bfbafbd12a39a791dc7d972ba42d95fe7d132a103ed36c

Observation 15bc84be-5671-4d28-8a77-668b9b38bd04 · outbound

This paper cites DeepSeek-V3 Technical Report.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism DeepSeek-V3 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:26.574840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:26.574840Z digest=sha256:14bb97bbd84aad99835243ed9b3324e0fb31e9313cf5c254ec4b66f1f20d2696

Observation 987fb9ad-f1b0-40b2-8a3d-0e4f7f93e3ab · outbound

This paper cites AdaBatch: Adaptive Batch Sizes for Training Deep Neural Networks.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism AdaBatch: Adaptive Batch Sizes for Training Deep Neural Networks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:26.676251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:26.676251Z digest=sha256:68027d0bc8935a4da5f58a72e254b7d6a102e5124ee70a7b42bba8eb07262825

Observation 13543c64-9340-4ab2-b218-39a27f7781cf · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism An image is worth 16x16 words: Transformers for image recognition at scale

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:27.007580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:27.007580Z digest=sha256:3c9117e58bc5fd041fd59cba0142663d05c350dd244de99a00db71163bb712af

Observation b961533f-5b32-49cb-a584-d36ecb1e2447 · outbound

This paper cites Susskind, and Armand Joulin.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Susskind, and Armand Joulin

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:31.486551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:27.787032Z digest=sha256:8c58b4df98d4e85e73d54a5afdb7bf79f4a7d12fb835065af95d71c1cf2b88e2

Observation 29e69ced-3080-4adc-8562-8aea5b81ec8c · outbound

This paper cites PyTorch Lightning, 2019.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism PyTorch Lightning, 2019

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:31.416593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:27.838868Z digest=sha256:543a11e3b45431f8f7b5cb44158c0afd28e9651fc497e9625dd353bdd6fcac54

Observation 38fb4cec-205f-4ba5-b869-2103668853ad · outbound

This paper cites Friedlander and Mark Schmidt.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Friedlander and Mark Schmidt

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:31.402273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:27.843948Z digest=sha256:d5ef909db59603ef1bd3d542278898b11a4ead49508f19d02e97658031352de2

Observation 13baffce-448c-4f9c-9c6c-e347e157cb1a · outbound

This paper cites Compiling machine learning programs via high-level tracing.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Compiling machine learning programs via high-level tracing

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:31.303387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:27.849740Z digest=sha256:0ef75254ede5e1287e12820afb3a7b4d8b046cf4bb5a5cfbf14a08bd2d0d13bc

Observation cf72c474-a5cb-42c7-9018-3fb9dcfd2392 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Gemma 2: Improving Open Language Models at a Practical Size

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:27.854625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:27.854625Z digest=sha256:7fdf09d5c9c188a239dd5799a2675741e4411a02a5fab1cf309d79d3326a23a7

Observation 9ac85a57-6330-4013-bb12-7df6911a50a6 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Gemma: Open Models Based on Gemini Research and Technology

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:27.860027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:27.860027Z digest=sha256:be929333d7b8596c82c9ca84b06203b78bd37e35e9bf80a4a91ee760d201431b

Observation a3ed2a93-05a1-4f3d-aa6c-667d7e12fdf4 · outbound

This paper cites OpenLLaMA: An open reproduction of LLaMA, May 2023.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism OpenLLaMA: An open reproduction of LLaMA, May 2023

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:31.288817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:27.863843Z digest=sha256:0b57ecaff013ad715ae88cf4a2a969bd357c124d4b80eb8e17ce7645f00657f5

Observation c6ef5449-5106-4836-bb71-1b19318437a1 · outbound

This paper cites Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:27.940336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:27.940336Z digest=sha256:c674ecececfb2abce779b73296a3d57a30873809b0a040a5321279680699a079

Observation ac423a62-c31f-41a4-906f-d01ac3142681 · outbound

This paper cites OLMo: Accelerating the Science of Language Models.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism OLMo: Accelerating the Science of Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:27.945364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:27.945364Z digest=sha256:c475115492d61bd2a36df8aaeb0ca1aa9d3cb5f1706f80e9d755f81849781698

Observation b5c3fcb0-70c9-4acf-9a10-3e43e0d12e93 · outbound

This paper cites Scaling laws for single-agent reinforcement learning.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Scaling laws for single-agent reinforcement learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:27.949748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:27.949748Z digest=sha256:ef7448464e018c630091646d88e458bdd36a257d5fbde8f700da92d8ead6e5f9

Observation 51935083-cbe8-4bab-932a-60ee23cdffa6 · outbound

This paper cites Train longer, generalize better: closing the generalization gap in large batch training of neural networks.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Train longer, generalize better: closing the generalization gap in large batch training of neural networks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:31.275809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:27.954773Z digest=sha256:1ce011cc419eae080676dd1a39fc28540e959fb802824a31066da878d22725d2

Observation 78be173c-05ad-4d19-b32c-bd48f41bc3c7 · outbound

This paper cites Mistral 7B.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Mistral 7B

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:27.958688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:27.958688Z digest=sha256:d4eb267ee7123cce00fc45e2c6acbba064d87c9ddecd96f7176cb2ea5c3502ed

Observation 17c167d4-30b6-4bdd-819a-27934a167b44 · outbound

This paper cites GrowLength: Accelerating LLMs Pretraining by Progressively Growing Training Length.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism GrowLength: Accelerating LLMs Pretraining by Progressively Growing Training Length

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:08:29.256142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:27.964430Z digest=sha256:c42776195c4ca0fb249381869bdccf19dbbf32d0e43a6f54251e1395444d9f40

Observation 1760b5be-98d9-4526-b657-3081a272bab5 · outbound

This paper cites AdaScale SGD: A user-friendly algorithm for distributed training.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism AdaScale SGD: A user-friendly algorithm for distributed training

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:31.172939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.018228Z digest=sha256:862854baa82cb001aea51b93e5aff375e4fdefbdf46bcbb614b072197e86fab5

Observation f48c3594-b81e-4b68-8d15-4c9f33f392e6 · outbound

This paper cites Scaling Laws for Neural Language Models.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Scaling Laws for Neural Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.022357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.022357Z digest=sha256:d4e005c5ba5bc7ab2d5755596800416297a61b9c4608b7f967af3920f1685607

Observation 17647574-453a-49ec-af1d-8bfb0abe4257 · outbound

This paper cites On large-batch training for deep learning: Generalization gap and sharp minima.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism On large-batch training for deep learning: Generalization gap and sharp minima

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:31.146639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.026043Z digest=sha256:d4d6053d3a78960c8ae030c6d41a09071e898feb7995e27b709184edfc8fb018

Observation a94984a0-e9fb-43e4-b2d8-10affe2a0b38 · outbound

This paper cites Better theory for SGD in the nonconvex world.Transactions on Machine Learning Research, 2023.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Better theory for SGD in the nonconvex world.Transactions on Machine Learning Research, 2023

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.030194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.030194Z digest=sha256:ef940a05d241093c2c277efc32b7f576c24173760b88e67f45b14e05a6683867

Observation a5e987de-2f6f-4743-89b1-93ad0cfb3414 · outbound

This paper cites Kingma and Jimmy Lei Ba.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Kingma and Jimmy Lei Ba

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:31.060307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.033499Z digest=sha256:35165ffb743386b64ce84e18d6e4aaf9a9efc4524e3cca1b482ecc2f571d6ece

Observation 347cc37f-44ae-4a79-b544-a41c08c30de7 · outbound

This paper cites Reducing activation recomputation in large transformer models.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Reducing activation recomputation in large transformer models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:31.045296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.037092Z digest=sha256:d4fc7eeed161fcb8e6c7ea9ea340ff9b250bdd481bd1303ab0f57d6aa6aae6a4

Observation 6c6e722c-e118-4edf-b158-1dec81a7baf7 · outbound

This paper cites an unresolved cited work.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:08:30.949976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.041594Z digest=sha256:2102c92b5708c1a571658a10eeead3f1df8b5e945b231dc0c6774a1be4c865b5

Observation c3a741b7-853d-4ea0-a2b0-fe8418af1843 · outbound

This paper cites Noise is not the main factor behind the gap between SGD and Adam on transformers, but sign descent might be.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Noise is not the main factor behind the gap between SGD and Adam on transformers, but sign descent might be

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:30.936710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.086417Z digest=sha256:1cccfc1e3a9570aad0193bc22faa05cf4966395166ad2eccf30210bad28d5d8e

Observation 2f3cd7b1-45a6-4f41-be3f-4401f52edc0a · outbound

This paper cites Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.116499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.116499Z digest=sha256:4adf24a8cb36e3754b7765eb4f8fab499c156d6684ff79c148bafaaa5fd44a3e

Observation 6560157b-eab2-4e15-89df-67a2ec763523 · outbound

This paper cites Communication-Efficient Adaptive Batch Size Strategies for Distributed Local Gradient Methods.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Communication-Efficient Adaptive Batch Size Strategies for Distributed Local Gradient Methods

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.119718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.119718Z digest=sha256:49686676fb1975e7c574c263ec2b70ba94a486276a88aff75ed7908cd931f686

Observation 2928c122-0501-4787-9a67-f490766b7b7a · outbound

This paper cites AdAdaGrad: Adaptive Batch Size Schemes for Adaptive Gradient Methods.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism AdAdaGrad: Adaptive Batch Size Schemes for Adaptive Gradient Methods

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.123888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.123888Z digest=sha256:07b558e237bc4c3af49346e1fcdb6f9ae71e340340b89d0c1ebdc5c162ea7ec0

Observation 0085f04e-753d-40ad-9c68-87cb0bf298a9 · outbound

This paper cites Orr, and Klaus Robert Müller.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Orr, and Klaus Robert Müller

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:30.901128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.128430Z digest=sha256:c179ecabfa64f5695cc30a7f0d4b0dd1e997b6eeaff9651a1a41bcbb6bcb22bd

Observation 38d7e523-8b65-408a-bd78-b257912d9c85 · outbound

This paper cites GShard: Scaling giant models with conditional computation and automatic sharding.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism GShard: Scaling giant models with conditional computation and automatic sharding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.132217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.132217Z digest=sha256:f737ff24339b88113f56f821d39fc241785347102c0a5076a927008880bd08b8

Observation e8fd0fe8-57ef-4aa5-92c7-0ac612b66797 · outbound

This paper cites The stability-efficiency dilemma: Investigating sequence length warmup for training GPT models.Advances in Neural Information Processing Systems (NeurIPS), 2022.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism The stability-efficiency dilemma: Investigating sequence length warmup for training GPT models.Advances in Neural Information Processing Systems (NeurIPS), 2022

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:30.788644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.135256Z digest=sha256:60412d9896ecdbf00b31b3d13757e4ee4440b131849ac769bc9f5660f94eaa7b

Observation fd20966e-19e9-431f-9293-2f6afc57c197 · outbound

This paper cites PyTorch distributed: Experiences on accelerating data parallel training.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism PyTorch distributed: Experiences on accelerating data parallel training

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:30.776699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.139259Z digest=sha256:dc6119f300fdd3c023d53fdb74b211c754fb6d53684d10712aeda1cd2258b4b8

Observation 80ff2617-719a-41af-9d88-a06457a4f1c6 · outbound

This paper cites TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:30.762450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.142617Z digest=sha256:c95973aa46be7cc80c765a8c2cfddd54f8234b5b3ff16d9ddb7cb73725626b05

Observation 58b284de-9ed3-4d16-8f14-8f780d34b680 · outbound

This paper cites LitGPT, 2023.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism LitGPT, 2023

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:30.685452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.191665Z digest=sha256:b50d5a4223dd457cdfba14bca27cdb90ab6dc3731155e62f3181b2b54e8e0749

Observation 9a41a052-d67d-4158-85f3-5931beda7774 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.197709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.197709Z digest=sha256:f82c0efa762ff1e6918e4f8464818998c16d521a97df30e704192c0828c0af52

Observation c644634f-0da3-4edd-82d2-d513fe3408c7 · outbound

This paper cites The Llama 3 Herd of Models.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism The Llama 3 Herd of Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.203669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.203669Z digest=sha256:a3f58008f73f495376d7270c012523565bb293a6626e41341a6ec2f2c7369ca9

Observation 34f011ae-18a7-4607-a505-725a14048efd · outbound

This paper cites Decoupled weight decay regularization.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Decoupled weight decay regularization

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.208234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.208234Z digest=sha256:2c23a59895d93a4e20def89d5a86ba97cc0ece7d1400e6a4d6377c4262d9f813

Observation d34ffa08-e6a3-40e1-b3e0-5cd1ba38d35a · outbound

This paper cites An Empirical Model of Large-Batch Training.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism An Empirical Model of Large-Batch Training

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.211454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.211454Z digest=sha256:e1c8e00f743e1bf8d47872975c55efc29cb7752ce01401da6b3369203b0cda5b

Observation 074aab96-b4b6-4f21-aac4-0bb83b2515f9 · outbound

This paper cites Efficient large-scale language model training on GPU clusters using Megatron-LM.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Efficient large-scale language model training on GPU clusters using Megatron-LM

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:30.563989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.215521Z digest=sha256:8145ac9c0b5c3ce2ebdb30e625b7fa7683901c8decc26f0b5f5547aa749c9c8f

Observation 22200e34-27dc-4aa2-874f-5bee70eadc5f · outbound

This paper cites AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:08:29.105925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.219044Z digest=sha256:642c061cb8279606f11802aa8998407329e04e3f9f4cc1491b3eb6b553e585f9

Observation de7bce79-ad55-4a69-8474-051b04a9546c · outbound

This paper cites Toward Understanding Why Adam Converges Faster Than SGD for Transformers.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Toward Understanding Why Adam Converges Faster Than SGD for Transformers

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.222610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.222610Z digest=sha256:55de91f5a580c7a62cd2535904467732221ece7f4a418b5bdf5f4b247fbb6b93

Observation a7f506ca-a3ad-4760-862d-0eda79403039 · outbound

This paper cites Nemotron-4 15B Technical Report.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Nemotron-4 15B Technical Report

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.313606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.313606Z digest=sha256:7d363f2834de71b21bab2d54c4ba9c55d2d742301954a2f21a0f756bd9c8ec77

Observation 8a8dd1b2-4148-4ef9-b20f-d797a19836dc · outbound

This paper cites PyTorch: An imperative style, high-performance deep learning library.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism PyTorch: An imperative style, high-performance deep learning library

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:30.479112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.354771Z digest=sha256:784600fe4448f48db7cd0db74d9f9896283d514fd1a2c7046165a1cd99f9ccfa

Observation 4137c201-4efe-4573-8233-8ca7b0e2ad11 · outbound

This paper cites Large scale language modeling: Converging on 40GB of text in four hours.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Large scale language modeling: Converging on 40GB of text in four hours

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:30.441402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.359098Z digest=sha256:ac9e13b6c07d737f1d1fc733c7e0111454068f1cf9fd80373b49fb544243a746

Observation bf88acb2-dc51-432b-a251-2a59cc317473 · outbound

This paper cites SimiGrad: Fine-grained adaptive batching for large scale training using gradient similarity measurement.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism SimiGrad: Fine-grained adaptive batching for large scale training using gradient similarity measurement

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:30.356127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.362697Z digest=sha256:6b466b0a6bf49a4f764d6afdea08b0dfafa736f809f778cb5c91b6ec7cb309d3

Observation 9dc0ab40-ecdb-4c15-adbb-fb68b7f9e0b9 · outbound

This paper cites an unresolved cited work.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.367120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.367120Z digest=sha256:ea989754ecb9c7d095fba69d201918e054f102bdd48697f5e3f5ce7c223385a6

Observation 6f68f0b6-4b00-4022-8e2e-3133f6e05054 · outbound

This paper cites ZeRO: Memory optimizations toward training trillion parameter models.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism ZeRO: Memory optimizations toward training trillion parameter models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:30.334289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.370025Z digest=sha256:7f622fb4b47e7716146cc86d316139e124d3cf76c4c80c1902b75c832fc52313

Observation 435361fa-002d-4e5c-8082-e025b9ff84c9 · outbound

This paper cites DeepSpeed: System optimizations enable training deep learning models with over 100 billion parameters.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism DeepSpeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:30.275184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.373751Z digest=sha256:40bccd48941bf1d9397099cec3a54a8b3b69ee16a34a4df0f492f799f1658387

Observation 413252af-5483-45ae-801c-a90863ecd40b · outbound

This paper cites ZeRO-Offload: Democratizing billion-scale model training.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism ZeRO-Offload: Democratizing billion-scale model training

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:30.232150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.378205Z digest=sha256:6e89f9100bb2f5546c3b9bd191ee3996a2f396b4a32e08c8e67bb67f8b3461ad

Observation 284f4916-6eaa-4116-a4f4-5957e00e21ae · outbound

This paper cites On the different regimes of stochastic gradient descent.Proceedings of the National Academy of Sciences, 121(9):e2316301121, 2024.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism On the different regimes of stochastic gradient descent.Proceedings of the National Academy of Sciences, 121(9):e2316301121, 2024

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:30.219609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.382219Z digest=sha256:8ade9a828b5c3def8f49d0882cbf6ed134c981762deba6772a6c589ff2f9427e

Observation b5b9e4a3-0569-436b-837d-6e4b093c86d9 · outbound

This paper cites Shallue, Jaehoon Lee, Joseph Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Shallue, Jaehoon Lee, Joseph Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:30.125126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.401956Z digest=sha256:c36660dcae318a87a918a9d916915725ea5c90f9dc03c93b691890bab76fd6ab

Observation 2c58e921-f588-416c-9b0f-8721c32559fe · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.459144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.459144Z digest=sha256:98f1a6f0fe719dcefc3737aa974d266e2ac8b2250807ddc76f4dfb3f75a4e8b5

Observation 63a15759-513c-44eb-a4a2-e140610e06c2 · outbound

This paper cites Smith and Quoc V.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Smith and Quoc V

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:30.031318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.513312Z digest=sha256:6d9a25191e74f904c07bdd36b2cd2cdb8e8eac467ee82c5c21e357a06cf94b23

Observation 234a99ec-997c-4c5c-81e8-00d68670c9f9 · outbound

This paper cites Smith, Pieter-Jan Kindermans, and Quoc V.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Smith, Pieter-Jan Kindermans, and Quoc V

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:29.994206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.518013Z digest=sha256:27fd5b667c04e42cc39236934fbacb6ec77ff9e7287b47ec84450fdcce8997e1

Observation 366ce334-1037-486c-a24f-3c327afabc8c · outbound

This paper cites Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.522233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.522233Z digest=sha256:397781c0e1c8f312c1299b9931fbe0f7b38d19aba58515c63da3c9daac1b0914

Observation 6140cea0-75ee-42cf-ad62-feac0f1759e2 · outbound

This paper cites Unraveling the Mystery of Scaling Laws: Part I.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Unraveling the Mystery of Scaling Laws: Part I

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.527070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.527070Z digest=sha256:9b67740a67c876839cb6893f7ab3ca585e6538d08b1a3698e67f866b957dc097

Observation 40cd2d9e-5b48-473f-8c59-f296d3c16d4b · outbound

This paper cites 2 OLMo 2 Furious.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism 2 OLMo 2 Furious

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.532201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.532201Z digest=sha256:cf87ef39c637e607c94eae243bc6d73368e962cd07860ad19172c9e712fd45ab

Observation 3a15bb9d-3ca6-4f94-9491-526cf29f2130 · outbound

This paper cites Introducing DBRX: A new state-of-the-art open LLM.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Introducing DBRX: A new state-of-the-art open LLM

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:29.973315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.536308Z digest=sha256:ae023a26aa743d34be94b4ea0bb2024bae752da06e4748ddfa2dc8f47e70e12d

Observation 68c1c185-1b47-43b3-a13b-42f30562146f · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.626060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.626060Z digest=sha256:a34b5dd397772a62453736854c9d6625963f2051dde5222b3a6a161c90d35bf3

Observation ff6ea462-df90-4351-9f61-f16f054ce4e9 · outbound

This paper cites Gomez, Łukasz Kaiser, and Illia Polosukhin.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Gomez, Łukasz Kaiser, and Illia Polosukhin

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:29.839689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.663765Z digest=sha256:c874d5d0dc11c9e45bf182868cef3f97605065ac5aa6a4ed400507a311012933

Observation 38068dd4-107f-4687-83b4-8cf6ff6ba778 · outbound

This paper cites Meta Lingua: A minimal PyTorch LLM training library, 2024.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Meta Lingua: A minimal PyTorch LLM training library, 2024

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:29.827044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.668757Z digest=sha256:0d3c76b8687d14d768cc556b92934f15c58f929b41ae30d7dd0138554de7b911

Observation 1f15343f-3aa1-4764-a255-45564d336b2c · outbound

This paper cites Closing the gap between the upper bound and lower bound of Adam’s iteration complexity.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Closing the gap between the upper bound and lower bound of Adam’s iteration complexity

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:29.812872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.672935Z digest=sha256:be3c1283a014ebe38eae2ea98884aa1641ebfbb89fccdfd713ef1d14515d5fa2

Observation 86ddbe66-ccb9-4fec-8b2f-07de7a059f34 · outbound

This paper cites MicroLlama-300M.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism MicroLlama-300M

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:29.799649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.675883Z digest=sha256:5c9a1c9e6ad0eae4a7b35dcb56873e40b1a14a30a8a98c5e58e41330bd46cc99

Observation ed4e73fc-e733-4218-9e62-787e408a693e · outbound

This paper cites Liu, Lechao Xiao, Katie Everett, Alex Alemi, Ben Adlam, John D.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Liu, Lechao Xiao, Katie Everett, Alex Alemi, Ben Adlam, John D

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:29.663870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.679793Z digest=sha256:3a2d595b2b89365b8ac4add1d130efc8b9be7bf8d0451a582c06c651a4aaaed8

Observation b954a9c7-ac0e-4775-8057-031783abea46 · outbound

This paper cites Baichuan 2: Open Large-scale Language Models.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Baichuan 2: Open Large-scale Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.684510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.684510Z digest=sha256:06c2594a328b703e452d883348f85fdf2e0910716a92da4d77e0d4072c63c35c

Observation 52d360f0-e63b-4c57-abf8-3f126710905b · outbound

This paper cites Qwen2 Technical Report.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Qwen2 Technical Report

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.733863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.733863Z digest=sha256:c07792927d63d7a17e144e18d5340afecb39672ef85e3c06da37c0a80f845312

Observation 6cdb2a43-e351-4882-944a-b155be40774b · outbound

This paper cites Large batch optimization for deep learning: Training BERT in 76 minutes.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Large batch optimization for deep learning: Training BERT in 76 minutes

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:29.620830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.778076Z digest=sha256:c85e161b50934b2d580d5dbef5c560d2bc009a84ae1e874a33bcdc2f81a4087b

Observation 041b3c68-5256-4840-819c-0e3d2633bbbc · outbound

This paper cites Susskind.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Susskind

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:29.606769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.809521Z digest=sha256:f15e64ed56c8cbe2dd2bb7a87a97a5b5e131e25578b9b33b700b46fb0d1e1572

Observation aca9d4ba-94cd-475e-9e9b-83fb0df939aa · outbound

This paper cites How does critical batch size scale in pre-training? InInternational Conference on Learning Representations (ICLR), 2025.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism How does critical batch size scale in pre-training? InInternational Conference on Learning Representations (ICLR), 2025

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:29.593274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.813683Z digest=sha256:365150e219690b4541eec4286bdd9cfd9c6eb62888fd9bcacedaafa493580ecd

Observation c1b70b2a-0fa1-4414-9fd0-5e12aa5c7f3a · outbound

This paper cites TinyLlama: An Open-Source Small Language Model.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism TinyLlama: An Open-Source Small Language Model

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.819083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.819083Z digest=sha256:4c27500a674baba62e29a24c3129126ac4838f064edf5bb5fe8fb48db72eab6f

Observation 8a79759a-0bd9-41a0-986e-8732853ebe77 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism OPT: Open Pre-trained Transformer Language Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.823254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.823254Z digest=sha256:149a9a15b99a5f9e3d3fea8c8e63a7b03884d8e20747f203d35bcc3118137a0e

Observation 2a1a938c-4694-44ef-8f10-c7c54ec86094 · outbound

This paper cites Why Transformers Need Adam: A Hessian Perspective.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Why Transformers Need Adam: A Hessian Perspective

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.827620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.827620Z digest=sha256:84a812e87973beef01c3b1d875b9c69a5dd955e3b3bfd154c7d96f6d7aa2ce8c

Observation 26383e5e-bdd6-4643-b43f-11e84065244a · outbound

This paper cites Picotron: Distributed training framework for education and research experimentation, 2025.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Picotron: Distributed training framework for education and research experimentation, 2025

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:29.500772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.831666Z digest=sha256:5512dc814abd9a87a85f188db5bf2d81c5d4bffd9d6615d3ed2b7f18ce9ab3fe

Observation 3103d643-bf42-404a-9da8-2952792fec22 · outbound

This paper cites first-order term.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism first-order term

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:08:29.489027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:08:28.873206Z digest=sha256:d3633b51dce6baa54a63de164a84b1c9cff64b58c9bc033ed69ebdf936a84946

Pith citing papers

Observation 3ef300f3-0353-4936-a83f-cac4e59f93c4 · inbound

Convergence of Riemannian Stochastic Gradient Descents: Varying Batch Sizes And Nonstandard Batch Forming cites this paper.

Convergence of Riemannian Stochastic Gradient Descents: Varying Batch Sizes And Nonstandard Batch Forming Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:10:53.387410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T18:38:58.883757Z digest=sha256:a14d666f3dfe98ec5a6a3a3a081389585f83949b7e23cae1bc280197ab03de19