Pith. sign in

Paper Citation Record · LEDGER

How Does Critical Batch Size Scale in Pre-training?

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2410.21676.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.21676 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T22:05:02.683938Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T23:59:07.430461Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 57bdad84-9478-4da6-a432-c3db2ea83848 · inbound

Scaling Laws for Differentially Private Language Models cites this paper.

Scaling Laws for Differentially Private Language Models How Does Critical Batch Size Scale in Pre-training?

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-09T22:05:02.683938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:05:02.683938Z digest=sha256:1d2f2e0c7e08de55927cedb6f09b22ce3bce5e49415f7ecf39912e1ac44a6ef7

Observation 37a5bffc-f69f-4701-ab14-eb9cf9c9450d · inbound

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling cites this paper.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling How Does Critical Batch Size Scale in Pre-training?

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.630441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.630441Z digest=sha256:ba5a046878978aebbd4dcbca8594e3dbb5e638a8e9159fa1cde8ae4de4ac34fa

Observation 22cce5e8-b8cc-4de9-8795-51ff963b6b16 · inbound

Language Models Improve When Pretraining Data Matches Target Tasks cites this paper.

Language Models Improve When Pretraining Data Matches Target Tasks How Does Critical Batch Size Scale in Pre-training?

Reference 119

Resolution
unresolved
no resolver link, observed 2026-08-06T16:53:16.446210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:53:16.446210Z digest=sha256:3a4351dbda665df36c05595d22432c82f104654cd32cdac363efd8e3ec53e1ff

Observation 31cb124b-8863-4e1e-92e1-f0b93d1a78dd · inbound

EMA Without the Lag: Bias-Corrected Iterate Averaging Schemes cites this paper.

EMA Without the Lag: Bias-Corrected Iterate Averaging Schemes How Does Critical Batch Size Scale in Pre-training?

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T10:23:52.031002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:23:52.031002Z digest=sha256:9981586faa10ebaeea5e0b1a337c2d5abc637e68d00ceba35b2df66821b6613b

Observation 8cbfc8fd-1707-4130-9735-3cf18834e1c2 · inbound

Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling cites this paper.

Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling How Does Critical Batch Size Scale in Pre-training?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T09:38:47.135094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:38:47.135094Z digest=sha256:4be4d6e1066c35a3461d27d63f4b625c087d8150d00bb8eec3748abf566735e7

Observation 9db54e40-6023-4347-b63c-506c461349fc · inbound

Predicting Large Model Test Losses with a Noisy Quadratic System cites this paper.

Predicting Large Model Test Losses with a Noisy Quadratic System How Does Critical Batch Size Scale in Pre-training?

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:01:25.573125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T04:40:05.006583Z digest=sha256:aa47a5dfb66b4112ffaaf4cb4ff0b97f7a1e0d615015132c94b93ae4283532e7

Observation e4dfbb5e-9efa-42c3-9c40-a3717612f892 · inbound

ReCoVer: Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile Workload cites this paper.

ReCoVer: Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile Workload How Does Critical Batch Size Scale in Pre-training?

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:02:06.468306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T01:59:28.408860Z digest=sha256:8b8e3ca0b12a3149fc4fe329dd44599475e6c996f08286d2c62105631c3ba405

Observation 144ac853-5870-4d81-9e16-c5c21cdf1c26 · inbound

ReCoVer: Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile Workload cites this paper.

ReCoVer: Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile Workload How Does Critical Batch Size Scale in Pre-training?

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:55:24.625480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:52:30.402478Z digest=sha256:6a1e8c5f9db5fb5f817f2e2167a7f44da0e4223bafe2f9b55cce5463e35621cf

Observation dcfc156b-b9a1-4849-a977-f084f4a441a1 · inbound

ScheduleFree+: Scaling Learning-Rate-Free & Schedule-Free Learning to Large Language Models cites this paper.

ScheduleFree+: Scaling Learning-Rate-Free & Schedule-Free Learning to Large Language Models How Does Critical Batch Size Scale in Pre-training?

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:23:16.775669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T12:22:30.263086Z digest=sha256:66004a83621c82f734ffd58816cf3852df38c4e126ccd7397936ef281652748f

Observation 7e7692b0-c2a6-4ad9-9bb0-d8a88d208115 · inbound

Compute Efficiency and Serial Runtime Tradeoffs for Stochastic Momentum Methods cites this paper.

Compute Efficiency and Serial Runtime Tradeoffs for Stochastic Momentum Methods How Does Critical Batch Size Scale in Pre-training?

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:59:07.432881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T21:31:49.871520Z digest=sha256:a188c8fa5a4e2c24a29bba746d344be6416921df7047af2ab11484c71fb65ef9

Observation 0a6b880c-f5b9-4eb1-9285-6969038660ea · inbound

One-Step Gradient Delay is Not a Barrier for Large-Scale Asynchronous Pipeline Parallel LLM Pretraining cites this paper.

One-Step Gradient Delay is Not a Barrier for Large-Scale Asynchronous Pipeline Parallel LLM Pretraining How Does Critical Batch Size Scale in Pre-training?

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:44:18.850288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T06:41:55.732230Z digest=sha256:6bad5ba5f2145a0994c33d8c373e1b50329addc2a87d9fd054445c0028542c60

Observation 4b332e82-0af7-4de7-a849-44aa9359d05d · inbound

Bridging Compute- and Data-Optimal Pretraining cites this paper.

Bridging Compute- and Data-Optimal Pretraining How Does Critical Batch Size Scale in Pre-training?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T03:02:01.026199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:02:01.026199Z digest=sha256:44df8ced513936df9be0ec949493944a278c6746b763e17c5badec7505a1d485