Pith. sign in

Paper Citation Record · LEDGER

Communication-Efficient Language Model Training Scales Reliably and Robustly: Scaling Laws for DiLoCo

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2503.09799.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.09799 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:13:20.674049Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 85f54d5b-9c0d-4c0a-915c-3e936ec6cb40 · inbound

DES-LOC: Desynced Low Communication Adaptive Optimizers for Training Foundation Models cites this paper.

DES-LOC: Desynced Low Communication Adaptive Optimizers for Training Foundation Models Communication-Efficient Language Model Training Scales Reliably and Robustly: Scaling Laws for DiLoCo

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:20.674049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:20.674049Z digest=sha256:3f70378e71113c2f90a5791379e9f10748819522be3a71c791990c38d9c9b5d1

Observation 95eb2ffb-6fa5-4b1f-94a8-092f89d79c01 · inbound

MuLoCo: Muon is a practical inner optimizer for DiLoCo cites this paper.

MuLoCo: Muon is a practical inner optimizer for DiLoCo Communication-Efficient Language Model Training Scales Reliably and Robustly: Scaling Laws for DiLoCo

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:26.394889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:26.394889Z digest=sha256:d87707224e47fe42a3e37c3a8f16b866839f3c59f7a30f20e95bb81ea779517a

Observation b44aafe4-29b1-4f1c-b8fb-76420366909f · inbound

NoLoCo: No-all-reduce Low Communication Training Method for Large Models cites this paper.

NoLoCo: No-all-reduce Low Communication Training Method for Large Models Communication-Efficient Language Model Training Scales Reliably and Robustly: Scaling Laws for DiLoCo

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:01.105792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:01.105792Z digest=sha256:a742135988c2b3ed3bc891d825227f1b9517e55f3294917c693119f89ac57dbb

Observation 40bc5464-e465-4d85-81fb-b670de0f5cfb · inbound

Predictable Scale: Part II, Farseer: A Refined Scaling Law in Large Language Models cites this paper.

Predictable Scale: Part II, Farseer: A Refined Scaling Law in Large Language Models Communication-Efficient Language Model Training Scales Reliably and Robustly: Scaling Laws for DiLoCo

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:29.282341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:29.282341Z digest=sha256:2ce8e6c82efafc4448c634d0012a0f275f3cdb5c6099d972445da01fb74a80b6

Observation 2b36b660-1f2c-4c85-9e40-cb9c4846fbc0 · inbound

Distributed and Decentralised Training: Technical Governance Challenges in a Shifting AI Landscape cites this paper.

Distributed and Decentralised Training: Technical Governance Challenges in a Shifting AI Landscape Communication-Efficient Language Model Training Scales Reliably and Robustly: Scaling Laws for DiLoCo

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:45.308895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:45.308895Z digest=sha256:63b61fa5ad1e6e12700561acb8714ee7a17a1e46a8bf939a37d33595d0e7ef5b

Observation 765a78bc-e74b-4f82-8c6b-cefcc58c91b1 · inbound

LoRDO: Distributed Low-Rank Optimization with Infrequent Communication cites this paper.

LoRDO: Distributed Low-Rank Optimization with Infrequent Communication Communication-Efficient Language Model Training Scales Reliably and Robustly: Scaling Laws for DiLoCo

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-03T04:44:55.124746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:44:55.124746Z digest=sha256:bc186af417531f9aea85e76000df089c7a7cf02d88117722259920da18c587f6

Observation fb106583-4e2a-4ae2-b1b4-e15b21b4b1ee · inbound

Outer-Momentum Restarting in High-Dimensional Two-Phase Optimization cites this paper.

Outer-Momentum Restarting in High-Dimensional Two-Phase Optimization Communication-Efficient Language Model Training Scales Reliably and Robustly: Scaling Laws for DiLoCo

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:33:27.503188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T13:31:23.664332Z digest=sha256:9fbcade3371eed259f8fdad71b1191c734ab213ce0767b51eb245ec046d9c995

Observation 9cba3c9d-1afb-4bc0-b6a3-a72792f25c6f · inbound

Does Distributed Training Undermine Compute Governance? cites this paper.

Does Distributed Training Undermine Compute Governance? Communication-Efficient Language Model Training Scales Reliably and Robustly: Scaling Laws for DiLoCo

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:06:03.317541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T00:54:00.828578Z digest=sha256:675865487b488a21454fcddf5beb320ef8393db73f2f02369f846b353dbbe628

Observation 9d4c4883-2ef4-42c6-94d2-3dc86aade6e1 · inbound

Echelon: Auditable Aggregate-Only Language-Model Adaptation Across Privacy Boundaries cites this paper.

Echelon: Auditable Aggregate-Only Language-Model Adaptation Across Privacy Boundaries Communication-Efficient Language Model Training Scales Reliably and Robustly: Scaling Laws for DiLoCo

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:06:23.691305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T13:38:48.333498Z digest=sha256:093f35a576d17b536267f9c7f1d2111f13c4939647aacc808f46485b138348da

Observation 6811ed37-4c4c-4afc-a87e-a9ada013e3fe · inbound

What's in a Smoothness Constant? Tighter Rates for Local SGD with Bounded Second-order Heterogeneity cites this paper.

What's in a Smoothness Constant? Tighter Rates for Local SGD with Bounded Second-order Heterogeneity Communication-Efficient Language Model Training Scales Reliably and Robustly: Scaling Laws for DiLoCo

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T01:23:35.023489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:23:35.023489Z digest=sha256:4a02aa5eceef45647dcc04d5036e5fc28e44819e9128b05f2e637a081c5ee9db