Pith. sign in

Paper Citation Record · LEDGER

Distributed Sign Momentum with Local Steps for Training Transformers

As of 15 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2411.17866.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17866 v2

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:52:40.239399Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6d0e721a-bb86-4806-ac83-b4fdceb0a6ad · outbound

This paper cites B.2 Proof of Theorem 2 Proof.

Distributed Sign Momentum with Local Steps for Training Transformers B.2 Proof of Theorem 2 Proof

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:40.375344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T11:52:40.232476Z digest=sha256:d67cba41d7fbda82b3685fc1d855157a2bd83451525f62db144b2879fe3ea277

Observation 61450cc5-0962-48bd-aedd-0fe47d821e58 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Distributed Sign Momentum with Local Steps for Training Transformers Training Compute-Optimal Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T11:52:40.189105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:52:40.189105Z digest=sha256:1a6ad7c9ebd6470dec498cf74e1399feea5dd95993aed771a0d59b5b7c6d02b8

Observation 7ae358af-52de-437b-9dc2-89549154bf32 · outbound

This paper cites Communication Efficient Distributed Training with Distributed Lion.

Distributed Sign Momentum with Local Steps for Training Transformers Communication Efficient Distributed Training with Distributed Lion

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T11:52:40.204858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:52:40.204858Z digest=sha256:4626d427cbd9b45775c3369992ce651430274e2ac78f1f87f2f90f782bf0a310

Observation 24c53192-5b69-43ee-8f93-0e148fb7d541 · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Distributed Sign Momentum with Local Steps for Training Transformers SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T11:52:40.208306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:52:40.208306Z digest=sha256:d9ca01334d9625846454eb2b7e87b063cce6103e4c789298f62eaab8d2e57163

Observation 2dc2c6f2-8bac-4c4c-bd29-18cd79a5b4c9 · outbound

This paper cites SlowMo: Improving Communication-Efficient Distributed SGD with Slow Momentum.

Distributed Sign Momentum with Local Steps for Training Transformers SlowMo: Improving Communication-Efficient Distributed SGD with Slow Momentum

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T11:52:40.224459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:52:40.224459Z digest=sha256:3e4cfd37f4840e1f520eededfb23ea1673c22849cc1103f3d03b4a6248a2b202

Observation d93a4334-a18f-4903-913c-b78584af68dd · outbound

This paper cites Component-wise vector multi- plication g2 t = gt ⊙gt, and √bvt means component- wise square root.

Distributed Sign Momentum with Local Steps for Training Transformers Component-wise vector multi- plication g2 t = gt ⊙gt, and √bvt means component- wise square root

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:40.387240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T11:52:40.228224Z digest=sha256:c48c7243ca742fab29d99e1a5a7471ed1baded1a1019314d4a189698728944f8

Observation 68c5216b-0cb4-41cd-865c-7b632db19664 · outbound

This paper cites (38) Proof.

Distributed Sign Momentum with Local Steps for Training Transformers (38) Proof

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:40.364753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T11:52:40.236107Z digest=sha256:1a2e7ebafd99ae2f2bf5bb22214094218376983c4ae3dd2249f23d30be1396ee

Observation c5c0e246-f380-40ce-a537-ffb182308105 · outbound

This paper cites an unresolved cited work.

Distributed Sign Momentum with Local Steps for Training Transformers Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:52:40.353998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T11:52:40.239399Z digest=sha256:e9b6dd98f8aed9ff1370bbc183bcfc81f6a1363d200bbe6637a3fb23ab78c176

Observation 4eb2d010-1737-4871-ab99-21df792f0506 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Distributed Sign Momentum with Local Steps for Training Transformers An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-12T11:52:40.185201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:52:40.185201Z digest=sha256:c185917f55763139e57db859e85194f2d9e491fdf1bd0355a881adea10bf4e0f

Observation ab8ea634-c246-4fbe-8ea6-12cf0246c586 · outbound

This paper cites Fedlion: Faster adaptive federated optimization with fewer communication.

Distributed Sign Momentum with Local Steps for Training Transformers Fedlion: Faster adaptive federated optimization with fewer communication

Reference 2013

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:40.398010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T11:52:40.220896Z digest=sha256:ec9cd90f37bf3c704b7a6ac2f5d955728fdbb1a86166446f59059ef9e06d6d1d

Observation 45969739-6588-4308-8448-b4d3af42cff9 · outbound

This paper cites Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be.

Distributed Sign Momentum with Local Steps for Training Transformers Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-12T11:52:40.200833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:52:40.200833Z digest=sha256:d5cc2b112f3a69ec1b8407f118fdcb92a4dd22160fdcc83cbdf34f9697a5fa46

Observation 9c3be546-6f44-4312-8a1a-ebd0bc63e72a · outbound

This paper cites How to scale distributed deep learning?.

Distributed Sign Momentum with Local Steps for Training Transformers How to scale distributed deep learning?

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-12T11:52:40.193192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:52:40.193192Z digest=sha256:f1d55845eb4fb5c6701b281695eb262176b2ee8b58fdfabb49aab66dc4605000

Observation 5a041328-1799-4128-826c-5fed2738dfc4 · outbound

This paper cites Scalable training of deep learning machines by incremental block training with intra-block parallel optimization and blockwise model-update filtering.

Distributed Sign Momentum with Local Steps for Training Transformers Scalable training of deep learning machines by incremental block training with intra-block parallel optimization and blockwise model-update filtering

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:40.421073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T11:52:40.180850Z digest=sha256:4fcb5266f98199c6ca2ea1277b81046e5b7c12235003bb18754b1b2d81c2f6a2

Observation c0f2abf1-8916-49a6-b3df-6bc4d073c8a7 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Distributed Sign Momentum with Local Steps for Training Transformers Adam: A Method for Stochastic Optimization

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-12T11:52:40.196909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:52:40.196909Z digest=sha256:588925c58c3167ba271b8f3c780feea2a84074d984e167a35682d1ad62b7ae58

Observation 0805b0f3-12ef-46ef-a7a7-a85b7a3d5f47 · outbound

This paper cites 1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns.

Distributed Sign Momentum with Local Steps for Training Transformers 1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:40.410273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T11:52:40.213397Z digest=sha256:049d1fe2dd72950027c139a7dfd5544be790c20e1382f9fc3fd41b77d926138d

Observation 5fad4a0a-d5b0-4dd0-bea0-65a68086b0c7 · outbound

This paper cites CO2: Efficient Distributed Training with Full Communication-Computation Overlap.

Distributed Sign Momentum with Local Steps for Training Transformers CO2: Efficient Distributed Training with Full Communication-Computation Overlap

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T11:52:40.216892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:52:40.216892Z digest=sha256:1c63db9f4f82137e9ea2b5db1257273ce941258d5de42e60e3e6194be9213712

Pith citing papers

No inbound Pith citation observations are available.