Pith. sign in

Paper Citation Record · LEDGER

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism

As of 15 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 3 inbound Pith citation observations for arXiv:2506.01260.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01260 v3

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:55:15.447206Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T18:53:47.187437Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T19:42:36.073867Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact1
  • verified fuzzy32
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 592c6e93-48c5-42ed-a329-6175452b8bf2 · outbound

This paper cites Transformers learn through gradual rank increase.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Transformers learn through gradual rank increase

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:19.678630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:55:12.762172Z digest=sha256:a7e2982a78ca28aefce639d2cc98c46c42bdf2907941adb7a1dda0e2ba517f6b

Observation c64f9b57-fbda-4ae6-9d11-f765a78e9e18 · outbound

This paper cites Qsgd: Communication-efficient sgd via gradient quantization and encoding.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Qsgd: Communication-efficient sgd via gradient quantization and encoding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:19.504881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:55:12.792678Z digest=sha256:8d503fad48f36017c7028b2b073e31e1375509957f3955e9fd7fade3290b614a

Observation d9f4a23d-1c97-4aff-8476-1549f87bd169 · outbound

This paper cites Dissecting adam: The sign, magnitude and variance of stochastic gradients.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Dissecting adam: The sign, magnitude and variance of stochastic gradients

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:19.330916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:55:12.887755Z digest=sha256:34e80d83eedb47154dd7350b3cafdf96c4895e672a9174348211cae7250a8305

Observation b0970cf0-a55e-4c8e-bd80-4d94511bc6d9 · outbound

This paper cites signsgd: Compressed optimisation for non-convex problems.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism signsgd: Compressed optimisation for non-convex problems

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:19.199270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:55:12.938754Z digest=sha256:6dc4fdcc32a21f71802355e477ffec4140dbde8a9c678ca16f40c43d889c8b23

Observation f0b24b53-ce89-4eeb-8b57-0b237e83940b · outbound

This paper cites Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.978236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:12.978236Z digest=sha256:a0e2b6fbd67556d33f64f8a67ec6a32c591515efcf7b34b6edf04178ff711cee

Observation 1add04da-defb-470c-90cd-02b6b5ca622c · outbound

This paper cites Does compressing activations help model parallel training? Proceedings of Machine Learning and Systems, 6: 0 239--252, 2024.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Does compressing activations help model parallel training? Proceedings of Machine Learning and Systems, 6: 0 239--252, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:19.045514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:55:13.031459Z digest=sha256:5055afa0a402c290f87ae4a6a050924de9a748d81072e887b2ce3c68400a7a4d

Observation c8e9404a-916d-4c17-804c-81fdebbac5f7 · outbound

This paper cites Low-rank gradient descent.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Low-rank gradient descent

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.889782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:55:13.092064Z digest=sha256:7d99b0eda2b9fb6cbb9f1c2e149dec8ccf778b89f637760d80baf92de79e0a63

Observation 68b31344-16d3-455d-bc0d-ca08bcbba0c4 · outbound

This paper cites DeepSeek LLMs , 2023.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism DeepSeek LLMs , 2023

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.694587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:55:13.159428Z digest=sha256:5c2f4512e61a20179f28969d7b5b85fa9bbddd44446bd205aae78bcd4ff2d572

Observation 4b22ed7a-9bb6-4124-bbc1-0c5aaec5abd2 · outbound

This paper cites A Simple Convergence Proof of Adam and Adagrad.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism A Simple Convergence Proof of Adam and Adagrad

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.202772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:13.202772Z digest=sha256:09eccfdaab00da305650e0f6a39b0d1b3b39a3c313acbf620d92a81737c3ba80

Observation f320d63e-c91c-46e4-9a45-535ee198c85f · outbound

This paper cites 8-bit Optimizers via Block-wise Quantization.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism 8-bit Optimizers via Block-wise Quantization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.250462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:13.250462Z digest=sha256:99e9a302f2874a7fd0b26244d04f393c6f6f25c9c9b6ec6b2f168082f76aa0f6

Observation 0ca8007e-5add-4001-a83b-f107719004ae · outbound

This paper cites Distributed deep learning in open collaborations.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Distributed deep learning in open collaborations

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.526941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:55:13.281551Z digest=sha256:c7cff3fa90622c5d0b5ab1efbc1d01e18a4fc386b4f0e4d119a2e19eafe27459

Observation f036de15-21bb-48ec-810e-d27ca7cb503e · outbound

This paper cites Attention is not all you need: Pure attention loses rank doubly exponentially with depth.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Attention is not all you need: Pure attention loses rank doubly exponentially with depth

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.456406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:55:13.326951Z digest=sha256:9b2a8878c90598e64a104a4e99a723c8697168c1ff412dc67cb6e83f7b50beff

Observation 6e906f3c-ff5a-495d-ab76-86472ce8777e · outbound

This paper cites DiLoCo: Distributed Low-Communication Training of Language Models.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism DiLoCo: Distributed Low-Communication Training of Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.360801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:13.360801Z digest=sha256:07aa423e9a770e05c37299d99d429927e6de89b442098103f4d93b2728a3db0c

Observation 657f8464-4b43-4ef3-a757-c4b350906cce · outbound

This paper cites The Llama 3 Herd of Models.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.407145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:13.407145Z digest=sha256:f902b81b765e7ac54243410a91ac9df8dba617ceacbcdad0babe3d8d7f6cdf11

Observation e1517c73-b813-4383-9d2b-c86d6c3edb31 · outbound

This paper cites Openwebtext corpus.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Openwebtext corpus

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.449970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:13.449970Z digest=sha256:d9856b561bf7f64c181bee080ff705d49a0209411f4b591c8669a44b23af1d2a

Observation 4dd99cc9-a68d-4bba-b2a7-7be5f5266f9b · outbound

This paper cites Gradient Descent Happens in a Tiny Subspace.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Gradient Descent Happens in a Tiny Subspace

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.490228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:13.490228Z digest=sha256:4fa2f23f5ef28241bc2cf44f5d46d262f6ebe19cf6ceb7dc8fb2dd9a03b8ea85

Observation af20aac7-bd93-48cd-8526-2d1f2499c4b3 · outbound

This paper cites Training compute-optimal large language models.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Training compute-optimal large language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.325287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:55:13.527372Z digest=sha256:ec07863084be07260cc88404153186418635e4c88e9393033cd899aa27ea05fb

Observation 683d281f-56b2-437a-8a1e-24729c94ee26 · outbound

This paper cites Gpipe: Efficient training of giant neural networks using pipeline parallelism.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Gpipe: Efficient training of giant neural networks using pipeline parallelism

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.561151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:13.561151Z digest=sha256:d8b32602c50249d5f633cf8efda21a381608bf9eb417da9ab80d38edac3b2227

Observation b4c27129-cdab-49cd-99c9-b8aa62ec37ca · outbound

This paper cites Error feedback fixes signsgd and other gradient compression schemes.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Error feedback fixes signsgd and other gradient compression schemes

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.601856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:13.601856Z digest=sha256:7bd2ee9252f7346d48a7276be0d061eb2999b39b85c8ddcbcd28c19f8895618d

Observation 64d9c80c-1fe9-42e7-bb18-22149f1a0fc3 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Adam: A Method for Stochastic Optimization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.644921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:13.644921Z digest=sha256:c0c8784fe4de37f37fb43cfed03dab2b2cfc3a7116710ff1a473940ce7cdb750

Observation c9e3c9a4-ab2e-456e-aeaa-ce89aaa02670 · outbound

This paper cites Big transfer (bit): General visual representation learning.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Big transfer (bit): General visual representation learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.191509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:55:13.685383Z digest=sha256:622e5d8cad0f88f3617cd658403e358638b6b6a7d9580539ef60f902aee9ed5d

Observation 28c35edf-fce5-4f31-9f2e-739c15297eb7 · outbound

This paper cites Decentralized stochastic optimization and gossip algorithms with compressed communication.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Decentralized stochastic optimization and gossip algorithms with compressed communication

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.111648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:55:13.722281Z digest=sha256:c2e07778750e2b55c99a252596818ac2e32fed476873c25f3dcfef419b599fbe

Observation 4e2ef206-c2d0-4d33-bc94-9ab270d35c10 · outbound

This paper cites A unified theory of decentralized sgd with changing topology and local updates.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism A unified theory of decentralized sgd with changing topology and local updates

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.022141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:55:13.774841Z digest=sha256:b4b1b47341432df1123ca1830e2094117ca3c54448cbf5dda4774913fa08451a

Observation 886cc201-60f9-4848-93b9-49a731d1a1f7 · outbound

This paper cites Imagenet classification with deep convolutional neural networks.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Imagenet classification with deep convolutional neural networks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.802110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:13.802110Z digest=sha256:8c07bc67a672d5a05a09951aabde624df71d1de1f0d6a4e265934ce3a457b10e

Observation 0d731396-90cc-43ae-bf02-63c8765e4443 · outbound

This paper cites Convergence of adam under relaxed assumptions.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Convergence of adam under relaxed assumptions

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.906167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:55:13.825919Z digest=sha256:2fba16711cfe0b8289844c2252ea2a8f454f6453c67b373b55af428a565e9fa8

Observation 4cc5152f-e7a0-4997-a5f0-5e9f547c6b78 · outbound

This paper cites Learning on transformers is provable low-rank and sparse: A one-layer analysis.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Learning on transformers is provable low-rank and sparse: A one-layer analysis

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.727061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:55:13.868539Z digest=sha256:21554f9ce9b0c1c0743694ff67aa01c88fd568ef04d3f425aeeea33c292b0472

Observation 19dd66f6-69f6-4bb8-a90d-8468618ce0ff · outbound

This paper cites PyTorch Distributed: Experiences on Accelerating Data Parallel Training.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism PyTorch Distributed: Experiences on Accelerating Data Parallel Training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.906478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:13.906478Z digest=sha256:67306e7ccd533dea1be80131fe67383396644a83a044d254c7be3e93e094a8a7

Observation 732b3de9-bc91-4a0d-a6ca-47b1d0f59ec7 · outbound

This paper cites Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.641992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:55:13.949599Z digest=sha256:cee10cff32cfcb732fa9d69c1437d98aaaf3cd8674b32ab0176d27ca52a9b6b6

Observation 0bd52271-10b8-4d50-a78c-22c4a00bb911 · outbound

This paper cites TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.990201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:13.990201Z digest=sha256:d1c25f025afd7553cf9378bf38dd56149b7ff0059dedf403c71f4321d44701a9

Observation e8cf2029-b266-4302-bfb0-99f8904247e0 · outbound

This paper cites Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.033583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:14.033583Z digest=sha256:7b38690b2c17604afc85b5412f7168725ae1d075bdabe834b9a09383808ffe5c

Observation 6f8f90bb-c2ea-4452-9de8-6966bae1c399 · outbound

This paper cites Adam$^+$: A Stochastic Method with Adaptive Variance Reduction.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Adam$^+$: A Stochastic Method with Adaptive Variance Reduction

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:55:15.848572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:55:14.067024Z digest=sha256:6a6c070d0edd11b4be64a00582022eef1a209e2f35f9768b417a5ee16ac87ddd

Observation 9d2acdea-e1c8-4755-8b44-50e70a49009d · outbound

This paper cites Decoupled Weight Decay Regularization.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Decoupled Weight Decay Regularization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.108295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:14.108295Z digest=sha256:144a826635e75818bea6a42b379c8599c0e03b1781bd89d0b349a9bf79fe3759

Observation 0dc712ca-9b3e-4e6c-99e1-0292bccb2473 · outbound

This paper cites Pointer sentinel mixture models, 2016.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Pointer sentinel mixture models, 2016

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.137281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:14.137281Z digest=sha256:48f4deed8a44f3adcdc1340a6c2bb88bb5964e0b6b8474a7d4064675c771a805

Observation 8c77da94-f270-4761-bd07-7264e30cae9f · outbound

This paper cites Efficient large-scale language model training on gpu clusters using megatron-lm.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Efficient large-scale language model training on gpu clusters using megatron-lm

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.176931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:14.176931Z digest=sha256:df34a2c257facfd88e90875bf7e7b783af2b628da516ae5969546a17cc37d25d

Observation fc818491-3547-47da-94c9-deac8862c8f2 · outbound

This paper cites Decoupled momentum optimization.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Decoupled momentum optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.217550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:14.217550Z digest=sha256:ed7b890464cc8f7de2bfe16d97a2e4026aea103741995a9be1dd2bb6318028d5

Observation 23dbdad6-0ba3-4c34-afb7-f26c30eb8305 · outbound

This paper cites Ai and compute.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Ai and compute

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.532560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:55:14.258302Z digest=sha256:c338afe4d2b6fe828424ceddc2e35bbcd297d90407fa09b33e4bd129e6bb967c

Observation 06aeac6e-0e51-469d-9d5e-ff5d00d31a4f · outbound

This paper cites an unresolved cited work.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.298139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:14.298139Z digest=sha256:418de070fefe970e7fbe8cc5759478d5901aac8240d0c48041eba781e6dd1368

Observation 084b05a6-4f5a-47db-801b-fb9d6b952adf · outbound

This paper cites PanGu-{\Sigma}: Towards Trillion Parameter Language Model with Sparse Heterogeneous Computing.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism PanGu-{\Sigma}: Towards Trillion Parameter Language Model with Sparse Heterogeneous Computing

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.338207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:14.338207Z digest=sha256:a56d5b4b96f3e0578f07dcdb25d5791afce26498f95813ecc20c6b2dbfc28c0d

Observation 95d26b16-c41d-4c86-b777-8211317f30dd · outbound

This paper cites Activations and gradients compression for model-parallel training.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Activations and gradients compression for model-parallel training

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.445422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:55:14.379343Z digest=sha256:ab0f9008e61cb547a463b4846b48f55ceab8d3dd09a24ad96f300cbf0e792030

Observation 847e1442-37ef-42f9-b5bd-b2febdfc801f · outbound

This paper cites Towards crowdsourced training of large neural networks using decentralized mixture-of-experts.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Towards crowdsourced training of large neural networks using decentralized mixture-of-experts

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.356432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:55:14.444090Z digest=sha256:0cadb61246ac73c48cb9f9e0d317f576dc49e022b3ab6b576e92d0daae03b561

Observation ea379800-0903-4da6-9da1-a2b0399d6d36 · outbound

This paper cites Moshpit sgd: Communication-efficient decentralized training on heterogeneous unreliable devices.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Moshpit sgd: Communication-efficient decentralized training on heterogeneous unreliable devices

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.271867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:55:14.490832Z digest=sha256:7dd32235aa496ad384cd2ed18391471ba2106fd569412baf5ea750fc8360f500

Observation ea093586-11fc-473b-a71d-e4a9a660c855 · outbound

This paper cites Swarm parallelism: Training large models can be surprisingly communication-efficient.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Swarm parallelism: Training large models can be surprisingly communication-efficient

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.187559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:55:14.539044Z digest=sha256:9962e575ce4e0726c52fb6691f0a8a01e7ffb2db60d4e68330850525fdf7ec28

Observation 01da2141-3467-4517-86e1-04d69921597a · outbound

This paper cites Inheritune: Training smaller yet more attentive language models, 2024.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Inheritune: Training smaller yet more attentive language models, 2024

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.585859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:14.585859Z digest=sha256:1f0743d7dad1f48cc8d22035b627b4b9e67ad5cd88907e9ddf8fbb4e899e11ab

Observation 28adb1f0-5240-469f-a239-f90147d83e63 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.633452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:14.633452Z digest=sha256:beeb20f6d09cf2a6e1c6f319f65f9a6b235f0dae8088f924acf9a5dbe13e515b

Observation 0c5ed812-b146-4583-acd0-2f391285cda7 · outbound

This paper cites 1-bit adam: Communication efficient large-scale training with adam’s convergence speed.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism 1-bit adam: Communication efficient large-scale training with adam’s convergence speed

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.120045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:55:14.675721Z digest=sha256:75bfc32e843822c91b859e7f416bef7e2028bfc109339e92cb2069315bfb64ca

Observation d3f47848-25af-49fb-8b0c-f25666683362 · outbound

This paper cites Scaling the summit: deploying the world’s fastest supercomputer.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Scaling the summit: deploying the world’s fastest supercomputer

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.055192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:55:14.729222Z digest=sha256:3e767da5def0d4ec8949d95764a78eb605ddaeafb0547b6c68385d993ae46d0f

Observation 3670813a-87e0-4f51-ae8f-e6a7fff0a382 · outbound

This paper cites Powersgd: Practical low-rank gradient compression for distributed optimization.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Powersgd: Practical low-rank gradient compression for distributed optimization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.772989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:14.772989Z digest=sha256:d137ac69f42d8c35003c5827412f659e3780ce03864eaef6d5a46ea38db064a9

Observation dcaaf8b2-61dc-427c-bae4-f0c9bf3f473a · outbound

This paper cites Kalman Gradient Descent: Adaptive Variance Reduction in Stochastic Optimization.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Kalman Gradient Descent: Adaptive Variance Reduction in Stochastic Optimization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.838055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:14.838055Z digest=sha256:c9ca781fe9b37d619d5192a3d7f75d56eef72b2e6924248aa902b309c62a9729

Observation 6874d0b4-192d-415b-96a5-3531593fc3fa · outbound

This paper cites Variance reduction for stochastic gradient optimization.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Variance reduction for stochastic gradient optimization

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:16.949401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:55:14.864069Z digest=sha256:2bf1b665815bb6a90d67377104ea1e3371597f324bae2ade851a90035f29431c

Observation 7f84e35e-cba5-42b3-8109-1a64f96e43ab · outbound

This paper cites Atomo: Communication-efficient learning via atomic sparsification.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Atomo: Communication-efficient learning via atomic sparsification

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:16.890495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:55:14.902786Z digest=sha256:d630c6de00dd1636e064e3eedd2362b122dc7e77abb62c2d46744360cad7fd60

Observation cc829027-62a7-422f-915c-8e12a8b7961c · outbound

This paper cites Pufferfish: Communication-efficient models at no extra cost.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Pufferfish: Communication-efficient models at no extra cost

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:16.804411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:55:14.942915Z digest=sha256:938f2f5f6afde3b0033cc8943d5fa0ea98899538bc27a9b3b669aaf49f2dde34

Observation 2193dd3a-4213-465e-b126-d8717960a4fe · outbound

This paper cites Efficient distributed learning with sparsity.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Efficient distributed learning with sparsity

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:16.716469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:55:15.004690Z digest=sha256:240b1518c4fa451eca4183ea82fcd4ba5666a49d7a42b26e94132009078fe661

Observation d4b889cc-8acb-4882-82c0-a01bebc72cff · outbound

This paper cites Cocktailsgd: Fine-tuning foundation models over 500mbps networks.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Cocktailsgd: Fine-tuning foundation models over 500mbps networks

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:16.633501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:55:15.041815Z digest=sha256:9339e8bc5752d9ca5a1f1e09e000e52a247d8f7685668b305ed7b2401526cd30

Observation bc257cec-5b28-4b92-8d2c-e30af183a949 · outbound

This paper cites Gradient sparsification for communication-efficient distributed optimization.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Gradient sparsification for communication-efficient distributed optimization

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:16.537448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:55:15.092180Z digest=sha256:b978ebbf9982f7b832682735832a7a8dd7e200f96553bae417a04dc4ba754bff

Observation 74a6b4c1-4926-4340-bc05-6b093c11153f · outbound

This paper cites Error compensated quantized sgd and its applications to large-scale distributed optimization.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Error compensated quantized sgd and its applications to large-scale distributed optimization

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:16.430575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:55:15.139453Z digest=sha256:4c1aa714359504af74717323878c38972364b07b1ee38b5ba0e633208db30ea6

Observation d054a76c-f635-49b6-b8ac-0faffc2dc425 · outbound

This paper cites A Spectral Condition for Feature Learning.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism A Spectral Condition for Feature Learning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:15.171300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:15.171300Z digest=sha256:9a2b2bd1aba80ed14bdc3988ea9875f3dac289fbce84b5d502731e1ef5d7f418

Observation fc8965be-0e92-426b-80f8-55b176dddbf8 · outbound

This paper cites Stochastic Gradient Variance Reduction by Solving a Filtering Problem.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Stochastic Gradient Variance Reduction by Solving a Filtering Problem

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:15.205516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:15.205516Z digest=sha256:5079586ca4de133f67e09ef52f4eb93fd451603482c6a01630ed6b3476e030e6

Observation e20f2efd-df6b-429e-8765-f24f65a631e5 · outbound

This paper cites Decentralized training of foundation models in heterogeneous environments.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Decentralized training of foundation models in heterogeneous environments

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:16.241381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:55:15.244018Z digest=sha256:0dc7f4c385cd0718502809865604a248e0e6382884d74d758e5c08032cc8ae1c

Observation 62d9db7f-7614-4cbd-9382-b7706f10294e · outbound

This paper cites Convergence Guarantees for RMSProp and Adam in Generalized-smooth Non-convex Optimization with Affine Noise Variance.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Convergence Guarantees for RMSProp and Adam in Generalized-smooth Non-convex Optimization with Affine Noise Variance

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:15.284168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:15.284168Z digest=sha256:ea1237fc81f04d532310ac4fd569fafd393ecffef30cc68864c9c451a3115353

Observation 54ddbd28-fa93-40ef-a94b-65c092a6f745 · outbound

This paper cites ZerO Initialization: Initializing Neural Networks with only Zeros and Ones.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism ZerO Initialization: Initializing Neural Networks with only Zeros and Ones

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:15.314868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:15.314868Z digest=sha256:f7347ff00190bc6a6b80c706f139f59d9df0dc168a7bc7feb1de9144a7765881

Observation 529c2762-7a6b-4c2d-b08d-0186b91e35ff · outbound

This paper cites GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:15.366875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:15.366875Z digest=sha256:dbd9529db5b7ff975c4c59d2eecffbc7eb9e5e3e463362fe4b560097ce126cea

Observation 03303144-9887-4265-b371-af3dcdde92c0 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:15.412077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:15.412077Z digest=sha256:83b0aee2f8e6e8e85bc8e57369479fca65ba8a5f7461d61b42b96e41115423e6

Observation 84ae9fa8-ad4c-46f6-8881-7f811796074d · outbound

This paper cites Aligning books and movies: Towards story-like visual explanations by watching movies and reading books.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Aligning books and movies: Towards story-like visual explanations by watching movies and reading books

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:16.124732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:55:15.447206Z digest=sha256:ff5f5165112c864841cfa50512a576cc43871fdab0727f67cb49b889fe5e3deb

Pith citing papers

Observation 9a3fc158-0571-435a-9652-263a421bd0b8 · inbound

On the Surprising Effectiveness of a Single Global Merging in Decentralized Learning cites this paper.

On the Surprising Effectiveness of a Single Global Merging in Decentralized Learning Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-07-09T00:19:35.965961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-19T05:39:53.088948Z digest=sha256:11d73cfc0b6359ddc0e4e950aaac1f23e0570411125dc9213bbd2adb43be9dd2

Observation 5ab30c17-9f1e-48dc-b734-49a5b9295181 · inbound

ResBM: Residual Bottleneck Models for Low-Bandwidth Pipeline Parallelism cites this paper.

ResBM: Residual Bottleneck Models for Low-Bandwidth Pipeline Parallelism Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-09T00:19:35.965961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T16:36:15.760164Z digest=sha256:8114b896e11c179622cf70fc390acc2987a630d080543ebfdc3733af8f6f25eb

Observation d3d7d422-f931-43f2-85ba-457aed58aacd · inbound

GNMR: Runtime Stability Control for Low-Precision Large Language Model Training cites this paper.

GNMR: Runtime Stability Control for Low-Precision Large Language Model Training Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T00:19:35.965961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T18:53:47.187437Z digest=sha256:2b4e3387fb4c11b58a760e7fc1af353c0aeeda7fd3aa9186e736a46544ae6275