Pith. sign in

Paper Citation Record · LEDGER

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees

As of 20 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2509.11254.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.11254 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T17:01:42.078459Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 83872f07-4923-4fbd-b9a1-03f44b2a66f2 · outbound

This paper cites Scaling distributed machine learning with the parameter server.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Scaling distributed machine learning with the parameter server

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.007942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.007942Z digest=sha256:488e5436419c1c29625a3eb70eb6fe67ea22e8bdb73ffe4b6fe6568c4836b025

Observation d9539af7-b0c0-4af2-8de0-8bd35ea37279 · outbound

This paper cites Bandwidth optimal all-reduce algorithms for clusters of workstations.Journal of Parallel and Distributed Computing, 69(2):117–124, 2009.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Bandwidth optimal all-reduce algorithms for clusters of workstations.Journal of Parallel and Distributed Computing, 69(2):117–124, 2009

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.011004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.011004Z digest=sha256:a7f55b326f531ac23a7efa35af4db72eb5703ae5420410b8ece6f5f9773d1ead

Observation fa0af9d4-aecb-4806-93dc-acd646b067de · outbound

This paper cites 1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees 1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.014246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.014246Z digest=sha256:4dcaacf85b26945f8e30938c9afa2228ea74ccacf3306b362c4180fc09f35c80

Observation 8f70ee7f-a8b0-4246-935a-3d8be6330014 · outbound

This paper cites Project adam: Building an efficient and scalable deep learning training system.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Project adam: Building an efficient and scalable deep learning training system

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.017193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.017193Z digest=sha256:b09310a4fdc78be6cde41d1ad366cfeaa13a7de91d8162f8e329a370d1b38d75

Observation a9ad12bd-9b49-40d6-a8b1-d79cdc2b55ae · outbound

This paper cites Qsgd: Communication-efficient sgd via gradient quantization and encoding.Advances in neural information processing systems, 30, 2017.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Qsgd: Communication-efficient sgd via gradient quantization and encoding.Advances in neural information processing systems, 30, 2017

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.020282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.020282Z digest=sha256:edf57087a862b3c64c81dbcbb5861cd5f15b03b64f28ab3c920fc883fd7bbf0a

Observation c090a467-3b2d-4520-b04e-df041fe5ddac · outbound

This paper cites Terngrad: Ternary gradients to reduce communication in distributed deep learning.Advances in neural information processing systems, 30, 2017.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Terngrad: Ternary gradients to reduce communication in distributed deep learning.Advances in neural information processing systems, 30, 2017

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.023417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.023417Z digest=sha256:49afbead72d0835fe33dfa9fff4cb65f61ef5e3c3a15cb7aef34224d9d2b4dda

Observation 70983f45-eb54-46d5-a7e9-b4e2e50b3c62 · outbound

This paper cites Federated Learning: Strategies for Improving Communication Efficiency.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Federated Learning: Strategies for Improving Communication Efficiency

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.026238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.026238Z digest=sha256:2a948893e703fc52ae67b63fdc7ac3c03e3c3c4325c6ac9aeb5a3fa457560fb4

Observation 5e083bab-dcb5-48a1-bbde-113d798d832d · outbound

This paper cites Gradient sparsification for communication-efficient distributed optimization.Advances in Neural Information Processing Systems, 31, 2018.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Gradient sparsification for communication-efficient distributed optimization.Advances in Neural Information Processing Systems, 31, 2018

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.029318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.029318Z digest=sha256:14ff931815c05f83eef4b812b114c876822a0028afa5f2c494129a0357265b43

Observation 48c065bf-e9f9-47f4-bc2f-212fd1f91d5f · outbound

This paper cites Sparsified sgd with memory.Advances in neural information processing systems, 31, 2018.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Sparsified sgd with memory.Advances in neural information processing systems, 31, 2018

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.031993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.031993Z digest=sha256:b4ffe10f0a3a91b107cc48739aefcece15c74c1abbbc41a57cf2d5330710bd65

Observation 7d3beb5f-c2e7-4fcc-871e-a82c46bb0527 · outbound

This paper cites Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.034491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.034491Z digest=sha256:dc110f451ac6e9314e71aaa9ad14f4e9d560d11169e9a551d67d7bfc8957f01f

Observation 1a3dd5dc-f83f-4698-8b3f-271a689313bf · outbound

This paper cites GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.037020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.037020Z digest=sha256:cbd9962a649145f629ebd5b434921df580a3ceedc3f8de51fc10db430e263915

Observation ad88d29b-8031-4eb9-a05a-f450f1f0e019 · outbound

This paper cites Sltrain: a sparse plus low rank approach for parameter and memory efficient pretraining.Advances in Neural Information Processing Systems, 37:118267–118295, 2024.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Sltrain: a sparse plus low rank approach for parameter and memory efficient pretraining.Advances in Neural Information Processing Systems, 37:118267–118295, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.039844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.039844Z digest=sha256:0d3de4a44680b6bf830b70c078b2d8461a99160266e0aac18fb8b7d845d937a7

Observation 95b7e784-4094-479f-b626-8868187e7af5 · outbound

This paper cites Enhancing Zeroth-order Fine-tuning for Language Models with Low-rank Structures.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Enhancing Zeroth-order Fine-tuning for Language Models with Low-rank Structures

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.042592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.042592Z digest=sha256:9da709110b8f5e6fbcef6284de40e2de1d9a1e4d747eed426cc4c333b1e507e8

Observation 6b27ab68-3b4f-41ab-9af7-1b0c2fe8f4ea · outbound

This paper cites Powersgd: Practical low-rank gradient compression for distributed optimization.Advances in Neural Information Processing Systems, 32, 2019.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Powersgd: Practical low-rank gradient compression for distributed optimization.Advances in Neural Information Processing Systems, 32, 2019

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.045428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.045428Z digest=sha256:52c6fca978ede42f72ae54dc2142aac5cc985d84ebfef8e6131b3c1d4eede619

Observation 4411f083-7653-4de7-8fea-d230e4f16aee · outbound

This paper cites Zero-shot text-to-image generation.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Zero-shot text-to-image generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.048192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.048192Z digest=sha256:7f76c4b0f47169ded6a33af2e450d7914c64369958efa2af6211d378d2ad1cce

Observation 8a12f758-da24-43c7-be5d-46f6984f79e3 · outbound

This paper cites Angel-PTM: A Scalable and Economical Large-scale Pre-training System in Tencent.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Angel-PTM: A Scalable and Economical Large-scale Pre-training System in Tencent

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.051031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.051031Z digest=sha256:b751b544b1914b23df1ab22f59c3bb1f1af04d216fd7e781668969e7bcd27f3e

Observation 2b48aaf0-3f7c-4f10-a112-c695b0483338 · outbound

This paper cites Error feedback fixes signsgd and other gradient compression schemes.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Error feedback fixes signsgd and other gradient compression schemes

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.053851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.053851Z digest=sha256:9e5b11559f72b12318d64fb21de23c54421fa5ad7e89201b39e660a45d82acff

Observation 638fb470-f9ea-4018-be12-ffd430cfb395 · outbound

This paper cites single-step power iteration.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees single-step power iteration

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.003318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.003318Z digest=sha256:73aa47231345d40955ab015bbe7fdaa6e49862731b20eac51bc173158dd75c46

Observation 81bde3cf-94be-45fb-89bf-8b8cf9e25e79 · outbound

This paper cites Ef21: A new, simpler, theoretically better, and practically faster error feedback.Advances in Neural Information Processing Systems, 34:4384–4396, 2021.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Ef21: A new, simpler, theoretically better, and practically faster error feedback.Advances in Neural Information Processing Systems, 34:4384–4396, 2021

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.056500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.056500Z digest=sha256:b6c887a033d9668f367a12fcdf40427a84d9f07230c7dbb69139cdf682552233

Observation 5fad0cd2-3d78-43d2-a6af-c9d062d96753 · outbound

This paper cites Lower bounds and nearly optimal algorithms in distributed learning with communication compression.Advances in Neural Information Processing Systems, 35:18955–18969, 2022.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Lower bounds and nearly optimal algorithms in distributed learning with communication compression.Advances in Neural Information Processing Systems, 35:18955–18969, 2022

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.059202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.059202Z digest=sha256:d4ad8fbe6b4e636922aab35a3e4f166f7430f9a9e2213a869720cb7c04a58eb9

Observation e65ccc3e-13ba-41a3-98b1-4059729db6ff · outbound

This paper cites The Error-Feedback Framework: Better Rates for SGD with Delayed Gradients and Compressed Communication.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees The Error-Feedback Framework: Better Rates for SGD with Delayed Gradients and Compressed Communication

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.061587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.061587Z digest=sha256:85a5238b5287e58170b819fcc00d470940c9ab157c752220f1e5776ba6eed8b6

Observation c7395f9c-76d2-4d0c-8269-09c78faacc0a · outbound

This paper cites Momentum provably improves error feedback!Ad- vances in Neural Information Processing Systems, 36:76444–76495, 2023.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Momentum provably improves error feedback!Ad- vances in Neural Information Processing Systems, 36:76444–76495, 2023

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.064823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.064823Z digest=sha256:41715caf5937bf21ab118eac22745a8c204facb87c2532f1b32e0bbb388fc260

Observation 16680f9c-4920-4dce-bb4a-12a797d159a6 · outbound

This paper cites an unresolved cited work.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.067290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.067290Z digest=sha256:2d8fbe169240cc9be39a31cb3e8ef10127b3b1345fb3b622e232d2542b0b29aa

Observation e651f1b9-d25d-44b1-8392-24c2c159f715 · outbound

This paper cites Atomo: Communication-efficient learning via atomic sparsification.Advances in neural information processing systems, 31, 2018.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Atomo: Communication-efficient learning via atomic sparsification.Advances in neural information processing systems, 31, 2018

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.069707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.069707Z digest=sha256:dc997552d41001fb00bd2e701b2b41dc1b2a6543c4dfbc913dfcfd2e2ecd966e

Observation 80269079-ba91-48c4-918b-88d47f9e1f0c · outbound

This paper cites Greedy low-rank gradient compression for distributed learning with convergence guarantees.arXiv preprint arXiv:2507.08784, 2025.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Greedy low-rank gradient compression for distributed learning with convergence guarantees.arXiv preprint arXiv:2507.08784, 2025

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.072449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.072449Z digest=sha256:6dde532e77939fd2096bcb73eef73446dc59cc3747ebe0ca458ee9938244d6d2

Observation ccd0a2ae-c1ed-4255-8051-ae80ee6c9889 · outbound

This paper cites Subspace Optimization for Large Language Models with Convergence Guarantees.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Subspace Optimization for Large Language Models with Convergence Guarantees

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.075076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.075076Z digest=sha256:1e5f3a4ce02cdb65d0481b6fbf7d7d34f973e4a7fa4f1d77dec0767aae828d75

Observation 3971d6e2-aa22-4d9d-b675-6ac80d1f8972 · outbound

This paper cites an unresolved cited work.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.078459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.078459Z digest=sha256:5d37d6acb4ef9a4bef44982a4658948f69a20b5be02c9ec52478c05fc53f5e75

Pith citing papers

No inbound Pith citation observations are available.