Pith. sign in

Paper Citation Record · LEDGER

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees

As of 8 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2509.11254.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.11254 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T17:01:42.078459Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 83872f07-4923-4fbd-b9a1-03f44b2a66f2 · outbound

This paper cites Scaling distributed machine learning with the parameter server.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Scaling distributed machine learning with the parameter server

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.007942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.007942Z digest=sha256:d4fb8728291971b8fb78c59568b42d1c2328c9ace28e922986424de3d7153b21

Observation d9539af7-b0c0-4af2-8de0-8bd35ea37279 · outbound

This paper cites Bandwidth optimal all-reduce algorithms for clusters of workstations.Journal of Parallel and Distributed Computing, 69(2):117–124, 2009.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Bandwidth optimal all-reduce algorithms for clusters of workstations.Journal of Parallel and Distributed Computing, 69(2):117–124, 2009

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.011004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.011004Z digest=sha256:89419e6cccecc42f4b2a2703d7f2480c8f3b465d08088967d2aa90c8c58ef8d4

Observation fa0af9d4-aecb-4806-93dc-acd646b067de · outbound

This paper cites 1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees 1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.014246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.014246Z digest=sha256:3ee1661817f73e77b8a423e9f5d743304abc9e16cacbe48195dc40085af6e47b

Observation 8f70ee7f-a8b0-4246-935a-3d8be6330014 · outbound

This paper cites Project adam: Building an efficient and scalable deep learning training system.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Project adam: Building an efficient and scalable deep learning training system

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.017193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.017193Z digest=sha256:9e940272e55bdec67ecb603f72d4b3143051a645f910e027ccac701588106266

Observation a9ad12bd-9b49-40d6-a8b1-d79cdc2b55ae · outbound

This paper cites Qsgd: Communication-efficient sgd via gradient quantization and encoding.Advances in neural information processing systems, 30, 2017.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Qsgd: Communication-efficient sgd via gradient quantization and encoding.Advances in neural information processing systems, 30, 2017

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.020282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.020282Z digest=sha256:8ffe9c3592ce18e703b00848fffbc8869ed51900a8c57567450af64b11017dc4

Observation c090a467-3b2d-4520-b04e-df041fe5ddac · outbound

This paper cites Terngrad: Ternary gradients to reduce communication in distributed deep learning.Advances in neural information processing systems, 30, 2017.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Terngrad: Ternary gradients to reduce communication in distributed deep learning.Advances in neural information processing systems, 30, 2017

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.023417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.023417Z digest=sha256:eb29dbdfc578f01dc37282b30ccc6180e2d5b49091beb31a5a08e2fc26588e61

Observation 70983f45-eb54-46d5-a7e9-b4e2e50b3c62 · outbound

This paper cites Federated Learning: Strategies for Improving Communication Efficiency.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Federated Learning: Strategies for Improving Communication Efficiency

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.026238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.026238Z digest=sha256:6165a422816d754567f1e321d477e55eb994612f42dc815f078c0db937e44b31

Observation 5e083bab-dcb5-48a1-bbde-113d798d832d · outbound

This paper cites Gradient sparsification for communication-efficient distributed optimization.Advances in Neural Information Processing Systems, 31, 2018.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Gradient sparsification for communication-efficient distributed optimization.Advances in Neural Information Processing Systems, 31, 2018

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.029318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.029318Z digest=sha256:7ac6da83979da95e6e034ae56287492cb2a7ad906e5bc698e3710b8429c151f9

Observation 48c065bf-e9f9-47f4-bc2f-212fd1f91d5f · outbound

This paper cites Sparsified sgd with memory.Advances in neural information processing systems, 31, 2018.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Sparsified sgd with memory.Advances in neural information processing systems, 31, 2018

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.031993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.031993Z digest=sha256:90985132d756cf7732f5b06d98c689ee91d69d6cc177f22a4538c7f305c3360e

Observation 7d3beb5f-c2e7-4fcc-871e-a82c46bb0527 · outbound

This paper cites Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.034491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.034491Z digest=sha256:4532b59ed74bd6a88b0db1a733b08efcd99561323e2b4581e822246e5753e9f0

Observation 1a3dd5dc-f83f-4698-8b3f-271a689313bf · outbound

This paper cites GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.037020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.037020Z digest=sha256:e56bbd0a7cf99849958a672bbbda598107c180e25b788d734fb05169308dcc04

Observation ad88d29b-8031-4eb9-a05a-f450f1f0e019 · outbound

This paper cites Sltrain: a sparse plus low rank approach for parameter and memory efficient pretraining.Advances in Neural Information Processing Systems, 37:118267–118295, 2024.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Sltrain: a sparse plus low rank approach for parameter and memory efficient pretraining.Advances in Neural Information Processing Systems, 37:118267–118295, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.039844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.039844Z digest=sha256:1ea65062c2a37ea49815601d88d3df37bf8c189e80fb0c1d42715f0ac329bc03

Observation 95b7e784-4094-479f-b626-8868187e7af5 · outbound

This paper cites Enhancing Zeroth-order Fine-tuning for Language Models with Low-rank Structures.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Enhancing Zeroth-order Fine-tuning for Language Models with Low-rank Structures

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.042592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.042592Z digest=sha256:3b95e054798a46251631b091d35c3d7e2b3735c01b9b3076043d5db0aac67c06

Observation 6b27ab68-3b4f-41ab-9af7-1b0c2fe8f4ea · outbound

This paper cites Powersgd: Practical low-rank gradient compression for distributed optimization.Advances in Neural Information Processing Systems, 32, 2019.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Powersgd: Practical low-rank gradient compression for distributed optimization.Advances in Neural Information Processing Systems, 32, 2019

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.045428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.045428Z digest=sha256:d11e4fcda1ae596d13278d9c6cccbb5b42c4f301853a19b4873d2ef86741a3c8

Observation 4411f083-7653-4de7-8fea-d230e4f16aee · outbound

This paper cites Zero-shot text-to-image generation.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Zero-shot text-to-image generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.048192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.048192Z digest=sha256:0f598ebaadddfa08676535606da89d3b74ec6d69773eac9980977d8de464a68d

Observation 8a12f758-da24-43c7-be5d-46f6984f79e3 · outbound

This paper cites Angel-PTM: A Scalable and Economical Large-scale Pre-training System in Tencent.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Angel-PTM: A Scalable and Economical Large-scale Pre-training System in Tencent

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.051031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.051031Z digest=sha256:af75cab6939a073c0e96a1955b50ad27bdfa03c3a865cadb776f9e8644de5585

Observation 2b48aaf0-3f7c-4f10-a112-c695b0483338 · outbound

This paper cites Error feedback fixes signsgd and other gradient compression schemes.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Error feedback fixes signsgd and other gradient compression schemes

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.053851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.053851Z digest=sha256:dfe99068bd888d3d59609d33ea2c8e94c22870fbede446591538822bce73f5d7

Observation 638fb470-f9ea-4018-be12-ffd430cfb395 · outbound

This paper cites single-step power iteration.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees single-step power iteration

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.003318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.003318Z digest=sha256:cc30b42e7b01759c4253ccd99ccd53d96b6a4ffc0dcfc2d6548f7a64d6a74c6b

Observation 81bde3cf-94be-45fb-89bf-8b8cf9e25e79 · outbound

This paper cites Ef21: A new, simpler, theoretically better, and practically faster error feedback.Advances in Neural Information Processing Systems, 34:4384–4396, 2021.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Ef21: A new, simpler, theoretically better, and practically faster error feedback.Advances in Neural Information Processing Systems, 34:4384–4396, 2021

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.056500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.056500Z digest=sha256:8c74528be9e37476943bfbf04df423a69aea961431113a73268a11c9ba125ad3

Observation 5fad0cd2-3d78-43d2-a6af-c9d062d96753 · outbound

This paper cites Lower bounds and nearly optimal algorithms in distributed learning with communication compression.Advances in Neural Information Processing Systems, 35:18955–18969, 2022.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Lower bounds and nearly optimal algorithms in distributed learning with communication compression.Advances in Neural Information Processing Systems, 35:18955–18969, 2022

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.059202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.059202Z digest=sha256:665d8a44008281ecdf4f59d8df10cffc09c07a09c1560eff357e519ea1938dd2

Observation e65ccc3e-13ba-41a3-98b1-4059729db6ff · outbound

This paper cites The Error-Feedback Framework: Better Rates for SGD with Delayed Gradients and Compressed Communication.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees The Error-Feedback Framework: Better Rates for SGD with Delayed Gradients and Compressed Communication

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.061587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.061587Z digest=sha256:81f2a28478302a3d1e4f3a39874e30e8193d4a2d40aa101b7db2b6d003a7ae21

Observation c7395f9c-76d2-4d0c-8269-09c78faacc0a · outbound

This paper cites Momentum provably improves error feedback!Ad- vances in Neural Information Processing Systems, 36:76444–76495, 2023.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Momentum provably improves error feedback!Ad- vances in Neural Information Processing Systems, 36:76444–76495, 2023

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.064823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.064823Z digest=sha256:243cd5da240c00c12b1c276d70f2ab0edaaead69f82469c47a3e00da73aa6073

Observation 16680f9c-4920-4dce-bb4a-12a797d159a6 · outbound

This paper cites an unresolved cited work.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.067290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.067290Z digest=sha256:19914e600011cce98fd3333e335082f6226fa12fff037c5dcb706bdd5be29641

Observation e651f1b9-d25d-44b1-8392-24c2c159f715 · outbound

This paper cites Atomo: Communication-efficient learning via atomic sparsification.Advances in neural information processing systems, 31, 2018.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Atomo: Communication-efficient learning via atomic sparsification.Advances in neural information processing systems, 31, 2018

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.069707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.069707Z digest=sha256:b5c732c190827bed8a9416092a41b167cfc28eaf976b327db8e2d804957bf538

Observation 80269079-ba91-48c4-918b-88d47f9e1f0c · outbound

This paper cites Greedy low-rank gradient compression for distributed learning with convergence guarantees.arXiv preprint arXiv:2507.08784, 2025.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Greedy low-rank gradient compression for distributed learning with convergence guarantees.arXiv preprint arXiv:2507.08784, 2025

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.072449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.072449Z digest=sha256:822111450e47a9cdf90f3513ec23e1d366d6cdee5887c024f5592a66bd1b5ff0

Observation ccd0a2ae-c1ed-4255-8051-ae80ee6c9889 · outbound

This paper cites Subspace Optimization for Large Language Models with Convergence Guarantees.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Subspace Optimization for Large Language Models with Convergence Guarantees

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.075076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.075076Z digest=sha256:2ae89d7fd6c3d4c0aad95561dedd0f62549675817ff87badf4b64abfeda3a5c8

Observation 3971d6e2-aa22-4d9d-b675-6ac80d1f8972 · outbound

This paper cites an unresolved cited work.

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:42.078459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:42.078459Z digest=sha256:d73b20daeec6801b1bfd24f00a63c3fa5018b791158fb9a3b43f8788dc611d00

Pith citing papers

No inbound Pith citation observations are available.