Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T17:01:42.078459Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2509.11254.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T17:01:42.078459Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
27 of 27 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 83872f07-4923-4fbd-b9a1-03f44b2a66f2 · outbound
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Scaling distributed machine learning with the parameter server
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9539af7-b0c0-4af2-8de0-8bd35ea37279 · outbound
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Bandwidth optimal all-reduce algorithms for clusters of workstations.Journal of Parallel and Distributed Computing, 69(2):117–124, 2009
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa0af9d4-aecb-4806-93dc-acd646b067de · outbound
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees 1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f70ee7f-a8b0-4246-935a-3d8be6330014 · outbound
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Project adam: Building an efficient and scalable deep learning training system
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9ad12bd-9b49-40d6-a8b1-d79cdc2b55ae · outbound
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Qsgd: Communication-efficient sgd via gradient quantization and encoding.Advances in neural information processing systems, 30, 2017
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c090a467-3b2d-4520-b04e-df041fe5ddac · outbound
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Terngrad: Ternary gradients to reduce communication in distributed deep learning.Advances in neural information processing systems, 30, 2017
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70983f45-eb54-46d5-a7e9-b4e2e50b3c62 · outbound
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Federated Learning: Strategies for Improving Communication Efficiency
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e083bab-dcb5-48a1-bbde-113d798d832d · outbound
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Gradient sparsification for communication-efficient distributed optimization.Advances in Neural Information Processing Systems, 31, 2018
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48c065bf-e9f9-47f4-bc2f-212fd1f91d5f · outbound
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Sparsified sgd with memory.Advances in neural information processing systems, 31, 2018
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d3beb5f-c2e7-4fcc-871e-a82c46bb0527 · outbound
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a3dd5dc-f83f-4698-8b3f-271a689313bf · outbound
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad88d29b-8031-4eb9-a05a-f450f1f0e019 · outbound
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Sltrain: a sparse plus low rank approach for parameter and memory efficient pretraining.Advances in Neural Information Processing Systems, 37:118267–118295, 2024
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95b7e784-4094-479f-b626-8868187e7af5 · outbound
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Enhancing Zeroth-order Fine-tuning for Language Models with Low-rank Structures
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b27ab68-3b4f-41ab-9af7-1b0c2fe8f4ea · outbound
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Powersgd: Practical low-rank gradient compression for distributed optimization.Advances in Neural Information Processing Systems, 32, 2019
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4411f083-7653-4de7-8fea-d230e4f16aee · outbound
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Zero-shot text-to-image generation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a12f758-da24-43c7-be5d-46f6984f79e3 · outbound
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Angel-PTM: A Scalable and Economical Large-scale Pre-training System in Tencent
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b48aaf0-3f7c-4f10-a112-c695b0483338 · outbound
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Error feedback fixes signsgd and other gradient compression schemes
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 638fb470-f9ea-4018-be12-ffd430cfb395 · outbound
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees single-step power iteration
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81bde3cf-94be-45fb-89bf-8b8cf9e25e79 · outbound
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Ef21: A new, simpler, theoretically better, and practically faster error feedback.Advances in Neural Information Processing Systems, 34:4384–4396, 2021
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fad0cd2-3d78-43d2-a6af-c9d062d96753 · outbound
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Lower bounds and nearly optimal algorithms in distributed learning with communication compression.Advances in Neural Information Processing Systems, 35:18955–18969, 2022
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e65ccc3e-13ba-41a3-98b1-4059729db6ff · outbound
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees The Error-Feedback Framework: Better Rates for SGD with Delayed Gradients and Compressed Communication
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7395f9c-76d2-4d0c-8269-09c78faacc0a · outbound
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Momentum provably improves error feedback!Ad- vances in Neural Information Processing Systems, 36:76444–76495, 2023
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16680f9c-4920-4dce-bb4a-12a797d159a6 · outbound
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e651f1b9-d25d-44b1-8392-24c2c159f715 · outbound
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Atomo: Communication-efficient learning via atomic sparsification.Advances in neural information processing systems, 31, 2018
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80269079-ba91-48c4-918b-88d47f9e1f0c · outbound
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Greedy low-rank gradient compression for distributed learning with convergence guarantees.arXiv preprint arXiv:2507.08784, 2025
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccd0a2ae-c1ed-4255-8051-ae80ee6c9889 · outbound
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Subspace Optimization for Large Language Models with Convergence Guarantees
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3971d6e2-aa22-4d9d-b675-6ac80d1f8972 · outbound
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.