Pith. sign in

Paper Citation Record · LEDGER

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended

As of 20 August 2026, this Paper Citation Record lists 100 of 100 outbound references and 0 inbound Pith citation observations for arXiv:2605.30728.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.30728 v1

Coverage vector

measured 100 of 100 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T23:31:27.783419Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 100 outbound references displayed

  • verified exact5
  • verified fuzzy0
  • unresolved95
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 58be9c57-58c9-4bc0-9f23-453c20e8f45a · outbound

This paper cites https: //docs.nvidia.com/cuda/cuda-c-programming-guide#compressible- memory.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended https: //docs.nvidia.com/cuda/cuda-c-programming-guide#compressible- memory

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:22020529a19022b996c478ee72cb1b07028baf541a440c70c9fe9779ed62e279

Observation 480b1c84-24c2-46a0-825c-395fcf9625ae · outbound

This paper cites https://docs.nvidia.com/cuda/cuda-c-programming-guide/#device- memory-accesses.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended https://docs.nvidia.com/cuda/cuda-c-programming-guide/#device- memory-accesses

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:cfa96d806d15868233006adacde0be358581211f42ae21b079f1d3e09cc40578

Observation fd5618a2-0317-459a-ac84-e1571df69b9a · outbound

This paper cites https://docs.nvidia.com/cuda/cuda-c-programming- guide/#hardware-implementation.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended https://docs.nvidia.com/cuda/cuda-c-programming- guide/#hardware-implementation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:6f50c46a4566d880970962186909538ec870a22ef15b7dcbd4365f7aab3ae8e3

Observation c6ce86f4-3a34-4b03-85e4-44e725cd109d · outbound

This paper cites https://docs.nvidia.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended https://docs.nvidia

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:39797573ab05eec64a7ee5bccccd453bcedf19b98da0f1092d01f9b734f87115

Observation edd49fed-8a93-4b38-b361-70fe5cd6f868 · outbound

This paper cites https://catalog.ngc.nvidia.com/orgs/nvidia/teams/ dle/models/dlrm_base_tf2_ckpt_ds-criteo-fl15.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended https://catalog.ngc.nvidia.com/orgs/nvidia/teams/ dle/models/dlrm_base_tf2_ckpt_ds-criteo-fl15

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:d19e5f3d4fb67be6f345697891d02b9de6a0e630e4461858849d9bb3db4b1183

Observation e3f46a52-7d0d-48ce-9f6d-4b13a972d88f · outbound

This paper cites https:// developer.nvidia.com/nvcomp.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended https:// developer.nvidia.com/nvcomp

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:f03591d9d83235739df31b67ee49487d2d7124ba54f6d21e4e4106034a33432d

Observation 9421c787-6802-4ed6-91f7-9811c919da51 · outbound

This paper cites an unresolved cited work.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:9ced8708fc424cff121213ec0d8c873e421dc7039dcd80dae305f89a9685a704

Observation fb1ede28-9520-4da8-93f0-304e4d149a7a · outbound

This paper cites https://resources.nvidia.com/en-us- tensor-core/nvidia-tensor-core-gpu-datasheet.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended https://resources.nvidia.com/en-us- tensor-core/nvidia-tensor-core-gpu-datasheet

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:7b6192ef0ffd20b855bde9783e811cd1c41a914396384d7ca58de5b645e7ba2a

Observation d811ba6d-25ec-41ca-8a15-f0e29468fd74 · outbound

This paper cites https://images.nvidia.com/content/ technologies/volta/pdf/volta-v100-datasheet-update-us-1165301- r5.pdf.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended https://images.nvidia.com/content/ technologies/volta/pdf/volta-v100-datasheet-update-us-1165301- r5.pdf

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:f3f04bdc704035585923bc648812fabf2cb95a3725032d27220eb664f492de5a

Observation c244c94d-5fa2-49c1-923d-becefcb4df64 · outbound

This paper cites https://www.nvidia.com/en-in/data-center/ nvlink/.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended https://www.nvidia.com/en-in/data-center/ nvlink/

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:671b2277e29f04bfaa1df1526fc12f17a68ae5b67305e68eecfdbd9d6129e738

Observation 933516ca-e192-4959-9604-4ca54d7e8e01 · outbound

This paper cites https://labs.criteo.com/2013/12/download- terabyte-click-logs-2/.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended https://labs.criteo.com/2013/12/download- terabyte-click-logs-2/

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:aee16aa0392504dae58047055f92a7671bbe101c8058d51e3da98739b65961fc

Observation fc2e4bdc-ca83-411b-bb8d-6dfa9371be66 · outbound

This paper cites Understanding training efficiency of deep learning recommendation models at scale.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Understanding training efficiency of deep learning recommendation models at scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:39a65b119d221567450cb8f7fcc8c99c84c699bee08b87c52c4e5d7def8fca42

Observation 5c138396-e6c8-4561-8843-48da8f83ef8c · outbound

This paper cites Accelerating gpu data processing using fastlanes compression.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Accelerating gpu data processing using fastlanes compression

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:02bd032cabfd85743b8ab43dba6a6da92ef92d92ccaa7592a2218d7a51b7fdf3

Observation d74e3866-f4d7-4aa5-8b46-90a9e0cb6582 · outbound

This paper cites Bagpipe: Accelerating deep recommendation model training.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Bagpipe: Accelerating deep recommendation model training

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:4ca76a1e05d86a92ea63784b9866b783749998a5da35674b1a2a417188b29bd5

Observation 78f89cf1-d351-4faa-8a8e-0d06b1ffe084 · outbound

This paper cites Graph neural network training systems: A performance comparison of full-graph and mini-batch.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Graph neural network training systems: A performance comparison of full-graph and mini-batch

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:45478f2aaa30c4ebad4a9363916e76956cd04f4c87724df1b083019e950185ee

Observation eff29eb0-b998-453a-842d-b52dbb8b014e · outbound

This paper cites Aware: Workload-aware, redundancy-exploiting linear algebra.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Aware: Workload-aware, redundancy-exploiting linear algebra

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:f8593b4f2403b371f23e7c9e7755d3616a640aa1e885fce2c144b84de2d39109

Observation 789f85ef-ebca-4da1-b721-e1a34d2b3bea · outbound

This paper cites Deep gaussian embedding of graphs: Unsupervised inductive learning via ranking.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Deep gaussian embedding of graphs: Unsupervised inductive learning via ranking

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:e1715b536fef4c2ee796c1481f142aa53192dfae684fb6e92c9ab2046861324b

Observation 3abfc8c3-2f8a-427a-9b8f-9b01c21cd119 · outbound

This paper cites Molecular generative graph neural networks for drug discovery.Neurocomputing, 450:242–252, 2021.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Molecular generative graph neural networks for drug discovery.Neurocomputing, 450:242–252, 2021

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:2a32a0b8172d1fec2e531ce478631f7fb3309328b3199c6d5c97f1a122bb1487

Observation 7d4e8955-bfff-4347-94b6-fcd9a52d7157 · outbound

This paper cites Fcbench: Cross-domain benchmarking of lossless compression for floating-point data.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Fcbench: Cross-domain benchmarking of lossless compression for floating-point data

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:f821979c0076249b6eaafe1ac2a0f63c1d5fd1d125782e86b4360f0fad5a4a15

Observation db428f5c-016d-426f-bbfc-03506bdcd199 · outbound

This paper cites Learned image compression with discretized gaussian mixture likeli- hoods and attention modules.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Learned image compression with discretized gaussian mixture likeli- hoods and attention modules

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:1b345b61e0199550dec3f0b2f0797c86270cfdeecbb4d76eb335df65192267db

Observation 0839d043-6779-4dd2-a5cd-91caea7330ea · outbound

This paper cites The trade-offs of model size in large recommendation models : 100gb to 10mb criteo-tb dlrm model.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended The trade-offs of model size in large recommendation models : 100gb to 10mb criteo-tb dlrm model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:3bb7c9a7a1a2ce8aea25a0a0dc23d0b4eaa1dae68b84f72b4518cf58bb47fef1

Observation e3425b2b-04c2-49db-85a9-76a9451c2f02 · outbound

This paper cites Gpt3.int8(): 8-bit matrix multiplication for transformers at scale.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Gpt3.int8(): 8-bit matrix multiplication for transformers at scale

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:ae001cd068a9a4da38004c8d1c850902b954249b15d560ac90bdd9d6b85a719b

Observation 7ad0f06f-e73c-455e-a7e8-b6fd9441dfe8 · outbound

This paper cites Qlora: Efficient finetuning of quantized llms.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Qlora: Efficient finetuning of quantized llms

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:49244a2f2a6703cf6ec8a594a51470a923799b8e6ddeb77de8abb48ea8759418

Observation 05a7fad3-6c85-4f36-99cf-b3fba532ce35 · outbound

This paper cites Accuracy is not all you need.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Accuracy is not all you need

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:184e98a0a867bfc63ef3484c3cce686c41e7cd19c66012fd851b521c2b28bdb6

Observation 3f38b486-b791-47bf-95d4-8253d4e00fd9 · outbound

This paper cites Haas, Frederick R.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Haas, Frederick R

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:c87901a752c9bfb7fcbff39e2019d97913bdde79b41d8f188916eac0e329bcf4

Observation 774f680c-1754-45d6-be16-5237fb861c2b · outbound

This paper cites Zipserv: Fast and memory-efficient llm inference with hardware-aware lossless compression.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Zipserv: Fast and memory-efficient llm inference with hardware-aware lossless compression

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:1477d6708db724cb3fe997ff564b23186eeb53182c00894f257a299606af54a6

Observation a8d4e7e4-6dd8-4ec6-ba07-e17ee40e0cd7 · outbound

This paper cites A Frequency-aware Software Cache for Large Recommendation System Embeddings.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended A Frequency-aware Software Cache for Large Recommendation System Embeddings

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:32:46.928700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:e27162415273d6474d8a3ea4a1c1e71323f3b986a02165a7182571855b7f9f25

Observation d956aa1e-e3d1-4c1d-81da-2739e64cccf0 · outbound

This paper cites Sahu, Marco Canini, and Amedeo Sapio.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Sahu, Marco Canini, and Amedeo Sapio

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:3994e8e33b41cc94689a7bafcf7c78531842036612c5d0deb7889911a41c1e51

Observation 3460820b-b559-4f5c-b189-27e3f38ae811 · outbound

This paper cites Accelerating Communication in Deep Learning Recommendation Model Training with Dual-Level Adaptive Lossy Compression.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Accelerating Communication in Deep Learning Recommendation Model Training with Dual-Level Adaptive Lossy Compression

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:6f8a0f89a44691541f681ef960a9b3f3564af9ca76f930beb5091670ec76e21a

Observation 45aa004e-7674-4d6a-920c-977a87340210 · outbound

This paper cites Mahoney, and Kurt Keutzer.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Mahoney, and Kurt Keutzer

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:be3030627d7c7226bc284bb5395240f8e7cdba75832c34ad004dd68647029fda

Observation e78ee4b7-bc97-4783-a6f0-1747a1a7a707 · outbound

This paper cites Lee, David Brooks, and Carole- Jean Wu.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Lee, David Brooks, and Carole- Jean Wu

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:22dac8510816e020bf740c73035a60a12e8671bbd9b93e61558df70a9d7bb046

Observation 81c6980d-5dec-4d42-aec9-338892be3f14 · outbound

This paper cites Inductive representa- tion learning on large graphs.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Inductive representa- tion learning on large graphs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:10a08f8444d9774d1bf3d16d6cfde544da59d15196a90737593482fd15d36ffa

Observation 7905fbc9-36b0-4c2e-a9d4-335cc43ea1ab · outbound

This paper cites How to optimize data transfers in cuda c/c++.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended How to optimize data transfers in cuda c/c++

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:d9a34676f61e83ef1210e891151adbe034247994c8dc19616a6cd486f991df18

Observation 3958feea-a664-4238-be44-abe85260ca37 · outbound

This paper cites Natural compression for distributed deep learning.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Natural compression for distributed deep learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:2f31be9b57de42946c305e747a56115ba5552814a503b979ee54c65d5f0302f6

Observation f32cc364-16fd-41b4-8f85-ca1632d963b2 · outbound

This paper cites Open graph benchmark: Datasets for machine learning on graphs.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Open graph benchmark: Datasets for machine learning on graphs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:9e4252b1d4031c5357e6e3ff4b995db9174f2cb98e1c50f6c11b81f5ac5481af

Observation b4d4085d-2806-4c81-aade-4ba18d4596ca · outbound

This paper cites an unresolved cited work.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:e241c9f282e2422ee758c72dbaa610f2fae1a13e87a9d9e70d32b40e66d64127

Observation 0d3aa107-1505-4a63-8729-ea11f554b084 · outbound

This paper cites Grow: A row-stationary sparse-dense gemm accelerator for memory-efficient graph convolutional neural networks.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Grow: A row-stationary sparse-dense gemm accelerator for memory-efficient graph convolutional neural networks

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:226a695b46963fcfb231ff74a64e9645d8680ea3bbd79567436f81f33b7ccf94

Observation ee1a08a2-4a0f-446b-bceb-801e6129435d · outbound

This paper cites An extended compression format for the optimization of sparse matrix-vector multiplication.IEEE Trans- actions on Parallel and Distributed Systems , 24(10):1930–1940, October 2013.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended An extended compression format for the optimization of sparse matrix-vector multiplication.IEEE Trans- actions on Parallel and Distributed Systems , 24(10):1930–1940, October 2013

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:68b1aee8334e29452c7107a6ce8e057dd118b46df0eed54e4a439e74c97ec693

Observation 76a167d5-cc2c-45f5-bcfa-bb4e648c50f5 · outbound

This paper cites Tensorfloat-32 in the a100 gpu accelerates ai training, hpc up to 20x.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Tensorfloat-32 in the a100 gpu accelerates ai training, hpc up to 20x

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:55f429fd639ada6de5c665d8df8ff4abd2a9b88652047cc9e7701483acd7ee3a

Observation 1d3ae069-df77-43e2-99ae-2a4a9b7fbd28 · outbound

This paper cites Datasets for benchmarking floating-point compressors, 2020.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Datasets for benchmarking floating-point compressors, 2020

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:a5f8f2b457cbc3a86225ccc507b5d7323d80d8897aa8439722394e4734239a50

Observation 4fa23bbb-2e43-410b-9024-6e9e46c9b815 · outbound

This paper cites ndzip-gpu: ef- ficient lossless compression of scientific floating-point data on gpus.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended ndzip-gpu: ef- ficient lossless compression of scientific floating-point data on gpus

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:b8cdffa83ad635a81ffa0f8f58a2e446af33e53018d2e02f254ad58ca13d9a3a

Observation fd6c6375-e65d-4a89-b5dc-47e929ce8943 · outbound

This paper cites Webb, Xin Wang, Marcel Nassar, Arjun K.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Webb, Xin Wang, Marcel Nassar, Arjun K

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:0b23325fc8e06cbad24ce14021a8590994692276c7bf96a15a7ec3c6631c64f5

Observation 06166886-e07d-4b9f-b8b1-41d2e548f4e2 · outbound

This paper cites Splitrpc: A Control + Data path splitting rpc stack for ml inference serving.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Splitrpc: A Control + Data path splitting rpc stack for ml inference serving

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:c3fbeecbb87cd98256348120605df8ef5288d294343df3d0e5d6498c132d0d64

Observation 243aeb13-81db-4714-aa62-4c1739260fe1 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Efficient memory management for large language model serving with pagedattention

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:ea5542208c53297052b665071e2eb57d26113939f29cddf3c2456af8c163235d

Observation 90515bfd-b669-44db-8b7c-527daed8ff94 · outbound

This paper cites InfiniGen: Efficient generative inference of large language models with dynamic KV cache management.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended InfiniGen: Efficient generative inference of large language models with dynamic KV cache management

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:db111c5dbd7b8769f861fa9062bb6e59c8a1b7f246929287779a4b2f30f8338b

Observation 7eaa13f5-e8f6-4bf1-a4f2-e39cd06a4583 · outbound

This paper cites Naughton, and Jignesh M.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Naughton, and Jignesh M

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:25d8f76f2c52474a684d1f16bc9ea0f754655e97929bf30ed310658de0ae4f87

Observation b6dc5995-db65-4a6b-a31b-5787a22441c2 · outbound

This paper cites THC: Accelerating Distributed Deep Learning Using Tensor Homomorphic Compression.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended THC: Accelerating Distributed Deep Learning Using Tensor Homomorphic Compression

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:5967ca7311e0d4a845ae8660e60063de22125087a843d4dd58408c9e8630cd27

Observation 47b697ea-fe90-4f34-bb49-121875260143 · outbound

This paper cites Colossal-ai: A unified deep learning system for large-scale parallel training.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Colossal-ai: A unified deep learning system for large-scale parallel training

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:109a2acf9cca5b8d10091221c50b0bdfb128d1fa989484ac752d195c8735724d

Observation c1fcb082-3cb3-438b-bc1b-015b8b5c7786 · outbound

This paper cites an unresolved cited work.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:ca8e070f091dc814a9bddff659ec626d75989c7df339741cfc6bb3a0b191f410

Observation 64a46fc8-ccb1-425b-9cdd-73fd58363cc9 · outbound

This paper cites Recoil: Parallel rans decoding with decoder-adaptive scalability.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Recoil: Parallel rans decoding with decoder-adaptive scalability

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:7189bfba249de7dbfc475f0a5a7b6e0cb28827bb8f13c8dc3bff51c8003ad9c4

Observation e858a7ae-6d0c-4d64-ae23-29f83b8b43ca · outbound

This paper cites Using cuda warp-level primitives.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Using cuda warp-level primitives

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:154d37370d943bbf9c3f0d477470131d636f3ebca9834ecf67d8d2f6840c4705

Observation a6cb9f79-afc7-4fb5-a368-35b07155cdde · outbound

This paper cites Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:c5c2b543ac87d35e14f3302b791d220b52a5efda3f594ebe1c53abc4d8f50bdf

Observation ccb87448-fd90-4747-a8ab-632fcd89b0f9 · outbound

This paper cites Pa- graph: Scaling gnn training on large graphs via computation-aware caching.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Pa- graph: Scaling gnn training on large graphs via computation-aware caching

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:40308a553f7d3e6ed6cc2f9d8774309230998f332751d8f2a28ecec762e0f9a9

Observation 03c8a26e-a4e1-4b31-b86e-e6d0ba77c44b · outbound

This paper cites Indigo: Gnn-based inductive knowledge graph completion using pair-wise encoding.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Indigo: Gnn-based inductive knowledge graph completion using pair-wise encoding

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:e7e0273e8a3e467552cda92bbbdf8fb5ceef04534a33bc5f75257620ad63bc33

Observation 5ba372c9-549c-4169-8b4e-ac1574d381ad · outbound

This paper cites BGL: GPU-Efficient GNN training by optimizing graph data I/O and preprocessing.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended BGL: GPU-Efficient GNN training by optimizing graph data I/O and preprocessing

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:37d49b8f0facfa3372a35732beafa8dce96e7577e315ff99b79f9f731452cf42

Observation 7acec39a-9a2e-47b7-8801-4e1dc12cf63c · outbound

This paper cites Pick and choose: A gnn-based imbalanced learning approach for fraud detection.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Pick and choose: A gnn-based imbalanced learning approach for fraud detection

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:bbc3fb86f05409b07c647c866a451313b45eb511b2a59061889c41ef1a032d92

Observation 388b03f8-fa5f-4a2f-bf1b-13e7391b0bd3 · outbound

This paper cites Cachegen: Kv cache compression and streaming for fast large language model serving.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Cachegen: Kv cache compression and streaming for fast large language model serving

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:969e6f242902cd6a15abac86233ce5d98f73985212694e75290edd1dfec01205

Observation 30868c71-8ae6-494c-8488-2327e840b407 · outbound

This paper cites Dvc: An end-to-end deep video compression framework.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Dvc: An end-to-end deep video compression framework

Reference 59

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:d16a41ef00565e3c411e08e74bd448c59c83a96f5c3c2d2d08a55769bfb4af12

Observation f31eabe6-84e5-4a95-a311-e0ccba012da0 · outbound

This paper cites Eliminating data processing bottlenecks in gnn training over large graphs via two-level feature compression.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Eliminating data processing bottlenecks in gnn training over large graphs via two-level feature compression

Reference 60

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:68767258a320d90d9f39217824d9adce149f726c56bf1f15c6918d7aa399a103

Observation ff28a625-a2f4-4629-a8ad-a492ad724d7e · outbound

This paper cites BiFeat: Supercharge GNN Training via Graph Feature Quantization.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended BiFeat: Supercharge GNN Training via Graph Feature Quantization

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:32:46.922930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:242bfb81404406c633b9041b05fcb941487ce650e40e4a5287f4faa87238df5f

Observation 53c41487-f68c-4495-8354-554d37f3d024 · outbound

This paper cites Emogi: efficient memory- access for out-of-memory graph-traversal in gpus.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Emogi: efficient memory- access for out-of-memory graph-traversal in gpus

Reference 62

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:8fadd03cd8e6fcc7c33a4dda0af27f623b460e22acb370bf6de0be39c83cd381

Observation 434671fd-1b3a-4c62-b9fb-e90ecab31030 · outbound

This paper cites an unresolved cited work.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:32d6c6f59d807f9fd8adb2ae25fe2af1b1bb8a9eca2be537f9812b41c4a27ebe

Observation a40dea05-0faf-480b-855c-9f5d5ae8d850 · outbound

This paper cites an unresolved cited work.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:56ab8798e1ff519e81035e696c20c4bfcb9c02c090fc37261eda89e079c53e57

Observation 9ddab374-b9a9-4f7e-b610-fe31c9c53caf · outbound

This paper cites Query-driven active surveying for collective classification.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Query-driven active surveying for collective classification

Reference 65

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:2f93c5f34feb722bc252dc4415fa94180bfbf35a93f56ad3361b671059bc992c

Observation a783d1e3-caea-43f3-ad0a-34df0e53354b · outbound

This paper cites Patel, Yao Zhang, Jason Mak, Andrew Davidson, and John D.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Patel, Yao Zhang, Jason Mak, Andrew Davidson, and John D

Reference 66

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:33988a75e5f48a250880f7befbd9118cc97a9f1d3a0a8beb75c66894d1931338

Observation d23c0cfd-cc64-49aa-93f4-70d90dca7ff1 · outbound

This paper cites Gpu-initiated on-demand high-throughput storage access in the BaM system architecture.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Gpu-initiated on-demand high-throughput storage access in the BaM system architecture

Reference 67

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:bc24db8708f9f491b2bf587b9693f21464e500caa6d560ccb71c56ba7b56b260

Observation 875c4bd6-3e20-4484-908b-de79aaf456c4 · outbound

This paper cites Real-time adaptive image com- pression.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Real-time adaptive image com- pression

Reference 68

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:2ccb7c95be55f4059b7e1614dcbfeb9cda9e01285f12ced2ba37f4e4d095c891

Observation 45c77b34-affa-4749-a7f1-afb9083063b4 · outbound

This paper cites Faster across the pcie bus: a gpu library for lightweight decompression: including support for patched compression schemes.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Faster across the pcie bus: a gpu library for lightweight decompression: including support for patched compression schemes

Reference 69

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:d61ee0f7996193a7bc2a01ee464c37b48af45b22cec7e76637294ef2b063ca39

Observation 7715261c-b98c-4ec9-ad5b-28f4e732f1c0 · outbound

This paper cites Collective classification in network data.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Collective classification in network data

Reference 70

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:aee4ff76070898f0cb60810d4418efdac3485c485495afe877012000aaae85a3

Observation 40d51d3c-ae4f-411f-b3ba-80e68ee7f4af · outbound

This paper cites Scalable graph neural network training: The case for sampling.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Scalable graph neural network training: The case for sampling

Reference 71

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:776e7c720333b11c75b8f794c64fc6374f46bc026b40ce6d576fbc30694a28c4

Observation 31c0e8c1-c916-442d-bc98-09c0c76c1955 · outbound

This paper cites Yogatama, Xiangyao Yu, and Samuel Madden.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Yogatama, Xiangyao Yu, and Samuel Madden

Reference 72

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:8ad06b1e64d85073a70d4e04a92201b0e8a8a52898bdb66f23ca3eb961715f16

Observation 9bc214dc-efb8-431b-8baf-4a2289901aea · outbound

This paper cites FlexGen: High-throughput generative inference of large language mod- els with a single GPU.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended FlexGen: High-throughput generative inference of large language mod- els with a single GPU

Reference 73

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:c6f31c75bac9313801fb02345538cf0ae84db0b4dc991adb08e5c1382049efd4

Observation dff30bca-721a-445c-bbfc-4963b008c41c · outbound

This paper cites Ugache: A unified gpu cache for embedding-based deep learning.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Ugache: A unified gpu cache for embedding-based deep learning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:0c1b6a46bd5e9447b89ac1820225a46c66903759bfdc4e421beb2720f0e7c745

Observation 649b7f85-b853-4efb-b489-f0fee07b7797 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Roformer: Enhanced transformer with rotary position embedding

Reference 75

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:e93d84b83d3911626d286e28ca923e81c07813bbb902895621764efeec803231

Observation bc495989-c00e-4e7a-8b4b-75abe4291eb5 · outbound

This paper cites Legion: Automatically pushing the envelope of Multi-GPU system for Billion- Scale GNN training.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Legion: Automatically pushing the envelope of Multi-GPU system for Billion- Scale GNN training

Reference 76

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:2ff08c0f2891cfeeb3785dc00b2b0e85d6f4b4ab71344061812b1d96b655a246

Observation c82750b2-14ce-4140-8579-06f10d847d15 · outbound

This paper cites Controlling data move- ment to boost performance on the nvidia ampere architec- ture.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Controlling data move- ment to boost performance on the nvidia ampere architec- ture

Reference 77

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:74c9664525ce51b3d5ca34d1ac3cc65b10a8b8535e60d7252d6a4c4d35190814

Observation 8789fc65-98b0-4656-9cf7-eea646aaa210 · outbound

This paper cites an unresolved cited work.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Unresolved cited work

Reference 78

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:c5a52961a0a74a8150d1f6ab69d42c03162e95f69fa7a6296e90afea434043a2

Observation f0c47773-57e0-413e-b0f5-0bd4c64c23c7 · outbound

This paper cites Attention is all you need.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Attention is all you need

Reference 79

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:70c4a2a90f64a81d364a0f4c600c13ceaf232216631e2450095858a01d235308

Observation 55836ff8-7788-4019-bbd9-e5dad7699ef4 · outbound

This paper cites Mariusgnn: Resource-efficient out-of-core training of graph neural networks.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Mariusgnn: Resource-efficient out-of-core training of graph neural networks

Reference 80

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:3230d0c0d47825e39238dc3dd6deef29195c260ba0039a41eee0a236e2eed258

Observation 4931efc0-0154-4b32-ac72-10a779bb8be0 · outbound

This paper cites ZeRO++: Extremely Efficient Collective Communication for Large Model Training.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended ZeRO++: Extremely Efficient Collective Communication for Large Model Training

Reference 81

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:18276e2a21e4983171ddc9e12b92c533ed46e80784f9bb7196d52a71831041cc

Observation 00eb0ab9-8e4e-4cf3-a011-439056655ce0 · outbound

This paper cites Deep Graph Library: A Graph-Centric, Highly-Performant Package for Graph Neural Networks.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Deep Graph Library: A Graph-Centric, Highly-Performant Package for Graph Neural Networks

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-06-28T23:32:46.920362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:e4fc1535bb7a316f7d20fa0b7181761a79fce14590c18082265e6d714d745195

Observation da586fe3-1ab0-4b7a-a99d-78cd59d14e5f · outbound

This paper cites Bfloat16: The secret to high per- formance on cloud tpus.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Bfloat16: The secret to high per- formance on cloud tpus

Reference 83

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:29c94a3f96c2ce1fd3ccb404e9de5f0a208473a844a6d019303a71833b6820aa

Observation 161265b2-6c2f-4a08-9bf7-aab0bd0725e6 · outbound

This paper cites A gpu-specialized inference parameter server for large-scale deep recommendation models.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended A gpu-specialized inference parameter server for large-scale deep recommendation models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:85d4d3cfa60bce801971a287770c7d607a27ae4859f8797b13b865e630fbebef

Observation edbd7f26-0ea9-4e15-9a5e-ff92b4e015c1 · outbound

This paper cites Massively parallel inverse block-sorting transforms for bzip2 decompression on gpus.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Massively parallel inverse block-sorting transforms for bzip2 decompression on gpus

Reference 85

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:1ad8cc5963968e989da0d1bc2ed4f7f915b1f95a8a8f169d9d2ca864883c8838

Observation 8bb53cea-deb6-4830-9a7e-aa3763d5216c · outbound

This paper cites MLaaS in the wild: Workload analysis and scheduling in Large-Scale heterogeneous GPU clusters.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended MLaaS in the wild: Workload analysis and scheduling in Large-Scale heterogeneous GPU clusters

Reference 86

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:069cfe0bd56fb74bc682d06bf7f6c9610f0a76f0d880b18082925cb07dc877c8

Observation 964e1d95-68d8-493f-b387-4e513ce72c62 · outbound

This paper cites Widrow, I.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Widrow, I

Reference 87

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:37a1a0508a5bd7f4fef9c431f4bd645c62eab48c45d6827fd57595c2e5044b53

Observation b088b425-1ec4-4f1a-8d78-9d937a094051 · outbound

This paper cites Abdelmoniem, Aritra Dutta, El Houcine Bergou, Konstantinos Karatsenidis, Marco Canini, and Panos Kalnis.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Abdelmoniem, Aritra Dutta, El Houcine Bergou, Konstantinos Karatsenidis, Marco Canini, and Panos Kalnis

Reference 88

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:a56fcb0c8d442b1030788fabbb7bc82eeb5366a35a182adfb34bb404c000aaea

Observation 5ec46941-92f6-4096-842c-4c0cab3a09d4 · outbound

This paper cites Mpc: A massively parallel compression algorithm for scientific data.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Mpc: A massively parallel compression algorithm for scientific data

Reference 89

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:5e014d69df36de991d6667c0fcd745ea3322486a504005c4656456eeaf4b5bd7

Observation 1917737e-b805-4d3b-aadf-ea01c6e07444 · outbound

This paper cites Gnnlab: a factored system for sample-based gnn training over gpus.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Gnnlab: a factored system for sample-based gnn training over gpus

Reference 90

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:1c5b56f6ae2fdfd4619af79aba6a242a1e2d222a73aa026169a5d41dbf355c68

Observation 1a6698af-23dc-441c-afae-4f51cc6da5d4 · outbound

This paper cites Lossy image compression with conditional diffusion models.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Lossy image compression with conditional diffusion models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:84804e8166b23f7a61882bd259685fdefd03936a83ec560a7cafcc6384ae149c

Observation dca569c2-248b-4097-8a2f-653f38319afb · outbound

This paper cites Spargnn: Efficient joint feature-model sparsity exploitation in graph neural network acceleration.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Spargnn: Efficient joint feature-model sparsity exploitation in graph neural network acceleration

Reference 92

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:74bc7596167d2a387c1eedc3c15928c621afe69fb895a114570bbf98f1f6e0be

Observation eade4f5c-45c0-40b5-9621-65cb4722a893 · outbound

This paper cites Hamilton, and Jure Leskovec.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Hamilton, and Jure Leskovec

Reference 93

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:53b6b9a338ded74b355df293e0093b749ee852fdd1a324fdaae74100eb2fe205

Observation 66ca2141-80e0-472f-848b-0594027d9ac1 · outbound

This paper cites SGCN: Exploiting Compressed-Sparse Features in Deep Graph Convolutional Network Accelerators.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended SGCN: Exploiting Compressed-Sparse Features in Deep Graph Convolutional Network Accelerators

Reference 94

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:8ad51f2e4f96cc8d71f86face7847a0ab0f98e48e642fcca531dcc9f23dce097

Observation e9c02e50-0c37-451d-8f89-9bd62460bc61 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended OPT: Open Pre-trained Transformer Language Models

Reference 95

Resolution
verified exact
local_arxiv, observed 2026-06-28T23:32:46.931201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:9f7c2666e982873e1b625d14de648ebd1653a887ff24093449ce7a6f338f5bc1

Observation b2dfe8ad-33e0-41a5-a700-4c19e1ebb98f · outbound

This paper cites 70% size, 100% accuracy: Lossless LLM compression for efficient GPU inference via dynamic- length float (DFloat11).

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended 70% size, 100% accuracy: Lossless LLM compression for efficient GPU inference via dynamic- length float (DFloat11)

Reference 96

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:718d4b486e34a7de870038bf135c804b0ca42d5f12a14620c1ce54e23bbe8166

Observation 942f0507-fa9a-4a9a-9730-f7b5dc42558d · outbound

This paper cites Ducati: A dual- cache training system for graph neural networks on giant graphs with the gpu.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Ducati: A dual- cache training system for graph neural networks on giant graphs with the gpu

Reference 97

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:7da1c4bcace423341f4586b550bb6d61e5c90742e33776c96414ffa9470d0253

Observation 08442e00-035e-479c-a085-b7b29f8f283e · outbound

This paper cites H2o: heavy-hitter oracle for efficient generative inference of large language models.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended H2o: heavy-hitter oracle for efficient generative inference of large language models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:7df0170ea5c00e88edc16a6aa2894b9de18be4a80ba21fd0bcb75b9b791c9f57

Observation 54eaea14-c732-4b68-81c6-b33253925b5f · outbound

This paper cites Silod: A co-design of caching and scheduling for deep learning clusters.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Silod: A co-design of caching and scheduling for deep learning clusters

Reference 99

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:a9c67093278b8e7682f0d3db7848f33d5ecc9292ac4c331e05db4877a6643fc0

Observation 95e2f71d-3e78-405c-bfad-436a4651580c · outbound

This paper cites Atom: Low-bit quantization for efficient and accurate llm serving.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Atom: Low-bit quantization for efficient and accurate llm serving

Reference 100

Resolution
unresolved
no resolver link, observed 2026-06-28T23:31:27.783419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:e8d8da8ca3666eece3edf3b306ccbbf78abd45ce462a5d76ce0a94774f985c58

Observation d4739a5d-6675-43b4-b467-68655052e63a · outbound

This paper cites ${CONDA_PREFIX}.

Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended ${CONDA_PREFIX}

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:32:46.925804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T23:31:27.783419Z digest=sha256:d527735879ebbd65d992bd62ddd6349601c717cca245934f3f6fbcece40b8d50

Pith citing papers

No inbound Pith citation observations are available.