Pith. sign in

Paper Citation Record · LEDGER

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding

As of 9 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2505.18758.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18758 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:32:20.014194Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy51
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 40af015c-b43e-4cbc-81e1-a8cdcc552ddf · outbound

This paper cites Croci, Bo Li, Pashmina Cameron, Martin Jaggi, Dan Alistarh, Torsten Hoefler, and James Hensman.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Croci, Bo Li, Pashmina Cameron, Martin Jaggi, Dan Alistarh, Torsten Hoefler, and James Hensman

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:28.797709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:16.190894Z digest=sha256:d913ec6932f2f7b2a16b60407d6f0882471d38e10d7ada0fa1394d9150657d07

Observation c51b9444-a97c-4766-a862-6c45b7db9d9c · outbound

This paper cites GPTVQ: The blessing of dimensionality for LLM quantization.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding GPTVQ: The blessing of dimensionality for LLM quantization

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:28.654010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:16.241300Z digest=sha256:e210286812be77d86bfd48e183f8fdc1117a23f26215c9447321309e953d73b0

Observation e4eba4ad-b0b8-47bc-a21b-5627f02ccf0b · outbound

This paper cites ONNX: Open neural network exchange, 2019.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding ONNX: Open neural network exchange, 2019

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:28.545358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:16.358834Z digest=sha256:0b44feea721347b5148f199a9e1bf8addebd4ed8c2dbd3512838f7d9789cddd1

Observation b84cb3ec-0010-460b-a007-edfffe9f8604 · outbound

This paper cites Understanding Entropy Coding With Asymmetric Numeral Systems (ANS): a Statistician's Perspective.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Understanding Entropy Coding With Asymmetric Numeral Systems (ANS): a Statistician's Perspective

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:16.426671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:32:16.426671Z digest=sha256:20cc50d353264ea685f5f603cda4d1d3c5d574ee412dd8d2fee976ed36aaa970

Observation c74d72d4-3de6-4958-b936-3e752748c3cb · outbound

This paper cites Bronstein, and Avi Mendelson.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Bronstein, and Avi Mendelson

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:28.368746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:16.511794Z digest=sha256:1793d83f57795faa1d8d93d51bc56a8cca373c6f51ab61373a6f9a4a91c6a4e8

Observation dc197d84-ff2f-4ee5-8f6e-447a5cd3dde0 · outbound

This paper cites NNCodec: An Open Source Software Implementation of the Neural Network Coding ISO/IEC Standard.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding NNCodec: An Open Source Software Implementation of the Neural Network Coding ISO/IEC Standard

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:28.181108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:16.583499Z digest=sha256:847268f51bdcfe46c68ae2f338dd89af3e945b8da9e4769207b56ca3bdec4273

Observation fcc570e4-20a8-4a20-ac16-835174b597f0 · outbound

This paper cites EfficientQAT: Efficient Quantization-Aware Training for Large Language Models, October 2024.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding EfficientQAT: Efficient Quantization-Aware Training for Large Language Models, October 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:27.999615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:16.640320Z digest=sha256:38f82632f927734845e0fa3994ca4d4b4aaa017ae620f8a2a9e1d16a6c401866

Observation 98f4c2c9-2ba9-42fb-823e-77ad995f07db · outbound

This paper cites Bronstein, and Avi Mendelson.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Bronstein, and Avi Mendelson

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:27.831968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:16.729492Z digest=sha256:af337e8525babe5b405d6709b4d9132e281c41d18ea316cf6834ffb6342bbfb9

Observation ba0676f6-8c18-4488-aa96-6dc25e01e812 · outbound

This paper cites Universal Deep Neural Network Compres- sion.IEEE Journal of Selected Topics in Signal Processing, 14(4):715–726, May 2020.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Universal Deep Neural Network Compres- sion.IEEE Journal of Selected Topics in Signal Processing, 14(4):715–726, May 2020

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:27.628117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:16.835954Z digest=sha256:67f647c64d334d68f22c01a061ea01e671f9daa42d1ee6be75d1fc11e85960e3

Observation 2c2e90cb-97c4-455d-b258-6aac7cd8bb31 · outbound

This paper cites Imagenet: A large- scale hierarchical image database.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Imagenet: A large- scale hierarchical image database

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:16.916800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:32:16.916800Z digest=sha256:18c3f019ceaf552d8c1d9a67c95ef83f73e788281f1f99598782c73669b7ac8d

Observation 2c31a623-0c1b-40ab-8873-f52a6e545d96 · outbound

This paper cites GPT3.int8(): 8-bit Matrix Multiplication for Transformers at Scale.Neural Information Processing Systems (NeurIPS), January 2022.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding GPT3.int8(): 8-bit Matrix Multiplication for Transformers at Scale.Neural Information Processing Systems (NeurIPS), January 2022

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:27.442812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:16.971706Z digest=sha256:3b693c17ee093e93297def3811ce113a6b2a2977fec3ca19d050ffedd66d490c

Observation b61ccbcd-d1ed-45a4-96b9-10487a13ebdc · outbound

This paper cites QLoRA: Efficient Finetuning of Quantized LLMs, May 2023.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding QLoRA: Efficient Finetuning of Quantized LLMs, May 2023

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:27.226382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:17.038296Z digest=sha256:ed41388ac423d4904913900e5af3571ccea0af8645713aea2b133eae3974b387

Observation 08b72a21-f583-4eba-b37f-99687fdc40e1 · outbound

This paper cites Svirschevski, Vage Egiazarian, Denis Kuznedelev, Elias Frantar, Saleh Ashkboos, Alexander Borzunov, Torsten Hoefler, and Dan Alistarh.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Svirschevski, Vage Egiazarian, Denis Kuznedelev, Elias Frantar, Saleh Ashkboos, Alexander Borzunov, Torsten Hoefler, and Dan Alistarh

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:27.050348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:17.101095Z digest=sha256:1cf5cdcedc3edef47605646594c01e572b3d2a710d78245446704d03aa39338c

Observation 7cb383c7-385d-49cc-8c08-2913a1704ae8 · outbound

This paper cites The case for 4-bit precision: K-bit Inference Scaling Laws.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding The case for 4-bit precision: K-bit Inference Scaling Laws

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:26.869093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:17.167832Z digest=sha256:83aed27e1c98237712cdcf512258e7fce15075f7b560d4aa299ad800abca37e9

Observation efdb50bb-466a-4a36-b625-99f1d7e75c22 · outbound

This paper cites STBLLM: Breaking the 1-Bit Barrier with Structured Binary LLMs, August 2024.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding STBLLM: Breaking the 1-Bit Barrier with Structured Binary LLMs, August 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:26.747986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:17.246176Z digest=sha256:bb149e61311b4ad546d31ff1688efa3444cf67b2aa59eff749b0e1c554923912

Observation 50f7a082-196c-4b17-b9ab-92101700f1d1 · outbound

This paper cites The use of asymmetric numeral systems as an accurate replacement for huffman coding.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding The use of asymmetric numeral systems as an accurate replacement for huffman coding

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:26.630506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:17.325605Z digest=sha256:f8dd51273e2f18611e449b2fa1982f174009c8aec8a13612aa51650a7bea7138

Observation 966c4542-3a45-410b-93a2-27b59bc06d25 · outbound

This paper cites Extreme Compression of Large Language Models via Additive Quantization, September 2024.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Extreme Compression of Large Language Models via Additive Quantization, September 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:26.543235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:17.399995Z digest=sha256:d12a45f0cb320ce38b8a20e5c741cf8a7d528840fba0373ddf69b570fb3120e5

Observation a7835023-4e6f-4a04-b6a4-30af7208a150 · outbound

This paper cites Optimal Brain Compression: A Framework for Accurate Post- Training Quantization and Pruning.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Optimal Brain Compression: A Framework for Accurate Post- Training Quantization and Pruning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:26.389323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:17.461915Z digest=sha256:7d005b57ae12368dc909c6f1669463e09ad4cdb3ba896fd59bd3f80718bb7545

Observation 97d2e906-e906-4fd3-a48c-fb5c6128f318 · outbound

This paper cites SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot, March 2023.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot, March 2023

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:26.279903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:17.553273Z digest=sha256:516bdf49fb9218dfacd9cd38bc9ca367644509c77ac128bd545ace936a4e81f4

Observation 11561f6f-8600-4d3d-bc06-6a2290aac634 · outbound

This paper cites OPTQ: Accurate quan- tization for generative pre-trained transformers.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding OPTQ: Accurate quan- tization for generative pre-trained transformers

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:26.200374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:17.633391Z digest=sha256:a7da19afb5445795b7a10999509b5ff4c1654cc808405099f5b60f5dcaec08a2

Observation a0f1d2a1-98dd-4ac6-8a23-f0827d426d36 · outbound

This paper cites Compression Scaling Laws:Unifying Sparsity and Quantization, February 2025.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Compression Scaling Laws:Unifying Sparsity and Quantization, February 2025

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:26.118695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:17.690251Z digest=sha256:fa9bab734e2517260c7ca5514847c7f969ccdaf643dd08cf0ec62bad6b4e6713

Observation 3805198f-2f45-4af3-a9e4-4bd38bd0eb37 · outbound

This paper cites MiniLLM: Knowledge Distillation of Large Language Models.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding MiniLLM: Knowledge Distillation of Large Language Models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:25.979526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:17.769125Z digest=sha256:c463f8cb1c3823a5718188a890928adf72e00b7031d8301048d49b9ef6a93f99

Observation e19cf533-d423-479f-9af8-f067ea7af6d8 · outbound

This paper cites an unresolved cited work.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:25.863479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:17.836993Z digest=sha256:b235f9e7baa3617cf09a0b394a0f3ee9e8761f7494b628b70167ec4b21abe3c1

Observation 29015da6-6314-4e3e-862d-e23f06486f50 · outbound

This paper cites NeuZip: Memory-Efficient Training and Inference with Dynamic Compression of Neural Networks, October 2024.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding NeuZip: Memory-Efficient Training and Inference with Dynamic Compression of Neural Networks, October 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:25.743356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:17.905015Z digest=sha256:5bd496c4d792ecc93312718672d2d734c9acb7da7e51dc17703e11aa082c9094

Observation b56910ea-1476-4b58-98fd-403b3e979413 · outbound

This paper cites Hassibi, D.G.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Hassibi, D.G

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:25.630947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:17.992323Z digest=sha256:482bcfe791a5e86705130e8483ff3d6b75de579fe6f0d12e8c5aa4b0101b7910

Observation c06c72ef-c27c-4b0d-9d36-b769f6272758 · outbound

This paper cites Deep Residual Learning for Image Recognition.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Deep Residual Learning for Image Recognition

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:25.469500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:18.053367Z digest=sha256:6aeca095ac705de4444d73dc9124c08749791367726450e1f2a86db4015533ce

Observation f52202bb-ac7e-4232-94c8-eca58327caa6 · outbound

This paper cites Model Compression in Practice: Lessons Learned from Practitioners Creating On-device Machine Learning Experi- ences.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Model Compression in Practice: Lessons Learned from Practitioners Creating On-device Machine Learning Experi- ences

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:25.359504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:18.110251Z digest=sha256:123b648f320f8353192f1b31da2978f7f74d312ddad9340353e3804734c27c89

Observation 04780eec-bda4-4ae5-a8b2-4e05c0ba1871 · outbound

This paper cites Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:25.189720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:18.187867Z digest=sha256:22513645d98e42b0ac375e15c75ac4997835e1fe68bbbc696ee5ee70dd3be57e

Observation 07cebcb8-7ea8-4c33-a4ed-fc6ea06072be · outbound

This paper cites Le, and Hartwig Adam.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Le, and Hartwig Adam

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:25.077667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:18.289283Z digest=sha256:20167162fa50e2b7c352a866499bafde090b5535fb8a3f5b492982a9da462554

Observation 9a27fcbb-35a0-41ce-8a89-2b75d6c075cc · outbound

This paper cites Accurate Post Training Quantization With Small Calibration Sets.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Accurate Post Training Quantization With Small Calibration Sets

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:24.916723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:18.361634Z digest=sha256:cfdc809dbaddcc1bd2667e44c82ab3389c5e4c512fbbf584295420d998760e12

Observation cb6914b4-e72e-44da-912b-88c05c479446 · outbound

This paper cites Mahoney, and Kurt Keutzer.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Mahoney, and Kurt Keutzer

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:24.770296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:18.411007Z digest=sha256:00391875c4458041d97ff802ea5c8573bc021d2921b619d03315ece79900e063

Observation 7b2c153c-0e14-4d83-8602-23851cdba37d · outbound

This paper cites Aksu, Miska M.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Aksu, Miska M

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:24.589078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:18.463769Z digest=sha256:11820fe19d6ddbd21a9672b5b33262e1b162144517cceaa9eb0a50bf930c0feb

Observation 678cbc9e-cf74-4e3f-9650-dce56d55ccca · outbound

This paper cites Adaptive weight compression for memory-efficient neural networks.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Adaptive weight compression for memory-efficient neural networks

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:24.467567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:18.535752Z digest=sha256:e2fb2f6e43e063c240dfcdb798a805e7c35365439097b06ddf9f1dc3dacde8f9

Observation 0c3747bb-9ab5-48b5-95a3-dcb1e6d96250 · outbound

This paper cites Cifar-10 (canadian institute for advanced research).

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Cifar-10 (canadian institute for advanced research)

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:18.611847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:32:18.611847Z digest=sha256:613456e96beb093708704b709078b349152e7f844d8282f78b24b9d015765031

Observation fedf0c03-7b2e-4efa-b589-abb654016c75 · outbound

This paper cites Energy-Efficient Model Compression and Splitting for Collaborative Inference Over Time-Varying Channels.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Energy-Efficient Model Compression and Splitting for Collaborative Inference Over Time-Varying Channels

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:24.309768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:18.661478Z digest=sha256:ccce9c1e9b18a8572436cbf2e668a78d4c3d5d08edc6f8d64ae0dfdeff8383b1

Observation f19237d7-95f6-4def-ad28-268ca16b2138 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Gonzalez, Hao Zhang, and Ion Stoica

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:18.725195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:32:18.725195Z digest=sha256:29fd962772c9b714859715ab698f370794cd5e119c784900d917bbca3b231156

Observation 49b08133-2996-49c4-9b5a-a3a463600a6f · outbound

This paper cites Memory Efficient Optimizers with 4-bit States.Advances in Neural Information Processing Systems, 36:15136–15171, December 2023.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Memory Efficient Optimizers with 4-bit States.Advances in Neural Information Processing Systems, 36:15136–15171, December 2023

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:24.193718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:18.777724Z digest=sha256:be651b9d40a36bbc3a985747ff555a2e695a13ba3c2afa605be23ea6ab061b0d

Observation ea0ff6d2-256e-4297-bb5e-61e19044e52c · outbound

This paper cites PENNI: Pruned Kernel Sharing for Efficient CNN Inference.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding PENNI: Pruned Kernel Sharing for Efficient CNN Inference

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:24.038562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:18.879989Z digest=sha256:c27a308655275c465a79be437cc1ab1c2b7c13d06e259703642f4cba13289094

Observation acbd46d8-c6fd-4c0b-9d62-f3bfd2794fb1 · outbound

This paper cites AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration, April 2024.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration, April 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:23.928089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:18.942144Z digest=sha256:aa93949fde50e0ee40f78fb10c82b5dfe1bc7575d72f80f9095907296993d9e7

Observation 47fc9a82-5229-4550-931c-928db5094e91 · outbound

This paper cites Cambridge university press, 2003.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Cambridge university press, 2003

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:23.749967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:19.009752Z digest=sha256:43e9336fb0b9eb230454bb9761a7852314d3bf9276204218e5a1119f77f11a1e

Observation 2812bc00-7f8b-45f5-a989-3b49903e3ae6 · outbound

This paper cites Range encoding: an algorithm for removing redundancy from a digitised message.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Range encoding: an algorithm for removing redundancy from a digitised message

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:23.570244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:19.078862Z digest=sha256:bc2d1b51d938d46e576e3aeaa73364e2ab7d58214a4612c1389cd2925c619e5a

Observation b6c3659e-7fcb-4b0e-baf2-a7b7ec128f50 · outbound

This paper cites Up or Down? Adaptive Rounding for Post-Training Quantization.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Up or Down? Adaptive Rounding for Post-Training Quantization

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:23.422163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:19.142752Z digest=sha256:9774c51bf7eb2cbe841092b623c8abdb99ac7ff1bfea7199d73635862c14fb14

Observation 82cadcdb-dc92-4d6c-a9ed-4b73d51897fe · outbound

This paper cites A White Paper on Neural Network Quantization, June 2021.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding A White Paper on Neural Network Quantization, June 2021

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:23.259329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:19.217951Z digest=sha256:23b297b1d07888a9d455e86efffaf0d9a2bacd20416fc4c10fff812344247e28

Observation 83cc05eb-0e6b-4970-9389-f3788a22d706 · outbound

This paper cites Kübler, Jiaji Huang, Matthäus Kleindessner, Jun Huan, V olkan Cevher, Yida Wang, and George Karypis.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Kübler, Jiaji Huang, Matthäus Kleindessner, Jun Huan, V olkan Cevher, Yida Wang, and George Karypis

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:23.089866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:19.307019Z digest=sha256:d398c50846d3c4ed965c952435aeb805737ef537a26286a4c1436cf782200bd2

Observation 0a6ff35c-4c2c-4566-8937-aa287d1d91a3 · outbound

This paper cites PhD thesis, Stanford University CA, 1976.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding PhD thesis, Stanford University CA, 1976

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:22.904461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:19.342907Z digest=sha256:8ccfa815b7da39dbb70c2e1d2e9f3e1d22f2ab247191c739e4d1f56fc2559ee5

Observation 1bd47c1e-0df3-47ac-8a80-3fd74b0e4fb3 · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library, December 2019.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding PyTorch: An Imperative Style, High-Performance Deep Learning Library, December 2019

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:22.735567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:19.398920Z digest=sha256:d633e7a6fce0f74a089b4afcfcaf957b121e16055027a09c5df625a259546755

Observation 8f6cb52b-a8e6-426b-84ce-e94e6f2d0d79 · outbound

This paper cites Accurate LoRA-Finetuning Quantization of LLMs via Information Retention, May 2024.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Accurate LoRA-Finetuning Quantization of LLMs via Information Retention, May 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:22.532664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:19.468784Z digest=sha256:25cc8bd3ed9a8248cc4258f096c1f28bebd9778d1df96e7e68cbabc63fe65109

Observation d191c42e-5b74-4010-8c30-85b192be3aa7 · outbound

This paper cites Arithmetic coding.IBM Journal of research and development, 23(2):149–162, 1979.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Arithmetic coding.IBM Journal of research and development, 23(2):149–162, 1979

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:22.305396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:19.525797Z digest=sha256:67a6a79303bf6169f30c5bbcf0ac4866e927205c29b42773312649c2d7eaa287

Observation 8e4c1547-574d-463d-8438-b1084a7d111c · outbound

This paper cites an unresolved cited work.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:22.087711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:19.581936Z digest=sha256:e8c8257a11ef8d02e63b50c8a9bd1534abfeb51e798c6cd09df60273ad4d0b3f

Observation 6a844d03-f3cc-46d2-b71b-58b23dce5c2c · outbound

This paper cites an unresolved cited work.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:19.697752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:32:19.697752Z digest=sha256:165d1be6fce2adeff45f04ae89ae1d41969a7852e3c31a2281f91cd2cb93b2ae

Observation dfbe890b-d91c-4e24-ab30-706c230288ce · outbound

This paper cites FlexGen: High-Throughput Generative Infer- ence of Large Language Models with a Single GPU.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding FlexGen: High-Throughput Generative Infer- ence of Large Language Models with a Single GPU

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:21.889948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:19.716104Z digest=sha256:855495ce70eecbc8aef4841a66510837f72324defedb7860cef0cc3ae5f8dd1c

Observation a30e67ad-c023-49aa-b021-828605eae64c · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:21.699086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:19.719781Z digest=sha256:f4578e921dd1c2c3656d24f932dabfcfcecdb13ccb5921b11c3371455fe8a106

Observation 72310331-4608-482e-aaa2-4249c9bdf1dc · outbound

This paper cites Zico Kolter.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Zico Kolter

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:21.525600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:19.756556Z digest=sha256:9e2e59695da1b662a8bb4ec25fbae5060f1c2461e1e51a5e9bfa0271be3ef444

Observation 7748fb8f-49b8-4b7e-9827-2715c1895a00 · outbound

This paper cites QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks, June 2024.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks, June 2024

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:21.316087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:19.797300Z digest=sha256:7723601f35eefdac11a3128f0721628e7d4c2ad4ee79d6fe3525005b2f196581

Observation caf8093d-eec4-4c4a-9e22-05a1ea09391a · outbound

This paper cites DeepCABAC: A Universal Compression Algorithm for Deep Neural Networks.IEEE Journal of Selected Topics in Signal Processing, 14(4):700–714, May 2020.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding DeepCABAC: A Universal Compression Algorithm for Deep Neural Networks.IEEE Journal of Selected Topics in Signal Processing, 14(4):700–714, May 2020

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:21.114329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:19.843328Z digest=sha256:48a4aa684681e34f0a2143ec5184eba1d886818510356013c16a0b54c7ccb3a3

Observation 11aceea7-f138-4598-bd53-90a71d68a4c1 · outbound

This paper cites Compact and computationally efficient representation of deep neural networks.IEEE Transactions on Neural Networks and Learning Systems, 31(3):772–785, 2020.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Compact and computationally efficient representation of deep neural networks.IEEE Transactions on Neural Networks and Learning Systems, 31(3):772–785, 2020

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:20.902614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:19.886008Z digest=sha256:980ee15546987ae5f0a6f3627d300be5592767465ba1facffeab6a5ff14cfd6f

Observation 085bb2b5-222b-4ca7-90dd-5f27afcb3923 · outbound

This paper cites Variational bayesian quantization.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Variational bayesian quantization

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:20.620742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:19.927150Z digest=sha256:6bcce39ec9077dbbf96cb2cd345552e39412577612204d4f35ff774b744e7e49

Observation cc116959-51fc-4bd1-8ffe-bdb6979cacb7 · outbound

This paper cites optimally compensate.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding optimally compensate

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:20.341654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:32:20.014194Z digest=sha256:ffe40583946073913251148c8b266989d1a30a82832a93b1e9ff7621b43736a6

Pith citing papers

No inbound Pith citation observations are available.