Pith. sign in

Paper Citation Record · LEDGER

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference

As of 10 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 7 inbound Pith citation observations for arXiv:2502.00922.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.00922 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T17:18:51.559897Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T13:34:08.346536Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T16:25:49.759745Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact9
  • verified fuzzy14
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0b6211d1-20cb-4f17-aac3-77fa40756c16 · outbound

This paper cites QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.401236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.401236Z digest=sha256:735b758a59a0a736a96e410a6c6e382bc281dbfd98424984fbc811a38aa7837c

Observation 497d3a6e-d881-4038-a9ba-c831f6a46599 · outbound

This paper cites B., Muralimanohar, N., Shafiee, A., and Srinivas, V.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference B., Muralimanohar, N., Shafiee, A., and Srinivas, V

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:18:54.153310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.406018Z digest=sha256:16a7c891f5f68094cea4a0f69147e1afad611e84c4313ddb149382a2ba3fc442

Observation 39175179-70b4-42d0-a441-12f5585cd0d6 · outbound

This paper cites S., and Sze, V.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference S., and Sze, V

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:18:54.142123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.410365Z digest=sha256:07baa7755f5ab0ed08abff5d5b75ef80c654c3024bed62997629a86b6314dcd6

Observation 0df24379-621e-4e64-bbd0-786c816a7e30 · outbound

This paper cites E., Stoica, I., and Xing, E.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference E., Stoica, I., and Xing, E

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.414132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.414132Z digest=sha256:e193c5ec96b0e0d569119243d0dc9b4816f5731ca48e8b3f371713248ab92497

Observation 186dc68c-d741-4e6d-9d13-4ae72cd31c86 · outbound

This paper cites B., O’Connor, M., Erez, M., Pool, J., Nellans, D., and Keckler, S.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference B., O’Connor, M., Erez, M., Pool, J., Nellans, D., and Keckler, S

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:18:54.123443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.417709Z digest=sha256:f0893da67dacabf63e63baee8efbd35e883cf866d5e50928f3958dcf704e42e7

Observation 5d581117-3a82-41e7-92a0-6d96e4d5d955 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.421395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.421395Z digest=sha256:c0d0d9441350e7b1c2481c5dd399ab523686df0c3be0c16e8dae3a255083f669

Observation d98a464e-7071-45a2-a9d3-a9a054ff5b11 · outbound

This paper cites an unresolved cited work.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.425873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.425873Z digest=sha256:221ffa801f8a17c9e004f6b158795588a50e9559f2fbc27ef51323c0fb6a6e79

Observation db0ee736-faeb-4c5d-bb65-6c7abe8b9028 · outbound

This paper cites The Llama 3 Herd of Models.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.429444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.429444Z digest=sha256:8444d59a5a46493b790930a8fd5a2570c4a63c1cec2179f947d323bbe66088e0

Observation f0da55fd-4ea4-4a84-965c-c9b1d98d74dd · outbound

This paper cites Accuracy is Not All You Need.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Accuracy is Not All You Need

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.432553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.432553Z digest=sha256:451359c0651f0e75480356098610702f5fe043e6283770b0f93ce889ed33fd4e

Observation 88038bfb-7d00-4cc3-a2d0-d1b65072ce7d · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.436104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.436104Z digest=sha256:0c2bfef4f4ad28cba854c261213e5df412b6a492b5a52b639f90b0986b2d43ba

Observation 39d89a7e-e826-4183-bfae-52ca8ac13d3f · outbound

This paper cites Does reduced precision hurt? Blog post, 2024.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Does reduced precision hurt? Blog post, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:18:54.104745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.440045Z digest=sha256:3d52aac4eae629ba1c699c021ef065d54489f928e062db202b4e3db26c57e09c

Observation 52288120-1fcf-4d33-8116-68bb996df7f7 · outbound

This paper cites Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.443586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.443586Z digest=sha256:31235038591904838837ca0bcd67abe7c49824c517050c05ad4ca745632d3dd2

Observation c5cb69f3-2a45-4c8c-a276-0dcbc753bf1a · outbound

This paper cites NeuZip: Memory-Efficient Training and Inference with Dynamic Compression of Neural Networks.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference NeuZip: Memory-Efficient Training and Inference with Dynamic Compression of Neural Networks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.447148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.447148Z digest=sha256:bdc861b870e6ef86a941ebfa879d06df30955439bbd4cd1ecfbf41fec586a324

Observation 24224d4b-9e44-4677-86db-9cca85038bbb · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Measuring Massive Multitask Language Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.450902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.450902Z digest=sha256:235edb79de5f402d2bff7d0a66e8364e0b5d9453563048ff66464ca4b65fd878

Observation 13f2b916-e420-45f9-86c6-38e9454a6290 · outbound

This paper cites ZipNN: Lossless Compression for AI Models.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference ZipNN: Lossless Compression for AI Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.455245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.455245Z digest=sha256:d1bce3985ac018e7954785c021bc74f1b56d2992cfb4f5dd6ca454d6561a11c9

Observation faf66d4a-73f0-4672-bcef-a72f4ce154ff · outbound

This paper cites Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.459650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.459650Z digest=sha256:cb50882d528a72955f53f389e464e30a8ffce6545bdc9e7fd973e32536938af6

Observation 9b83baf2-131f-4e76-bdd9-97215c84da27 · outbound

This paper cites S., Choi, Y., Kim, C., Kim, Y., Yu, H., Abdel-Aziz, H., Park, J.-S., Lee, H., Lee, D., Kim, M.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference S., Choi, Y., Kim, C., Kim, Y., Yu, H., Abdel-Aziz, H., Park, J.-S., Lee, H., Lee, D., Kim, M

Reference 17

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T17:18:53.851033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.464502Z digest=sha256:f7456e21e2c5f5179b7e0df28a9343528537dcb0d9a56d7cfdc701f7f5bfe693

Observation c5d10bca-e45a-493e-856c-d855a22457bf · outbound

This paper cites G., Zimmer, B., Dally, W.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference G., Zimmer, B., Dally, W

Reference 18

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T17:18:53.677618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.468359Z digest=sha256:9d199370627d8bd02a59843c41c8abaca2b530542b82deea1ab60eda3ccd5633

Observation 06fd14a6-0be5-47d4-8133-13c3a327a963 · outbound

This paper cites Bit-plane compression: Transforming data for better compression in many-core architectures.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Bit-plane compression: Transforming data for better compression in many-core architectures

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:18:54.093391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.472747Z digest=sha256:0eac423947e137d333fd939a50f27f833af038b874a4faa9a425275a6cd706b7

Observation 544b551d-9f22-46a9-92d3-77aac1f6f33c · outbound

This paper cites Cerebras architecture deep dive: First look inside the hardware/software co-design for deep learning.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Cerebras architecture deep dive: First look inside the hardware/software co-design for deep learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:18:54.082633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.476815Z digest=sha256:264ea71b71e7c7bf64dbabe3f74988e0a509935bbcd7c0b142c4588b41c8b0b8

Observation 89a7859f-4c67-47aa-8cbf-fcdf1e3a6091 · outbound

This paper cites Awq: Activation-aware weight quantization for on-device llm compression and acceleration.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Awq: Activation-aware weight quantization for on-device llm compression and acceleration

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.480864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.480864Z digest=sha256:c2d497139191d678aa072de3064f0cdfbd188e371c206606cd82dc66752546e8

Observation 76ad177f-f306-4596-b79c-60b751e0cc29 · outbound

This paper cites How Does Quantization Affect Multilingual LLMs?.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference How Does Quantization Affect Multilingual LLMs?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.484484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.484484Z digest=sha256:dcaee3ed8db099f40943dc1dc52e7c3625a7c440693bfd83abd042ff5525f87f

Observation 2c3a6913-b377-46a5-acec-ed3e0f8d94b7 · outbound

This paper cites and Mutyam, M.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference and Mutyam, M

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:18:54.065251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.488227Z digest=sha256:5a0201eb69aeb37588cf66bfb7093fd3b97868a4b9bc84ca2633b286c5e1633c

Observation eccdb6a4-6572-4238-b29b-1f48900cdbc1 · outbound

This paper cites S., Chen, Y.-H., Ying, V.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference S., Chen, Y.-H., Ying, V

Reference 24

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T17:18:53.514198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.491858Z digest=sha256:391aac2a9ed0a247aecce9cb17f2b96d79a552b9575bb1e6cb817e3d3b14904c

Observation dbd0fc17-7f86-4d95-ba8a-5b423af1eec2 · outbound

This paper cites Arrayflex: A systolic array architecture with configurable transparent pipelining.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Arrayflex: A systolic array architecture with configurable transparent pipelining

Reference 25

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T17:18:53.348934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.495866Z digest=sha256:6f2f68545da2cf29698ed94a3fd5bb32219e4d7186ead8b843ed14cb99726cf1

Observation 5ca734ab-e616-4df4-83f6-da58ce0372d8 · outbound

This paper cites M., Zhu, Y., Whatmough, P., Mattina, M., and Krishna, T.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference M., Zhu, Y., Whatmough, P., Mattina, M., and Krishna, T

Reference 26

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T17:18:53.196253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.499709Z digest=sha256:8303a4472d382adb125e1ee58a44bf35f828d5c8590042a7ba355b8b8bc986d2

Observation a4da60c4-3f85-4183-8837-66e36c6d6340 · outbound

This paper cites S., Reagen, B., Wei, G.-Y., and Brooks, D.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference S., Reagen, B., Wei, G.-Y., and Brooks, D

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:18:54.054810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.503608Z digest=sha256:cf846e00c06bc94f0e70b388540b2996512e18d5dac81b62fff61cb30e92a48b

Observation 48ac90d6-2c7c-4844-b235-e3c4b65aed84 · outbound

This paper cites S., Clemons, J., Venkatesan, R., Zimmer, B., Fojtik, M., Jiang, N., Keller, B., Klinefelter, A., Pinckney, N., Raina, P., Tell, S.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference S., Clemons, J., Venkatesan, R., Zimmer, B., Fojtik, M., Jiang, N., Keller, B., Klinefelter, A., Pinckney, N., Raina, P., Tell, S

Reference 28

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T17:18:53.032808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.507466Z digest=sha256:2a3c1ad9a5795e98d9a7afdcb47b828d639431d3274b0101cebadde076a2065a

Observation 6ddde6e5-8229-4f70-9da0-ad8bb8de5360 · outbound

This paper cites The nvidia deep learning accelerator.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference The nvidia deep learning accelerator

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:18:54.043556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.510907Z digest=sha256:8775c2205e766025958ae020c3b6ba26a739af110654e66ad0a204b0d3c59d12

Observation f3e699c8-7458-4341-bb9c-e14c2b39b7e5 · outbound

This paper cites Q., Gomez, J., Khwa, W.-S., Sarwar, S.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Q., Gomez, J., Khwa, W.-S., Sarwar, S

Reference 30

Resolution
verified exact
doi, observed 2026-08-09T17:18:51.593843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.514086Z digest=sha256:d0beea5e77537317f8b770817702d9ea32764739b942763a217fa92851a52d19

Observation 84d6e910-77b1-4637-8770-53fafc06b513 · outbound

This paper cites Google coral edge tpu board vs nvidia jetson nano dev board hardware comparison, 2020.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Google coral edge tpu board vs nvidia jetson nano dev board hardware comparison, 2020

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:18:54.032560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.518081Z digest=sha256:67264d79fa2a53e3585a965f14887521f3e405c8090c993d31ba6b8f41fe5584

Observation 28fef51e-1b3b-4d36-9759-1671562d4210 · outbound

This paper cites Llama3.1 model quality evaluation: Cerebras, groq, sambanova, together, and fireworks.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Llama3.1 model quality evaluation: Cerebras, groq, sambanova, together, and fireworks

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:18:54.020042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.521058Z digest=sha256:f7c3dfeb83abeeb521016ea17267f8a53847d91805fe7d5c19ac6ec9f3b75271

Observation 13900705-0c5e-4824-b43d-9cf2bc913a95 · outbound

This paper cites S., Wang, M., Clemons, J., Dai, S., Fojtik, M., Keller, B., Klinefelter, A., Pinckney, N., Raina, P., Zhang, Y., Zimmer, B., Dally, W.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference S., Wang, M., Clemons, J., Dai, S., Fojtik, M., Keller, B., Klinefelter, A., Pinckney, N., Raina, P., Zhang, Y., Zimmer, B., Dally, W

Reference 33

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T17:18:52.752480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.524265Z digest=sha256:c8e23a15b33f91230150cc0cc54a788dac9fec4972c82acae2abbe6584b8eed7

Observation f6d8a82c-08b5-4783-b062-57f30f0e6988 · outbound

This paper cites Spatten: Efficient sparse attention architecture with cascade token and head pruning.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Spatten: Efficient sparse attention architecture with cascade token and head pruning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:18:54.008849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.527838Z digest=sha256:c4b4fb1b2646b3d1a777be7fd00c95680acd93bfa9c9923f20352a70c6709a81

Observation 4aa3a4ab-ec52-4e97-a1b2-a5744c3ace34 · outbound

This paper cites an unresolved cited work.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-09T17:18:53.997080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.531468Z digest=sha256:f86e10e5cc71dd163ed3392ff3bcb7ad3f554324e4945dea7129f33d8856cff8

Observation 2add9484-be78-4852-8edb-f4ea43b8cc38 · outbound

This paper cites The roofline model: A pedagogical tool for program analysis and optimization.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference The roofline model: A pedagogical tool for program analysis and optimization

Reference 36

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T17:18:51.788060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.535139Z digest=sha256:0c12a3105595b457d09f50a467f55ef688053252d1c0834f3ef27e0990128135

Observation f05c421a-88d6-45f8-a4cf-923223872d26 · outbound

This paper cites Beyond Perplexity: Multi-dimensional Safety Evaluation of LLM Compression.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Beyond Perplexity: Multi-dimensional Safety Evaluation of LLM Compression

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.538529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.538529Z digest=sha256:7d294f6d2b4f1fc71a337615b6e1f02fc76bc150099e7b88b3dac550998bc251

Observation 294af4e8-a8bf-4134-b83c-e97dec1d7de7 · outbound

This paper cites Qwen2.5 Technical Report.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Qwen2.5 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.542155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.542155Z digest=sha256:763f99d20558754931b1556254473dacc5b61cf022441bccfe875824b924f84d

Observation 12ec46ef-eeb0-4e0b-9780-4187f5e6b99f · outbound

This paper cites Zeroquant: Efficient and affordable post-training quantization for large-scale transformers.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Zeroquant: Efficient and affordable post-training quantization for large-scale transformers

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:18:53.985520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.545813Z digest=sha256:4f239269a8972c66d4ceead4d099f0588b22aca3fc93165055079453029ecbae

Observation a00658f7-e833-4828-ae7b-f6cbf34132b1 · outbound

This paper cites 15.1 a 0.795 fj/bit physically-unclonable function-protected tcam for a software-defined networking switch.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference 15.1 a 0.795 fj/bit physically-unclonable function-protected tcam for a software-defined networking switch

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:18:53.974529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.549387Z digest=sha256:109274db08e2d48d1bc342d7e9015e2abd7cf4fae9cfc384bccc9536c59645fb

Observation a3fa07f6-ec20-41a8-b086-fd63779593c6 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference OPT: Open Pre-trained Transformer Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.552602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.552602Z digest=sha256:a03dee0a31114002f675826c963d568bf4ca5bc9e791547d1d1bba0e633b0b6c

Observation 8a9fe5e5-78ad-43ad-a1e4-79127df8e422 · outbound

This paper cites Catastrophic Failure of LLM Unlearning via Quantization.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Catastrophic Failure of LLM Unlearning via Quantization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.556165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.556165Z digest=sha256:8c70687ca65c93a762f584f58e616b5059b02a96a0263f9ed9b3c161914ebc18

Observation 3a7a4bc4-b2e2-4205-a348-932c9ca58099 · outbound

This paper cites write newline.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference write newline

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.559897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.559897Z digest=sha256:1f6b7d730460092813a9c9ac4ccce5f4848ad0b0d3f125db20074f31fbdd31d3

Pith citing papers

Observation e1dfcaf3-3b2c-476f-8f79-2ca47a43a852 · inbound

ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity Scheduling cites this paper.

ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity Scheduling Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:45:29.195963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T07:43:24.937953Z digest=sha256:8ac108450bad485cf264a865d458fbe01d6ee8a79b51ae158bf137de7be2507f

Observation fb484879-642c-4585-888e-3487ebe49ece · inbound

ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUs cites this paper.

ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUs Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:58:03.058558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T21:56:38.684104Z digest=sha256:10828a1f5990baaea190821af2ba80924d8d729ef33ca7cf1b3a450ca7143bbc

Observation 403395a6-6800-428f-b105-edc52095be6c · inbound

SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving cites this paper.

SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:26:08.754295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T16:59:58.809897Z digest=sha256:eb9222eb36bcec8678bae891c48055ee96b0f18c9e99457fbfdf69228f56fa95

Observation 8f590228-6923-4cd6-9777-d30678fc3d88 · inbound

SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving cites this paper.

SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:21:23.650906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:15:06.859289Z digest=sha256:074944cdcf89bba8698b16ef58491180158633c3a03bc6f936102568c06305f0

Observation 983e8689-bffa-4257-8f33-69b430b335a2 · inbound

SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving cites this paper.

SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T13:15:45.375310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T00:58:28.798369Z digest=sha256:f12e862eafe5e8808f9631d66102e2def39de87bd48b4917125763f38dbdc8de

Observation d2d39926-d8ed-4696-a386-0251e0c1bfae · inbound

Cassandra: Enabling Reasoning LLMs at Edge via Self-Speculative Decoding cites this paper.

Cassandra: Enabling Reasoning LLMs at Edge via Self-Speculative Decoding Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T16:25:49.761209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T16:19:54.910430Z digest=sha256:6becc095f36b711abe9a9fa0e20edf53b8db030d959212c57cd0942b07a55780

Observation 99a6d1d4-d55e-4205-8c80-fc89ebb29591 · inbound

Lossless Tensor Compression as Program Synthesis cites this paper.

Lossless Tensor Compression as Program Synthesis Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T13:34:08.346536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:34:08.346536Z digest=sha256:ed1b434403edd1a61a872c284e292564642d5706b601515fa73f6779364e098e