Pith. sign in

Paper Citation Record · LEDGER

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem

As of 14 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 7 inbound Pith citation observations for arXiv:2411.17525.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17525 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:07:30.282490Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T20:34:01.289119Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:29:45.503006Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8e931ca6-7b20-4c6a-8757-1ed6a701140d · outbound

This paper cites an unresolved cited work.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.084631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.084631Z digest=sha256:2945d0f8f6ee2c9ab1259606e740e7c662777412c108d5488cb29b764cabc14f

Observation cd9ef3cf-4c20-4eec-bbe1-ccf891b2d264 · outbound

This paper cites QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.089789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.089789Z digest=sha256:4f7f2603d57366e875b9abcf1b1a70055c3f6fc115d423738e6d7495881b47e0

Observation 9b1e3c6a-8c52-4a8c-8d63-19d25f3a233b · outbound

This paper cites an unresolved cited work.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.094572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.094572Z digest=sha256:ce6cc13e5b0bd391704ea008b67076df19579ceb285fedc26b387698feff7f82

Observation b39b94d3-7e74-404b-b65e-122b9f7c3a9b · outbound

This paper cites Qwen Technical Report.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Qwen Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.098277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.098277Z digest=sha256:36e7594b886ddae2990332d8fc34178f63396efed184ced104f14f24903ec01d

Observation eb61a05d-453e-4ee5-8f08-02196e7215cd · outbound

This paper cites QuIP: 2-Bit Quantization of Large Language Models With Guarantees.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem QuIP: 2-Bit Quantization of Large Language Models With Guarantees

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.102279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.102279Z digest=sha256:f7cca2643622ea4926cb94f78772a2940fb4524b06b6c594c076139734e4f879

Observation 3539d07e-54da-42dd-89c9-cf4862405e5f · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.106544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.106544Z digest=sha256:1bfdc889dbb6542a319d5f7517772ec3d0f7fecd177201dd449e533a3e0f62c0

Observation b10bcbcf-b2ac-4285-b821-b1c81616c9b6 · outbound

This paper cites New Bounds For Distributed Mean Estimation and Variance Reduction.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem New Bounds For Distributed Mean Estimation and Variance Reduction

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-12T12:07:30.689890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T12:07:30.111311Z digest=sha256:ee58e2c5cc525187a5bd0255df8a68902ff36115486f3385b254ca4a313b666f

Observation 25d5f7d4-c903-4678-a3e6-d34c4ec54ae3 · outbound

This paper cites an unresolved cited work.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:07:30.859827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T12:07:30.115289Z digest=sha256:4091ff641770f059f716650e247d4e870d9507dfedb1d02e4dc16a8ae6bebc9d

Observation 5077b3cc-9a4a-41df-9db4-ca61e7efb6e8 · outbound

This paper cites QLoRA: Efficient Finetuning of Quantized LLMs.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem QLoRA: Efficient Finetuning of Quantized LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.119506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.119506Z digest=sha256:a9d3de4670a86e15c54d5da856c1e379176a0c3b03d263e55f4f6adab6218920

Observation 30de5264-e0ff-415c-a055-75eb5ceffd8f · outbound

This paper cites SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.124299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.124299Z digest=sha256:aaa980150d4e66dd5f3d24954883cc5cc4e4fb7146afa64b4201c19ab8b91e5b

Observation 9f5544af-0a08-4f33-a644-cc588d639a2c · outbound

This paper cites an unresolved cited work.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:07:30.846789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T12:07:30.128470Z digest=sha256:8d247b20e0a0c36c7c4551d814e960c2423bfed13f88a8c31def3e03ad444c19

Observation c7fc47a2-7626-4146-8fcf-88264558226c · outbound

This paper cites The case for 4-bit precision: k-bit Inference Scaling Laws.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem The case for 4-bit precision: k-bit Inference Scaling Laws

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.132120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.132120Z digest=sha256:600b19ac3951584a2440c55b0866305b661c04ae7f9c7172c29b0ccc4f6c1d68

Observation 0006d9b4-06c8-4e03-b297-605716482785 · outbound

This paper cites The Llama 3 Herd of Models.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem The Llama 3 Herd of Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.135705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.135705Z digest=sha256:ea4d525d25817350a3b23668b89543187efe3a340678c599bcd968dd4b15e1cc

Observation ff19f2ad-17a8-454f-81c1-bed1deb2ca6f · outbound

This paper cites Extreme Compression of Large Language Models via Additive Quantization.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Extreme Compression of Large Language Models via Additive Quantization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.139505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.139505Z digest=sha256:9632da70a771132eb3b22f4355dddc137e206614a0b936efc6fd874934d9350f

Observation 8e39312b-8a66-43f7-adae-30a2ee1fee76 · outbound

This paper cites SPDY: Accurate Pruning with Speedup Guarantees.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem SPDY: Accurate Pruning with Speedup Guarantees

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.143185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.143185Z digest=sha256:c8db8ddc55461e694df6d76c34ff99f31d327c10e03df6663bce288732b66873

Observation b479fa4d-c00a-432f-ad42-855d461521a3 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.147064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.147064Z digest=sha256:b313df20f2754d5384facb40eea574f8ed891976bc96b475d7bc15c4e5af801f

Observation bae30d21-2384-42e1-abe4-7ed254df34aa · outbound

This paper cites MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.150804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.150804Z digest=sha256:eabac1d48b65bcd34a0618e637b4bc9722e1d422f8c88272bd8702606aa6c097

Observation c94f01cd-473e-477b-9b60-602e2acb6f47 · outbound

This paper cites an unresolved cited work.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.155084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.155084Z digest=sha256:522cba94f0b65c31a70c2437eb4fa89fbb8ad4d7dfe7ac8cf912a2fb4fa6e4b0

Observation 874cb9a2-565f-427a-82af-fd0c6a853de7 · outbound

This paper cites A Survey of Quantization Methods for Efficient Neural Network Inference.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem A Survey of Quantization Methods for Efficient Neural Network Inference

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.158934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.158934Z digest=sha256:2edc3ea3166a73f71a281938fd233b54876fa98a4b38f159bae8e6f3271a2033

Observation 0a5f54e7-d482-4d2c-bca2-bcd15536d4c9 · outbound

This paper cites Fast Matrix Multiplications for Lookup Table-Quantized LLMs.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Fast Matrix Multiplications for Lookup Table-Quantized LLMs

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-12T12:07:30.572708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T12:07:30.162916Z digest=sha256:87c7eb966c7fd3f31095280a69841671e612b35833a9c3cbd118f7bd14972a9d

Observation 9c02e29a-74bc-41e1-b3f4-ead42b8a1f00 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Measuring Massive Multitask Language Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.166853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.166853Z digest=sha256:017976d03d7e37ffcd0710fd4051781236f8c097aaaa9bdcca7bc563b4aa13c4

Observation 2944c500-c29d-442f-b280-76f842316fb5 · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem SqueezeLLM: Dense-and-Sparse Quantization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.170714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.170714Z digest=sha256:19b2f0cac610fa4c38e5e091b0b4aad8f93634ea9a8a00ef8a2cc5cb86979816

Observation 81ab99ba-97cb-4ae9-b6c5-eb9ecfe98788 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Gonzalez, Hao Zhang, and Ion Stoica

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.174567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.174567Z digest=sha256:359f8cc17d745a02fd3cce9bc27944fee0b602950320ea79b869215c005e1ea8

Observation de9f2010-1148-4651-b1f2-d76f010f89b6 · outbound

This paper cites OWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language Models.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem OWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.178086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.178086Z digest=sha256:bc7c19b67a76bc968912c932efc9a75c212e9e461a12971f0a90fd8e8ae4dedd

Observation f0d16a36-4a7b-417a-ab6e-6c5a67b6e53c · outbound

This paper cites AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.181926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.181926Z digest=sha256:321904d8b8a6725c6fbb570bb0b9a0131f4f3a8b1ba6a640c649f90ae1d85c5b

Observation 1b52193c-b804-4f98-afe0-6f7ffac0498d · outbound

This paper cites SpinQuant: LLM quantization with learned rotations.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem SpinQuant: LLM quantization with learned rotations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.186121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.186121Z digest=sha256:fb59c9ad2d8648ced82f5b90e53b76a699f9e800d20bc071c8eea339f43e2e25

Observation e7b03015-bed8-46fe-aa28-4d2cd2b48f6b · outbound

This paper cites Pointer Sentinel Mixture Models.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Pointer Sentinel Mixture Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.193833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.193833Z digest=sha256:05031d9fd4e293c3713059e3b00514eb37e7229093c700e1f482580420d080f4

Observation 6e6b4b6a-87e1-48e0-828c-8873a1046dc7 · outbound

This paper cites an unresolved cited work.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:07:30.828305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T12:07:30.197183Z digest=sha256:7a2045197524bcedb30f1aa13f4e25912b8bdf7e18687359df9038899332f656

Observation a27bc9cb-8365-45a3-a9e0-79a85bc36245 · outbound

This paper cites an unresolved cited work.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:07:30.816308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T12:07:30.200955Z digest=sha256:e4d0a248f1e31bc137fc324ab60b1b4c88c706c91b782b8d35108c41d610e147

Observation dc93881d-8e8d-4a22-9d54-0eaca56cdc63 · outbound

This paper cites LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.204508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.204508Z digest=sha256:d670d6af82449dc0254514641313142e9581e6bd126de7215bf67e7fcf2d2e1d

Observation 45c72733-76ad-4009-9971-acdd63cb1129 · outbound

This paper cites an unresolved cited work.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.208108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.208108Z digest=sha256:28367a5ae780c011e3d09f56414d90e332034e3ce32af725c064dc701a4f04a2

Observation 6d528b8a-5230-4e4d-970f-026ec6b1af1d · outbound

This paper cites an unresolved cited work.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:07:30.797729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T12:07:30.211589Z digest=sha256:f26c570d99be3a8b70c6097b98868cfd54c84b4a9d364d626b9686ed4e88654c

Observation 333c96cf-9aea-4074-8694-983832cf2a42 · outbound

This paper cites an unresolved cited work.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.216197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.216197Z digest=sha256:b232ae2c988285ac6014ba0488c3ec08cd1cf403f8d2b21f060d3a0854af42e5

Observation 4fc34261-3b22-415d-8261-5c19164091e0 · outbound

This paper cites Distributed Mean Estimation with Limited Communication.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Distributed Mean Estimation with Limited Communication

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T12:07:30.470908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T12:07:30.219966Z digest=sha256:1398abfb271098b3c92b9b0c92e88f8de69b2821d2e11cccadd5b412bee3dc24

Observation b3061b82-c5c0-4cdb-8cb3-c682b9622fa0 · outbound

This paper cites an unresolved cited work.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:07:30.785313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T12:07:30.224126Z digest=sha256:bb0e8521f4aea1bcceb3a3fc5048a76782867fb46bd027cc39dc0bb9378a76c4

Observation bf461b08-9a23-4f36-ae2f-beaabd37830b · outbound

This paper cites an unresolved cited work.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:07:30.773067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T12:07:30.227749Z digest=sha256:7134752659489e0447ad241cb12fb73b43bc5e190f34a790f0a6cf5648a8d9b0

Observation 17d524fe-736b-4d72-921d-460b2db5530e · outbound

This paper cites QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.231141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.231141Z digest=sha256:043d73e350f799c56d9a07d0e39187b837bfed4afe13c826cb7a12f9a7c89490

Observation 2a8598c7-5839-49c6-becd-12b696923a07 · outbound

This paper cites QTIP: Quantization with Trellises and Incoherence Processing.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem QTIP: Quantization with Trellises and Incoherence Processing

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.234824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.234824Z digest=sha256:079f7f12d77d3e08d97e73d4d313e5dc2885aaee146194f3db1355ec07d327b6

Observation f44b096f-0126-4412-8715-8e70487c6e83 · outbound

This paper cites GPTVQ: The Blessing of Dimensionality for LLM Quantization.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem GPTVQ: The Blessing of Dimensionality for LLM Quantization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.238902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.238902Z digest=sha256:04bb2cd440d5c62dcb14b05cc915b0cd1525e193e62e3b1cfeb2c2170603b7db

Observation 37994103-658c-4d4d-b324-61ea3983cde8 · outbound

This paper cites DRIVE: One-bit Distributed Mean Estimation.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem DRIVE: One-bit Distributed Mean Estimation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.242741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.242741Z digest=sha256:c07c3c46e6f4a3633569444ba63aa988356dd701925c28f24674daf793a4a10a

Observation f0226074-94dc-4eba-8178-2a37740906e0 · outbound

This paper cites EDEN: Communication-Efficient and Robust Distributed Mean Estimation for Federated Learning.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem EDEN: Communication-Efficient and Robust Distributed Mean Estimation for Federated Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.246967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.246967Z digest=sha256:38beeeeb76831098502acb14ab10d071026243c89bd21ff461e2f13fedfa37b3

Observation aae0845c-0b73-4211-8c8d-d97f2e12bd39 · outbound

This paper cites an unresolved cited work.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:07:30.761743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T12:07:30.250824Z digest=sha256:93909d3856d5c033e5cfd20c2d72eb9bbba3dae2543eff841b79d8b6e3a9e51a

Observation f1f6c9aa-6529-458e-96d3-e75f2c1c536c · outbound

This paper cites FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.254327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.254327Z digest=sha256:4219b9027871d8e29a977d331b18618d5adfe84c25835f0374d3e57ee43998df

Observation 9015d2e5-5189-4571-abb4-2871b0e365e5 · outbound

This paper cites ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.258416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.258416Z digest=sha256:40a3339584200551f27db3ebeb19b47b89ee992eddeef47e736a9b564765ada5

Observation f5c7acba-c9af-4b5d-a774-2f36e9dc24ba · outbound

This paper cites Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.262218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.262218Z digest=sha256:5ae80e208f66ebe5ff876dc922db47771256c73e4c011b766804d29b8d442bac

Observation 12283b9e-4a4e-4c96-8559-9caae4e85edd · outbound

This paper cites NF4 Isn't Information Theoretically Optimal (and that's Good).

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem NF4 Isn't Information Theoretically Optimal (and that's Good)

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.266478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.266478Z digest=sha256:011545b42af569881bcff5efe3388f18eae9f95d6c8d77927053712933762aa8

Observation 81f70250-5938-4f27-b9cd-df7b122b723f · outbound

This paper cites an unresolved cited work.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.270394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.270394Z digest=sha256:6f5a3c0301aaba345ead80498bf1f4fa379d49fe3a2af6e87d8f2adbc34cbb46

Observation 7bed2eb8-17e9-4f9a-9a1b-1a4dc8253917 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem OPT: Open Pre-trained Transformer Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.274074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.274074Z digest=sha256:b37c067e182f6c880ea9c67c93e31af037a8cf6089a2b80bb3314f0494116de4

Observation 665598d1-cad2-4c77-b593-78ad8b2f88f9 · outbound

This paper cites online" 'onlinestring :=.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem online" 'onlinestring :=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.277992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.277992Z digest=sha256:177499377a13e5c64e1a52a1b501b1d754de79ed1e6e59eb2d9240888c75cc2c

Observation ca5e6263-3184-4fe6-a383-540e4005972c · outbound

This paper cites write newline.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem write newline

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.282490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.282490Z digest=sha256:0924bb1f69d65b4e033fde96ca60775cfa711506d8265a48a2d40514fd553828

Pith citing papers

Observation f0f3d9b0-5619-4ccf-a814-6d412e9edcfc · inbound

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models cites this paper.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Pushing the Limits of Large Language Model Quantization via the Linearity Theorem

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.289119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.289119Z digest=sha256:6b8a98b3e9b826a085ca38f550e677648f43824f5e3816ff70dad2216dd8257f

Observation 26d22641-5be1-453b-bcca-af9996668f5f · inbound

CE-LoRA: Computation-Efficient LoRA Fine-Tuning for Language Models cites this paper.

CE-LoRA: Computation-Efficient LoRA Fine-Tuning for Language Models Pushing the Limits of Large Language Model Quantization via the Linearity Theorem

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T15:36:04.444303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:36:04.444303Z digest=sha256:5258ca6c0de67ee9657e04d3eb810d29d0d3e6953e8e3c6e57990b7041c64403

Observation 7660a5e6-f3c7-4897-aecd-a5e13d5e036c · inbound

KV Cache Offloading for Context-Intensive Tasks cites this paper.

KV Cache Offloading for Context-Intensive Tasks Pushing the Limits of Large Language Model Quantization via the Linearity Theorem

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:06:01.042670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T16:49:13.353700Z digest=sha256:1f6bf756f5fe7d1275abd2d98eac8b864e8ce12c96c4210f3adfc957720ae3a3

Observation 69b2aa2a-6198-48c1-b62a-6ea3ecd0af6c · inbound

KV Cache Offloading for Context-Intensive Tasks cites this paper.

KV Cache Offloading for Context-Intensive Tasks Pushing the Limits of Large Language Model Quantization via the Linearity Theorem

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:20:56.317662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:48:21.341462Z digest=sha256:1e4b32418cd4fa4455d310e18c7665d29a9c6128712f34689b60463914e31825

Observation efc361d6-9c2b-4631-9de5-5a2be0490d6e · inbound

KV Cache Offloading for Context-Intensive Tasks cites this paper.

KV Cache Offloading for Context-Intensive Tasks Pushing the Limits of Large Language Model Quantization via the Linearity Theorem

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:37:29.365075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T07:35:30.155592Z digest=sha256:905b55cb163103c9aa0f3afa23fae31e36b426110352947defd254144b8257d4

Observation a729e4b6-e5e5-43fa-9169-c18a4b781040 · inbound

KV Cache Offloading for Context-Intensive Tasks cites this paper.

KV Cache Offloading for Context-Intensive Tasks Pushing the Limits of Large Language Model Quantization via the Linearity Theorem

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:42:39.689196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T16:41:30.021302Z digest=sha256:222e026a6f9800561cdc22284c5a0279ab884f9d03bf5e16e913111e5400c217

Observation f8aab825-e78f-41ce-a5fd-3a528d9b95e3 · inbound

HyperQuant: A Rate-Distortion-Optimal Quantization Pipeline for Large Language and Diffusion Models cites this paper.

HyperQuant: A Rate-Distortion-Optimal Quantization Pipeline for Large Language and Diffusion Models Pushing the Limits of Large Language Model Quantization via the Linearity Theorem

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:45.504441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T08:46:01.500880Z digest=sha256:f6309f1a32e85ff620a806b5c8ac0d9a5a865d0d2840cf2ed366ac36dcde7563