Pith. sign in

Paper Citation Record · LEDGER

QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 53 inbound Pith citation observations for arXiv:2402.04396.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.04396 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 53 of 53 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:03:44.677581Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 28701e4b-a19a-4ee3-8cbf-e3743c6ad444 · inbound

The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits cites this paper.

The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:11:43.597896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T20:11:43.559035Z digest=sha256:9ff7325186c7a0f5eefb80e83d5679f1d1105d2cd6eeb974e2739c22985d69bf

Observation a9f02f46-12f3-4abb-b454-9281721a0e04 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 216

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:39:33.271712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:c15de363d24c6f73b1b117db399439fb2767060f0e2ed57747ac7be4433c455c

Observation 6952bc00-5e45-44c9-bd43-e9a65c78c65e · inbound

FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision cites this paper.

FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:45:36.456215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T19:45:36.337956Z digest=sha256:4ec9c8de6864fb2e6e05e413163488b6fd476d5f13e744d819f2624c3daaff65

Observation b18d4993-1044-44d9-a6ac-83d11e28ebb1 · inbound

RoSTE: An Efficient Quantization-Aware Supervised Fine-Tuning Approach for Large Language Models cites this paper.

RoSTE: An Efficient Quantization-Aware Supervised Fine-Tuning Approach for Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T23:03:44.677581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:03:44.677581Z digest=sha256:8e847e148c46550aa6465a0a2f8df6f08cbdc6abc9bdfd4169320e1652f3c1da

Observation 8f1435f5-aada-4861-9453-291e952d78cf · inbound

Rethinking the Outlier Distribution in Large Language Models: An In-depth Study cites this paper.

Rethinking the Outlier Distribution in Large Language Models: An In-depth Study QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:40.665688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:30:40.665688Z digest=sha256:518b4077a91126e7a6e7bbf4fb75a22ab3a91ce19b1155cb48f12473c42e44d7

Observation abf25eb5-8f0d-4593-965a-6a9964b135d3 · inbound

FPTQuant: Function-Preserving Transforms for LLM Quantization cites this paper.

FPTQuant: Function-Preserving Transforms for LLM Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:47.223987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:47.223987Z digest=sha256:54e72a8512a4f144730f2a853110290a920b4ce0ea87a6bfcdd33864cfdf041b

Observation f9803ddb-fe65-4a43-a073-ffdc263e8117 · inbound

PCDVQ: Enhancing Vector Quantization for Large Language Models via Polar Coordinate Decoupling cites this paper.

PCDVQ: Enhancing Vector Quantization for Large Language Models via Polar Coordinate Decoupling QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:42:41.017031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:42:41.017031Z digest=sha256:79fd547060c51315587a85c0e9aa7421e4c8fa6d4357d96a28bccb238118bed6

Observation 653af550-633e-4c12-99bf-49198969a537 · inbound

LCD: Advancing Extreme Low-Bit Clustering for Large Language Models via Knowledge Distillation cites this paper.

LCD: Advancing Extreme Low-Bit Clustering for Large Language Models via Knowledge Distillation QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:09.141264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:09.141264Z digest=sha256:13ab2c20c9a7644b589c3de391071104bb7ac866a7da68ee9f0fd257df45c497

Observation fa5f84dc-5fa2-48b3-b38d-220288c5e4af · inbound

BTC-LLM: Efficient Sub-1-Bit LLM Quantization via Learnable Transformation and Binary Codebook cites this paper.

BTC-LLM: Efficient Sub-1-Bit LLM Quantization via Learnable Transformation and Binary Codebook QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:07:20.504672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T14:03:35.214840Z digest=sha256:5c0313e0d13a2872745fcc939a5eb08b070cda25a237f9327f84197e28e2cf46

Observation 76718352-586f-4a20-830d-36091d474f31 · inbound

BASE-Q: Bias and Asymmetric Scaling Enhanced Rotational Quantization for Large Language Models cites this paper.

BASE-Q: Bias and Asymmetric Scaling Enhanced Rotational Quantization for Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:06:24.880486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:06:24.880486Z digest=sha256:12a329623d0552e8aa573b98c50b7f74f0906d6226160ba20cb8fcc852b56947

Observation 844df51a-8dcc-4883-97e2-cff04b4e25a6 · inbound

CCQ: Convolutional Code for Extreme Low-bit Quantization in LLMs cites this paper.

CCQ: Convolutional Code for Extreme Low-bit Quantization in LLMs QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:16.104514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:09:16.104514Z digest=sha256:081cca88daf9a241532ff20bdd36815d55ee2534eada3cdbdbf0b269c5671d57

Observation b468cf42-94a4-4060-b90c-7ff0b58c9d84 · inbound

Provable Post-Training Quantization: Theoretical Analysis of OPTQ and Qronos cites this paper.

Provable Post-Training Quantization: Theoretical Analysis of OPTQ and Qronos QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:52:53.059020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T23:52:06.036879Z digest=sha256:abe9162a63463e2172fce2958b1188f1d3ffc2c199d1a96d644d2787ddf85bec

Observation bc2d4450-97ef-4fca-ad2c-fd26dad50cbd · inbound

From Segments to Scenes: Temporal Understanding for Agentic Autonomous Driving via Vision-Language Models cites this paper.

From Segments to Scenes: Temporal Understanding for Agentic Autonomous Driving via Vision-Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T18:28:19.977137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:28:19.977137Z digest=sha256:c13a06d5d02bf489b0a21f7a1a13f6d6bc7193bec1ff39923d8108a8fe723747

Observation 334fd9c3-d10f-459a-86ea-b26c2e77af96 · inbound

High-Rate Quantized Matrix Multiplication I cites this paper.

High-Rate Quantized Matrix Multiplication I QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:20:53.235633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T11:18:41.456313Z digest=sha256:0c271472eb92e5f9cdcd684baa804b2a331bdbf30c24ace8000842d9afc2210f

Observation c4b4a7bf-04bf-43f8-9125-e543a17f50a8 · inbound

SALAAD: Sparse And Low-Rank Adaptation via ADMM for Large Language Model Inference cites this paper.

SALAAD: Sparse And Low-Rank Adaptation via ADMM for Large Language Model Inference QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T06:01:15.959552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:01:15.959552Z digest=sha256:40192725046b1f890cf7bf1bfd5cd315c32d402eb679d5fd4df803d07c8a338b

Observation 29b21c0b-812c-4736-98e7-520c6166b220 · inbound

Price of metric universality in vector quantization is at most 0.11 bit cites this paper.

Price of metric universality in vector quantization is at most 0.11 bit QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T04:10:43.337679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:10:43.337679Z digest=sha256:5b38c921e5338a3a34932f14fa44067898f3dac9a3e36435f3fce8e911c4ba44

Observation 5850f303-8b23-4b49-8c14-9523306ebce5 · inbound

Leech Lattice Vector Quantization for Efficient LLM Compression cites this paper.

Leech Lattice Vector Quantization for Efficient LLM Compression QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T23:10:46.775151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:10:46.775151Z digest=sha256:091f193dfee507a9f86b9f60ec9e1bf2db8cd35d8956a895b09f99720c166195

Observation 979a4726-743f-460c-995d-f138547b3c09 · inbound

Rethinking Residual Errors in Compensation-based LLM Quantization cites this paper.

Rethinking Residual Errors in Compensation-based LLM Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:21:01.232262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:10:54.439287Z digest=sha256:d08b12da454c52cc1b8b4a1a5affdbfd1a8af20eb4f3ae1d083a35d7eba864c1

Observation a82fc49e-4cc3-4433-a214-14351a05f5a7 · inbound

DuQuant++: Fine-grained Rotation Enhances Microscaling FP4 Quantization cites this paper.

DuQuant++: Fine-grained Rotation Enhances Microscaling FP4 Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:01:13.345365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T05:59:23.738476Z digest=sha256:302d649a5775244d355c75ee9e0f9f95b6c437619972bfba4e8139c01335c4d6

Observation 1c84b740-b831-4017-9c08-7a1429a8d62b · inbound

GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling cites this paper.

GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:41:02.529038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T05:29:51.182114Z digest=sha256:a21466c2d26f3157b01d692e2a14170ed965e1f4221c8911620765b11cb1330f

Observation ded0d6f2-9429-4b4d-9d74-40da99bc914d · inbound

GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling cites this paper.

GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-19T18:02:42.158526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T18:01:08.514022Z digest=sha256:2a3d0c0d21671c2300f44f660935251d6b7adaf4abe99b9a6ce8ba6208bf5a5a

Observation 7faa9924-73d0-49a2-ab8e-b55ec9c3b411 · inbound

From Signal Degradation to Computation Collapse: Uncovering the Two Failure Modes of LLM Quantization cites this paper.

From Signal Degradation to Computation Collapse: Uncovering the Two Failure Modes of LLM Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T02:38:17.179684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T02:32:50.182859Z digest=sha256:01dfc8c23fea5421b8edf4a99f0f09ca11fcb6ddcff982ade073e3b9ed1c85bc

Observation d90fc254-7403-4cc7-9e95-459904760672 · inbound

Variance Is Not Importance: Structural Analysis of Transformer Compressibility Across Model Scales cites this paper.

Variance Is Not Importance: Structural Analysis of Transformer Compressibility Across Model Scales QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:26:04.520977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T01:44:42.989053Z digest=sha256:a3990ebaf970d9155de2351f2cd47acdb655a1b99a60b6f5b38be22ab913a1dc

Observation 65084898-a928-49f3-8b7d-d021bd7234ed · inbound

BitRL: Reinforcement Learning with 1-bit Quantized Language Models for Resource-Constrained Edge Deployment cites this paper.

BitRL: Reinforcement Learning with 1-bit Quantized Language Models for Resource-Constrained Edge Deployment QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:46:43.051795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T04:23:26.079298Z digest=sha256:e7f9fbf1acff7d5e072914b01380d16252c5216ce72120c62d60eddd0783d95f

Observation 1277e456-2dab-422e-8fe1-41b5f467ee08 · inbound

When Quantization Is Free: An int4 KV Cache That Outruns fp16 on Apple Silicon cites this paper.

When Quantization Is Free: An int4 KV Cache That Outruns fp16 on Apple Silicon QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:11:14.373234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T03:16:00.868964Z digest=sha256:f73ca3c2422f4e2c8d6eecbbacab45c1afb3f5ec1f0d8bbb88cc69d8a1b2e39b

Observation e74c1b17-4971-445f-9e69-2044a2204c1e · inbound

Search Your Block Floating Point Scales! cites this paper.

Search Your Block Floating Point Scales! QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 157

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:24.139246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T05:52:26.558984Z digest=sha256:c0c2a721cf8d1ff80fa889f4cdeb6b77140d35cdb421c68f15332de186bb225a

Observation fa5e0692-584d-4ebb-a1c8-c1a8e06a9df5 · inbound

High-Rate Quantized Matrix Multiplication II cites this paper.

High-Rate Quantized Matrix Multiplication II QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:32:52.654039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T19:29:22.250363Z digest=sha256:fb980571457d2486567ad36dfab56428a03c92f34b95c7b2553e6cb1940f6548

Observation 8e2d83e8-3f0b-4d30-a204-174f4a9231be · inbound

High-Rate Quantized Matrix Multiplication II cites this paper.

High-Rate Quantized Matrix Multiplication II QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:35:04.991846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T21:25:26.862900Z digest=sha256:d5d070538731a41f49490d0e5e1c141351f9344a4c9eac003b6d51b4fdc8a4bd

Observation 69e0ff21-96d9-40f6-a763-1e2993afd757 · inbound

XFP: Quality-Targeted Adaptive Codebook Quantization with Sparse Outlier Separation for LLM Inference cites this paper.

XFP: Quality-Targeted Adaptive Codebook Quantization with Sparse Outlier Separation for LLM Inference QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:35:04.699300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T21:28:36.358474Z digest=sha256:f81d096284d65b8dc83c460e5baa9d1c8e400548610d3b41318d56cc9ce36c0b

Observation 7afb0a30-8a77-4ba0-ba68-76932eac5c88 · inbound

GAMMA: Global Bit Allocation for Mixed-Precision Models under Arbitrary Budgets cites this paper.

GAMMA: Global Bit Allocation for Mixed-Precision Models under Arbitrary Budgets QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:28:16.800354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T12:25:39.417436Z digest=sha256:42e8633795e826f0617643729f75bf31cfb90e16653f396f74fdb495add3592e

Observation 5c711e35-6b5a-40c7-aeaa-e3055fc33baa · inbound

Theory-optimal Quantization Based on Flatness cites this paper.

Theory-optimal Quantization Based on Flatness QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T22:39:10.010380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T22:38:06.888665Z digest=sha256:c5db5e086788a9ae8c02976116669cb3010d24212202f8bc3ec6fc023d7e1049

Observation 66a62db8-dd20-40f8-84a5-1e764908bff2 · inbound

Llamas on the Web: Memory-Efficient, Performance-Portable, and Multi-Precision LLM Inference with WebGPU cites this paper.

Llamas on the Web: Memory-Efficient, Performance-Portable, and Multi-Precision LLM Inference with WebGPU QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:53:55.288364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T02:52:40.923230Z digest=sha256:699840f9998b1a53a8155e742b2dbb9d1a1a88ca87c62ef5ce2b8d54a9656ef7

Observation 9a056b76-2bae-4243-beaf-d2045c7cc745 · inbound

GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs cites this paper.

GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:36:39.854470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T05:33:06.719954Z digest=sha256:48d9b6dc5c7394ccbe4336e41bb4b7dbeb7279d0d476ca2762aea27d3a310ef7

Observation a922941e-80d6-4ce0-b902-4094c5e34ef0 · inbound

MGVQ: Synergizing Multi-dimensional Sensitivity-Aware and Gradient-Hessian Fusion for Vector Quantization cites this paper.

MGVQ: Synergizing Multi-dimensional Sensitivity-Aware and Gradient-Hessian Fusion for Vector Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:05:48.338378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T17:36:45.807397Z digest=sha256:0a0ae7d412f26f77fa86ba2c613c559295682236a6321e1b9f19446e25e48c7a

Observation a829676f-80bb-44d5-9680-467e235be0fc · inbound

EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture cites this paper.

EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:44:45.610682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T14:35:02.484377Z digest=sha256:1b7be74f6f9899ca7a66a2660b4ef126a0f813b2236e48f774fb6a97e76892f1

Observation c625abbc-a448-4950-a656-559e37105655 · inbound

Influence-Inspired Spectral Rotations for Extreme Low-Bit LLM Quantization cites this paper.

Influence-Inspired Spectral Rotations for Extreme Low-Bit LLM Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:14:39.052951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T12:13:18.805668Z digest=sha256:6c33635854a6680d2e3af07f64c76fe74aa90476782fd8f833452bed2eb40ee9

Observation 66a62b39-59be-4d24-821d-bb8eb1562754 · inbound

Quantized Reasoning Models Think They Need to Think Longer, but They Do Not cites this paper.

Quantized Reasoning Models Think They Need to Think Longer, but They Do Not QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:16:00.108641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T23:05:00.401365Z digest=sha256:af97012891c3d0932dcb70409c8f2fc6319221e06e963307d486e3e8923f7edb

Observation c13b1c42-78a6-46b0-958b-7e8ed3a4b80c · inbound

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation cites this paper.

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:34.620068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T18:55:51.474956Z digest=sha256:68f7434ddf47b6d8c1940968c090b66875e7d5d968ac4ccc9c77e7695ef821b9

Observation c9ce07cf-3804-46ab-85a5-3a24dc1b8fa8 · inbound

LASER: Loss-Aware Singular-value Decomposition and Rank Allocation for Efficient Low-Precision Vision-Language Models cites this paper.

LASER: Loss-Aware Singular-value Decomposition and Rank Allocation for Efficient Low-Precision Vision-Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:12:34.668492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T19:09:29.347284Z digest=sha256:f903a796ddd0e2723caaf682bb66554707eb3bc301cd549ab756ac74d618a3ca

Observation bde79f78-77dc-4e2e-b4b9-1ee44d2b52f9 · inbound

Qift: Shift-Friendly No-Zero W2 Post-Training Quantization for Rotated W2A4/KV4 LLM Inference cites this paper.

Qift: Shift-Friendly No-Zero W2 Post-Training Quantization for Rotated W2A4/KV4 LLM Inference QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:06:15.982040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T15:49:41.836888Z digest=sha256:38418c76027e950e2f84cdad7472fb44874e76cbad18936aff43c540b0e6a65e

Observation d3d7e18c-cfd3-49cb-8c9a-7aed78fa2441 · inbound

LiftQuant: Continuous Bit-Width LLM via Dimensional Lifting and Projection cites this paper.

LiftQuant: Continuous Bit-Width LLM via Dimensional Lifting and Projection QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:26.980422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T11:14:03.535306Z digest=sha256:96b8f63290c46dcee737cb06e726ade983c40ae97219ac7336db0dd504b3c59e

Observation d37ff3ed-95ea-4597-a431-61d0017580b9 · inbound

LiftQuant: Continuous Bit-Width LLM via Dimensional Lifting and Projection cites this paper.

LiftQuant: Continuous Bit-Width LLM via Dimensional Lifting and Projection QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:24:38.273775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T11:17:53.736872Z digest=sha256:3e49874de44f938864224fd64639d7d03aad142a84adcb5e326e95443ad23e60

Observation 24c04631-cb20-4438-892f-108233728ee6 · inbound

Minimizing the Hidden Cost of Scales: Graph-Guided Ultra-Low-Bit Quantization for Large Language Models cites this paper.

Minimizing the Hidden Cost of Scales: Graph-Guided Ultra-Low-Bit Quantization for Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:16:48.288085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T06:09:42.838355Z digest=sha256:16bd46c5117779ca8ff36879ca093b82af1cba634942ee4a6eec69010445b1e1

Observation 93ff78ee-bb36-434a-b4af-36eaf89af560 · inbound

FAIR-Calib: Frontier-Aware Instability-Reweighted Calibration for Post-Training Quantization of Diffusion Large Language Models cites this paper.

FAIR-Calib: Frontier-Aware Instability-Reweighted Calibration for Post-Training Quantization of Diffusion Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:55.577768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T02:35:00.563647Z digest=sha256:24dd1bcc641998178a1133cda7a5a854c83e7338c1e11b49ddcf62cbf76766e6

Observation c0132601-3a29-4c84-96e1-d0977ae20abc · inbound

FAIR-Calib: Frontier-Aware Instability-Reweighted Calibration for Post-Training Quantization of Diffusion Large Language Models cites this paper.

FAIR-Calib: Frontier-Aware Instability-Reweighted Calibration for Post-Training Quantization of Diffusion Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T12:23:13.879878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:23:13.879878Z digest=sha256:47b6e546269fb35747892ca96af247b2ad45e664949565d056d4c63e642177f2

Observation 70b88107-7b85-4d66-a397-62ab0c9862d1 · inbound

HyperQuant: A Rate-Distortion-Optimal Quantization Pipeline for Large Language and Diffusion Models cites this paper.

HyperQuant: A Rate-Distortion-Optimal Quantization Pipeline for Large Language and Diffusion Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:45.532847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T08:46:01.500880Z digest=sha256:441ad5385139b5f5edb992355e171624b0497171fc293efac40a183aa80cad1e

Observation 2f57c8e7-0ff3-4570-96ce-312bc4869769 · inbound

GRINQH: Graded Input-based Quantization Hierarchy for Efficient LLM Generation cites this paper.

GRINQH: Graded Input-based Quantization Hierarchy for Efficient LLM Generation QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:39:45.441374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T08:38:32.577228Z digest=sha256:c079adfae46c4543be32f28176058fb9eafc7a4925dc9f9eabb74d6260d5537c

Observation 75cbd6cf-71ce-45fd-bc84-c577b9e28297 · inbound

Reliability Scaling Laws for Quantized Large Language Models cites this paper.

Reliability Scaling Laws for Quantized Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 160

Resolution
unresolved
no resolver link, observed 2026-07-14T08:45:52.855783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T08:45:52.855783Z digest=sha256:33c9093e730b06ee26046da5e29aec4fa8085351035c3545de419b3a1ed62107

Observation 6f7f2d25-95d6-4df6-a7d3-18d90b9d3ac6 · inbound

Cross-Layer Error Compensation and Finite-Sample Feature-Statistics Matching for Extreme Low-Bit Quantization of Large Language Models cites this paper.

Cross-Layer Error Compensation and Finite-Sample Feature-Statistics Matching for Extreme Low-Bit Quantization of Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T01:35:43.851034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:35:43.851034Z digest=sha256:dd600749facd96f6ce0f3d5099ed067d089ae0de9b9e1535666997a92a876a23

Observation 0d47cdbc-5d5d-4238-836f-77290468ea3e · inbound

Break Through the Compression Bottleneck: From Theory to Practice cites this paper.

Break Through the Compression Bottleneck: From Theory to Practice QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T14:29:28.509827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:29:28.509827Z digest=sha256:5bc2150ef5bcb39dde39229dcf321a8f94929ed44ecf1c2bb6a693d28eca289c

Observation 829af20d-2a27-471e-a618-badf62ddb760 · inbound

A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference cites this paper.

A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-31T22:36:29.341483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T22:36:29.341483Z digest=sha256:b75d31c8aa2807b6514575bc0b112c71b239ffd9d4d389461dd4b29f1990cf1d

Observation 34d4b659-a33f-4912-8ed3-172af34aa219 · inbound

GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference cites this paper.

GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T03:16:46.931878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:16:46.931878Z digest=sha256:2e939ab95e4c3cbf631ffe3a3664745c82522be44066072736cf449dce7af9f1

Observation c2ee7334-ce78-4440-80e0-3553412fa3c8 · inbound

TaskPress: Query-Agnostic KV Cache Compression via Task-Guided Pruning cites this paper.

TaskPress: Query-Agnostic KV Cache Compression via Task-Guided Pruning QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T21:59:39.992784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:59:39.992784Z digest=sha256:92d338e20dbb4cd3707c0835e26faaeee655291902bbdae9ab17aca622dce7d5