Pith. sign in

Paper Citation Record · LEDGER

QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 53 inbound Pith citation observations for arXiv:2402.04396.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.04396 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 53 of 53 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:03:44.677581Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 28701e4b-a19a-4ee3-8cbf-e3743c6ad444 · inbound

The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits cites this paper.

The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:11:43.597896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T20:11:43.559035Z digest=sha256:497bb0318980a64ca32ec85e63d9e4d071555569aaf1090227cdbca630d17e00

Observation a9f02f46-12f3-4abb-b454-9281721a0e04 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 216

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:39:33.271712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:c61648cf0ba2c5f33d9a44401a4c5c17f37fb0ef8a2a387e3f59a5bd431b8c65

Observation 6952bc00-5e45-44c9-bd43-e9a65c78c65e · inbound

FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision cites this paper.

FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:45:36.456215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T19:45:36.337956Z digest=sha256:fb74b1f74cf9814553e75ce85f71cc403668033431b7e02b6405a0af24439986

Observation b18d4993-1044-44d9-a6ac-83d11e28ebb1 · inbound

RoSTE: An Efficient Quantization-Aware Supervised Fine-Tuning Approach for Large Language Models cites this paper.

RoSTE: An Efficient Quantization-Aware Supervised Fine-Tuning Approach for Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T23:03:44.677581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:03:44.677581Z digest=sha256:ad253007f5abc8e2b0809e8e31e3987c9ad6e8b705eecb1cd9ae5c90d5d08b91

Observation 8f1435f5-aada-4861-9453-291e952d78cf · inbound

Rethinking the Outlier Distribution in Large Language Models: An In-depth Study cites this paper.

Rethinking the Outlier Distribution in Large Language Models: An In-depth Study QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:40.665688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:30:40.665688Z digest=sha256:1596661ff2595d209d6c9f9d7bd351ae0be130540803eb3718db54957124f240

Observation abf25eb5-8f0d-4593-965a-6a9964b135d3 · inbound

FPTQuant: Function-Preserving Transforms for LLM Quantization cites this paper.

FPTQuant: Function-Preserving Transforms for LLM Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:47.223987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:47.223987Z digest=sha256:7734c5cbdc42e68e9c6dcc942baf05054a95d1917485a16ce2ceba1d8ecd9c88

Observation f9803ddb-fe65-4a43-a073-ffdc263e8117 · inbound

PCDVQ: Enhancing Vector Quantization for Large Language Models via Polar Coordinate Decoupling cites this paper.

PCDVQ: Enhancing Vector Quantization for Large Language Models via Polar Coordinate Decoupling QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:42:41.017031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:42:41.017031Z digest=sha256:f0b9973d66581a48b73df666ebde9f01ddb1278d033f3d7ecada7575323b716d

Observation 653af550-633e-4c12-99bf-49198969a537 · inbound

LCD: Advancing Extreme Low-Bit Clustering for Large Language Models via Knowledge Distillation cites this paper.

LCD: Advancing Extreme Low-Bit Clustering for Large Language Models via Knowledge Distillation QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:09.141264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:09.141264Z digest=sha256:b66f2ddc456f18d577a2fc501487e0c33178f662d1e6a2ff45749fe0b3421328

Observation fa5f84dc-5fa2-48b3-b38d-220288c5e4af · inbound

BTC-LLM: Efficient Sub-1-Bit LLM Quantization via Learnable Transformation and Binary Codebook cites this paper.

BTC-LLM: Efficient Sub-1-Bit LLM Quantization via Learnable Transformation and Binary Codebook QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:07:20.504672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T14:03:35.214840Z digest=sha256:20c41fabe9cc4a88d02dae442b985b931a30d5a50bdd1545d03af3c8c5217dc6

Observation 76718352-586f-4a20-830d-36091d474f31 · inbound

BASE-Q: Bias and Asymmetric Scaling Enhanced Rotational Quantization for Large Language Models cites this paper.

BASE-Q: Bias and Asymmetric Scaling Enhanced Rotational Quantization for Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:06:24.880486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:06:24.880486Z digest=sha256:7af289d86455ea5549a83082f9802aa7953a7f775039c0d58ab21bef08ee1048

Observation 844df51a-8dcc-4883-97e2-cff04b4e25a6 · inbound

CCQ: Convolutional Code for Extreme Low-bit Quantization in LLMs cites this paper.

CCQ: Convolutional Code for Extreme Low-bit Quantization in LLMs QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:16.104514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:09:16.104514Z digest=sha256:ffa6fdd389a7fbb285c747536d7bc44394c8784f8112157a1417e757285fa6e7

Observation b468cf42-94a4-4060-b90c-7ff0b58c9d84 · inbound

Provable Post-Training Quantization: Theoretical Analysis of OPTQ and Qronos cites this paper.

Provable Post-Training Quantization: Theoretical Analysis of OPTQ and Qronos QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:52:53.059020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T23:52:06.036879Z digest=sha256:e00cd8c14f4c6cfe4cf53277d3fb804cb349dffb3e2343c21eb2f596338dd288

Observation bc2d4450-97ef-4fca-ad2c-fd26dad50cbd · inbound

From Segments to Scenes: Temporal Understanding for Agentic Autonomous Driving via Vision-Language Models cites this paper.

From Segments to Scenes: Temporal Understanding for Agentic Autonomous Driving via Vision-Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T18:28:19.977137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:28:19.977137Z digest=sha256:8af4e9a1b37b84691b70a705d83d04f87c15f26fd4c846a9d356441dd61977f2

Observation 334fd9c3-d10f-459a-86ea-b26c2e77af96 · inbound

High-Rate Quantized Matrix Multiplication I cites this paper.

High-Rate Quantized Matrix Multiplication I QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:20:53.235633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T11:18:41.456313Z digest=sha256:a6f625add964211987fe95f1a64bfb3604bdd059f533451de17acfb80e4f07a3

Observation c4b4a7bf-04bf-43f8-9125-e543a17f50a8 · inbound

SALAAD: Sparse And Low-Rank Adaptation via ADMM for Large Language Model Inference cites this paper.

SALAAD: Sparse And Low-Rank Adaptation via ADMM for Large Language Model Inference QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T06:01:15.959552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:01:15.959552Z digest=sha256:f5eac8642f2b1c5e2d5738f2c03186beb1542af0b59c6fb4e693c026386ac2fd

Observation 29b21c0b-812c-4736-98e7-520c6166b220 · inbound

Price of metric universality in vector quantization is at most 0.11 bit cites this paper.

Price of metric universality in vector quantization is at most 0.11 bit QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T04:10:43.337679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:10:43.337679Z digest=sha256:8070379e7de85ca3686890a72eefb48f1d6ce0851def9e3262d57d4672c7fc7b

Observation 5850f303-8b23-4b49-8c14-9523306ebce5 · inbound

Leech Lattice Vector Quantization for Efficient LLM Compression cites this paper.

Leech Lattice Vector Quantization for Efficient LLM Compression QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T23:10:46.775151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:10:46.775151Z digest=sha256:d00e070aadf6101f68db28e238ada41da9f6451a0df3be030a42a3880e2be04c

Observation 979a4726-743f-460c-995d-f138547b3c09 · inbound

Rethinking Residual Errors in Compensation-based LLM Quantization cites this paper.

Rethinking Residual Errors in Compensation-based LLM Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:21:01.232262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:10:54.439287Z digest=sha256:92d307514b300bc58dc97e0579abcce034ac0903fc7fdba506b4a2850006f006

Observation a82fc49e-4cc3-4433-a214-14351a05f5a7 · inbound

DuQuant++: Fine-grained Rotation Enhances Microscaling FP4 Quantization cites this paper.

DuQuant++: Fine-grained Rotation Enhances Microscaling FP4 Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:01:13.345365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T05:59:23.738476Z digest=sha256:4a48e9dcce7b3df465432b788b7f3a86cf9ddd69600b848a125823c90a81a2eb

Observation 1c84b740-b831-4017-9c08-7a1429a8d62b · inbound

GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling cites this paper.

GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:41:02.529038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T05:29:51.182114Z digest=sha256:14c7629009dcbb4007257becf7db8c4d4fb68e82df51e3b5ad05b770bd776835

Observation ded0d6f2-9429-4b4d-9d74-40da99bc914d · inbound

GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling cites this paper.

GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-19T18:02:42.158526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T18:01:08.514022Z digest=sha256:a3220ca41ee35413d2a0a541def05bde9467d28e4bd7bdfe53cfda79b86765e2

Observation 7faa9924-73d0-49a2-ab8e-b55ec9c3b411 · inbound

From Signal Degradation to Computation Collapse: Uncovering the Two Failure Modes of LLM Quantization cites this paper.

From Signal Degradation to Computation Collapse: Uncovering the Two Failure Modes of LLM Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T02:38:17.179684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T02:32:50.182859Z digest=sha256:abb0983054b801e895edf36125460a4eecd3b0180d5c1e03e5f08022f34a8687

Observation d90fc254-7403-4cc7-9e95-459904760672 · inbound

Variance Is Not Importance: Structural Analysis of Transformer Compressibility Across Model Scales cites this paper.

Variance Is Not Importance: Structural Analysis of Transformer Compressibility Across Model Scales QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:26:04.520977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T01:44:42.989053Z digest=sha256:706ad6e0fef3e25e6aaebb2a72ecf63f6035b0c1399aa6d9b17e11985ec9e5b5

Observation 65084898-a928-49f3-8b7d-d021bd7234ed · inbound

BitRL: Reinforcement Learning with 1-bit Quantized Language Models for Resource-Constrained Edge Deployment cites this paper.

BitRL: Reinforcement Learning with 1-bit Quantized Language Models for Resource-Constrained Edge Deployment QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:46:43.051795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T04:23:26.079298Z digest=sha256:0c4cf65e285fb54399e5e25e05ed5bbed21fb1436d8928f8faa9164d7f5c2212

Observation 1277e456-2dab-422e-8fe1-41b5f467ee08 · inbound

When Quantization Is Free: An int4 KV Cache That Outruns fp16 on Apple Silicon cites this paper.

When Quantization Is Free: An int4 KV Cache That Outruns fp16 on Apple Silicon QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:11:14.373234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T03:16:00.868964Z digest=sha256:54c405db6e3521b3370c460f429490c8a42a871a089bff180723b48ff28b7a74

Observation e74c1b17-4971-445f-9e69-2044a2204c1e · inbound

Search Your Block Floating Point Scales! cites this paper.

Search Your Block Floating Point Scales! QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 157

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:24.139246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T05:52:26.558984Z digest=sha256:7338775f8906897a2e8d6515828481f5987d108c7b4215470a9bfba64cfbdc1c

Observation fa5e0692-584d-4ebb-a1c8-c1a8e06a9df5 · inbound

High-Rate Quantized Matrix Multiplication II cites this paper.

High-Rate Quantized Matrix Multiplication II QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:32:52.654039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T19:29:22.250363Z digest=sha256:db0107c5c98af13c113fd8bd57c05c739c6acc33e3186d851629581988bc956f

Observation 8e2d83e8-3f0b-4d30-a204-174f4a9231be · inbound

High-Rate Quantized Matrix Multiplication II cites this paper.

High-Rate Quantized Matrix Multiplication II QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:35:04.991846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T21:25:26.862900Z digest=sha256:fa9c49b97283b92d5a2c5cd0b454408c259cdc4c4c47aaecd7fb8dde1584a7a8

Observation 69e0ff21-96d9-40f6-a763-1e2993afd757 · inbound

XFP: Quality-Targeted Adaptive Codebook Quantization with Sparse Outlier Separation for LLM Inference cites this paper.

XFP: Quality-Targeted Adaptive Codebook Quantization with Sparse Outlier Separation for LLM Inference QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:35:04.699300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T21:28:36.358474Z digest=sha256:e9b88cfb82ae328727fddd4c1d0b8d36e87f3e1cc4934ade3c128ea9c6162258

Observation 7afb0a30-8a77-4ba0-ba68-76932eac5c88 · inbound

GAMMA: Global Bit Allocation for Mixed-Precision Models under Arbitrary Budgets cites this paper.

GAMMA: Global Bit Allocation for Mixed-Precision Models under Arbitrary Budgets QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:28:16.800354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T12:25:39.417436Z digest=sha256:f13e4732106040433028d0255771ff83fe0c8ecec8fedcb4d6af46487150a952

Observation 5c711e35-6b5a-40c7-aeaa-e3055fc33baa · inbound

Theory-optimal Quantization Based on Flatness cites this paper.

Theory-optimal Quantization Based on Flatness QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T22:39:10.010380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T22:38:06.888665Z digest=sha256:745982ac726b7471f0d96d6b59864560e241423e489f75fb7b70f21089da3555

Observation 66a62db8-dd20-40f8-84a5-1e764908bff2 · inbound

Llamas on the Web: Memory-Efficient, Performance-Portable, and Multi-Precision LLM Inference with WebGPU cites this paper.

Llamas on the Web: Memory-Efficient, Performance-Portable, and Multi-Precision LLM Inference with WebGPU QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:53:55.288364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T02:52:40.923230Z digest=sha256:005f91678a70e6d4b0dee00a8049ce01a2a5fe7754575b6256a6c7c349fcb935

Observation 9a056b76-2bae-4243-beaf-d2045c7cc745 · inbound

GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs cites this paper.

GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:36:39.854470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:33:06.719954Z digest=sha256:fe6334cbb3bbbd9e57267a8398000366de9f1052154e9cac90548f4c45686cd1

Observation a922941e-80d6-4ce0-b902-4094c5e34ef0 · inbound

MGVQ: Synergizing Multi-dimensional Sensitivity-Aware and Gradient-Hessian Fusion for Vector Quantization cites this paper.

MGVQ: Synergizing Multi-dimensional Sensitivity-Aware and Gradient-Hessian Fusion for Vector Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:05:48.338378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T17:36:45.807397Z digest=sha256:246634b43559a0f5fc84e82f178ec103e60093c36eb47919e7b6ffc001472a03

Observation a829676f-80bb-44d5-9680-467e235be0fc · inbound

EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture cites this paper.

EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:44:45.610682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T14:35:02.484377Z digest=sha256:60f17100879eab2bc3542b27440d2f2a059e5831f358c0153f20738dbd36ba9d

Observation c625abbc-a448-4950-a656-559e37105655 · inbound

Influence-Inspired Spectral Rotations for Extreme Low-Bit LLM Quantization cites this paper.

Influence-Inspired Spectral Rotations for Extreme Low-Bit LLM Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:14:39.052951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T12:13:18.805668Z digest=sha256:4772560978be9a20316c18900f4d688c373853684f5138d9b60721e2313db16e

Observation 66a62b39-59be-4d24-821d-bb8eb1562754 · inbound

Quantized Reasoning Models Think They Need to Think Longer, but They Do Not cites this paper.

Quantized Reasoning Models Think They Need to Think Longer, but They Do Not QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:16:00.108641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T23:05:00.401365Z digest=sha256:ace2a468b5966852db0d7bb248065614988f2773cb0f6be53f1f805c2c2d235f

Observation c13b1c42-78a6-46b0-958b-7e8ed3a4b80c · inbound

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation cites this paper.

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:34.620068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T18:55:51.474956Z digest=sha256:ef13e5b2bc0509dcb7a83dc46cc860b7b1f3ca5b291b20e3c149f25fb2536ac7

Observation c9ce07cf-3804-46ab-85a5-3a24dc1b8fa8 · inbound

LASER: Loss-Aware Singular-value Decomposition and Rank Allocation for Efficient Low-Precision Vision-Language Models cites this paper.

LASER: Loss-Aware Singular-value Decomposition and Rank Allocation for Efficient Low-Precision Vision-Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:12:34.668492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T19:09:29.347284Z digest=sha256:d7dc3298ce95ffa2630075ca135461fab7eaa4cfc5772621a56857032885a3dc

Observation bde79f78-77dc-4e2e-b4b9-1ee44d2b52f9 · inbound

Qift: Shift-Friendly No-Zero W2 Post-Training Quantization for Rotated W2A4/KV4 LLM Inference cites this paper.

Qift: Shift-Friendly No-Zero W2 Post-Training Quantization for Rotated W2A4/KV4 LLM Inference QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:06:15.982040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T15:49:41.836888Z digest=sha256:d67e1cf658d6fc9935fa688695206b61e4cdb66b4210e0b835da965e253f41f8

Observation d3d7e18c-cfd3-49cb-8c9a-7aed78fa2441 · inbound

LiftQuant: Continuous Bit-Width LLM via Dimensional Lifting and Projection cites this paper.

LiftQuant: Continuous Bit-Width LLM via Dimensional Lifting and Projection QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:26.980422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T11:14:03.535306Z digest=sha256:65bbe98e9a951175bc63e98498d6b5ed992e25c0304b050b59b66a6ba8c4a100

Observation d37ff3ed-95ea-4597-a431-61d0017580b9 · inbound

LiftQuant: Continuous Bit-Width LLM via Dimensional Lifting and Projection cites this paper.

LiftQuant: Continuous Bit-Width LLM via Dimensional Lifting and Projection QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:24:38.273775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T11:17:53.736872Z digest=sha256:773732170817cfb78c2624e1c5796a5ff82ffa97f840e236bb56ff22934786a5

Observation 24c04631-cb20-4438-892f-108233728ee6 · inbound

Minimizing the Hidden Cost of Scales: Graph-Guided Ultra-Low-Bit Quantization for Large Language Models cites this paper.

Minimizing the Hidden Cost of Scales: Graph-Guided Ultra-Low-Bit Quantization for Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:16:48.288085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T06:09:42.838355Z digest=sha256:1cc57ca78bf94686ea1b45cbac134108f58bdb8aa30d7c5aa13b2bdb064f63ee

Observation 93ff78ee-bb36-434a-b4af-36eaf89af560 · inbound

FAIR-Calib: Frontier-Aware Instability-Reweighted Calibration for Post-Training Quantization of Diffusion Large Language Models cites this paper.

FAIR-Calib: Frontier-Aware Instability-Reweighted Calibration for Post-Training Quantization of Diffusion Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:55.577768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:35:00.563647Z digest=sha256:9e80c6919d9b910edb33767862c17c6444679840c61474b7db0f7695eda29b2f

Observation c0132601-3a29-4c84-96e1-d0977ae20abc · inbound

FAIR-Calib: Frontier-Aware Instability-Reweighted Calibration for Post-Training Quantization of Diffusion Large Language Models cites this paper.

FAIR-Calib: Frontier-Aware Instability-Reweighted Calibration for Post-Training Quantization of Diffusion Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T12:23:13.879878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:23:13.879878Z digest=sha256:39de70e844c1d8aa327a2bfdfc422f00877c22e2cbcc39c4fe7d6b9c7fe6ce38

Observation 70b88107-7b85-4d66-a397-62ab0c9862d1 · inbound

HyperQuant: A Rate-Distortion-Optimal Quantization Pipeline for Large Language and Diffusion Models cites this paper.

HyperQuant: A Rate-Distortion-Optimal Quantization Pipeline for Large Language and Diffusion Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:45.532847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T08:46:01.500880Z digest=sha256:ec7d53be94ade627180580b1bccacf91288683a4e2443432974260496c5ee68e

Observation 2f57c8e7-0ff3-4570-96ce-312bc4869769 · inbound

GRINQH: Graded Input-based Quantization Hierarchy for Efficient LLM Generation cites this paper.

GRINQH: Graded Input-based Quantization Hierarchy for Efficient LLM Generation QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:39:45.441374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T08:38:32.577228Z digest=sha256:d158b52cabcf1f33762f69e996cd8efec1c7f6c15d306dc8c45e1b5015174140

Observation 75cbd6cf-71ce-45fd-bc84-c577b9e28297 · inbound

Reliability Scaling Laws for Quantized Large Language Models cites this paper.

Reliability Scaling Laws for Quantized Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 160

Resolution
unresolved
no resolver link, observed 2026-07-14T08:45:52.855783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T08:45:52.855783Z digest=sha256:ab596875d86f739165dfb6e36c1e7c31ccf6438468f4454097c88091abf51c8f

Observation 6f7f2d25-95d6-4df6-a7d3-18d90b9d3ac6 · inbound

Cross-Layer Error Compensation and Finite-Sample Feature-Statistics Matching for Extreme Low-Bit Quantization of Large Language Models cites this paper.

Cross-Layer Error Compensation and Finite-Sample Feature-Statistics Matching for Extreme Low-Bit Quantization of Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T01:35:43.851034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:35:43.851034Z digest=sha256:183ff4630f74c291a180b24d26c93da9ed0a2e5abcc46408a3ed4a908d16e26f

Observation 0d47cdbc-5d5d-4238-836f-77290468ea3e · inbound

Break Through the Compression Bottleneck: From Theory to Practice cites this paper.

Break Through the Compression Bottleneck: From Theory to Practice QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T14:29:28.509827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:29:28.509827Z digest=sha256:8444a7191625a83b6aedcbadab855b8cd9e2f5b09b8a78cc245f3ab687f49a52

Observation 829af20d-2a27-471e-a618-badf62ddb760 · inbound

A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference cites this paper.

A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-31T22:36:29.341483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T22:36:29.341483Z digest=sha256:5de11c553b784b109896b5d3077eb7407941d1542dad4c0ff66f4a2e15e791fc

Observation 34d4b659-a33f-4912-8ed3-172af34aa219 · inbound

GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference cites this paper.

GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T03:16:46.931878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:16:46.931878Z digest=sha256:002969cf634a141cc1b5b099d8aa82dc507bbd50e9fe343392ba6daf806811d6

Observation c2ee7334-ce78-4440-80e0-3553412fa3c8 · inbound

TaskPress: Query-Agnostic KV Cache Compression via Task-Guided Pruning cites this paper.

TaskPress: Query-Agnostic KV Cache Compression via Task-Guided Pruning QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T21:59:39.992784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:59:39.992784Z digest=sha256:d31081adce21f4c3ae3b961e8d9654d6f71bd321583e7973a3ff0de0d1cb810a