Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T12:07:30.282490Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 7 inbound Pith citation observations for arXiv:2411.17525.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T12:07:30.282490Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-09T20:34:01.289119Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T10:29:45.503006Z
50 of 50 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8e931ca6-7b20-4c6a-8757-1ed6a701140d · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd9ef3cf-4c20-4eec-bbe1-ccf891b2d264 · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b1e3c6a-8c52-4a8c-8d63-19d25f3a233b · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b39b94d3-7e74-404b-b65e-122b9f7c3a9b · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Qwen Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb61a05d-453e-4ee5-8f08-02196e7215cd · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem QuIP: 2-Bit Quantization of Large Language Models With Guarantees
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3539d07e-54da-42dd-89c9-cf4862405e5f · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b10bcbcf-b2ac-4285-b821-b1c81616c9b6 · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem New Bounds For Distributed Mean Estimation and Variance Reduction
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 25d5f7d4-c903-4678-a3e6-d34c4ec54ae3 · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5077b3cc-9a4a-41df-9db4-ca61e7efb6e8 · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem QLoRA: Efficient Finetuning of Quantized LLMs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30de5264-e0ff-415c-a055-75eb5ceffd8f · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f5544af-0a08-4f33-a644-cc588d639a2c · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c7fc47a2-7626-4146-8fcf-88264558226c · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem The case for 4-bit precision: k-bit Inference Scaling Laws
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0006d9b4-06c8-4e03-b297-605716482785 · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem The Llama 3 Herd of Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff19f2ad-17a8-454f-81c1-bed1deb2ca6f · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Extreme Compression of Large Language Models via Additive Quantization
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e39312b-8a66-43f7-adae-30a2ee1fee76 · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem SPDY: Accurate Pruning with Speedup Guarantees
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b479fa4d-c00a-432f-ad42-855d461521a3 · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bae30d21-2384-42e1-abe4-7ed254df34aa · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c94f01cd-473e-477b-9b60-602e2acb6f47 · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 874cb9a2-565f-427a-82af-fd0c6a853de7 · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem A Survey of Quantization Methods for Efficient Neural Network Inference
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a5f54e7-d482-4d2c-bca2-bcd15536d4c9 · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Fast Matrix Multiplications for Lookup Table-Quantized LLMs
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9c02e29a-74bc-41e1-b3f4-ead42b8a1f00 · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Measuring Massive Multitask Language Understanding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2944c500-c29d-442f-b280-76f842316fb5 · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem SqueezeLLM: Dense-and-Sparse Quantization
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81ab99ba-97cb-4ae9-b6c5-eb9ecfe98788 · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Gonzalez, Hao Zhang, and Ion Stoica
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de9f2010-1148-4651-b1f2-d76f010f89b6 · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem OWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0d16a36-4a7b-417a-ab6e-6c5a67b6e53c · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b52193c-b804-4f98-afe0-6f7ffac0498d · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem SpinQuant: LLM quantization with learned rotations
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7b03015-bed8-46fe-aa28-4d2cd2b48f6b · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Pointer Sentinel Mixture Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e6b4b6a-87e1-48e0-828c-8873a1046dc7 · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a27bc9cb-8365-45a3-a9e0-79a85bc36245 · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation dc93881d-8e8d-4a22-9d54-0eaca56cdc63 · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45c72733-76ad-4009-9971-acdd63cb1129 · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d528b8a-5230-4e4d-970f-026ec6b1af1d · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 333c96cf-9aea-4074-8694-983832cf2a42 · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fc34261-3b22-415d-8261-5c19164091e0 · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Distributed Mean Estimation with Limited Communication
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b3061b82-c5c0-4cdb-8cb3-c682b9622fa0 · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bf461b08-9a23-4f36-ae2f-beaabd37830b · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 17d524fe-736b-4d72-921d-460b2db5530e · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a8598c7-5839-49c6-becd-12b696923a07 · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem QTIP: Quantization with Trellises and Incoherence Processing
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f44b096f-0126-4412-8715-8e70487c6e83 · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem GPTVQ: The Blessing of Dimensionality for LLM Quantization
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37994103-658c-4d4d-b324-61ea3983cde8 · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem DRIVE: One-bit Distributed Mean Estimation
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0226074-94dc-4eba-8178-2a37740906e0 · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem EDEN: Communication-Efficient and Robust Distributed Mean Estimation for Federated Learning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aae0845c-0b73-4211-8c8d-d97f2e12bd39 · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f1f6c9aa-6529-458e-96d3-e75f2c1c536c · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9015d2e5-5189-4571-abb4-2871b0e365e5 · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5c7acba-c9af-4b5d-a774-2f36e9dc24ba · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12283b9e-4a4e-4c96-8559-9caae4e85edd · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem NF4 Isn't Information Theoretically Optimal (and that's Good)
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81f70250-5938-4f27-b9cd-df7b122b723f · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem Unresolved cited work
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bed2eb8-17e9-4f9a-9a1b-1a4dc8253917 · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem OPT: Open Pre-trained Transformer Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 665598d1-cad2-4c77-b593-78ad8b2f88f9 · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem online" 'onlinestring :=
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca5e6263-3184-4fe6-a383-540e4005972c · outbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem write newline
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0f3d9b0-5619-4ccf-a814-6d412e9edcfc · inbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Pushing the Limits of Large Language Model Quantization via the Linearity Theorem
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26d22641-5be1-453b-bcca-af9996668f5f · inbound
CE-LoRA: Computation-Efficient LoRA Fine-Tuning for Language Models Pushing the Limits of Large Language Model Quantization via the Linearity Theorem
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7660a5e6-f3c7-4897-aecd-a5e13d5e036c · inbound
KV Cache Offloading for Context-Intensive Tasks Pushing the Limits of Large Language Model Quantization via the Linearity Theorem
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 69b2aa2a-6198-48c1-b62a-6ea3ecd0af6c · inbound
KV Cache Offloading for Context-Intensive Tasks Pushing the Limits of Large Language Model Quantization via the Linearity Theorem
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation efc361d6-9c2b-4631-9de5-5a2be0490d6e · inbound
KV Cache Offloading for Context-Intensive Tasks Pushing the Limits of Large Language Model Quantization via the Linearity Theorem
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a729e4b6-e5e5-43fa-9169-c18a4b781040 · inbound
KV Cache Offloading for Context-Intensive Tasks Pushing the Limits of Large Language Model Quantization via the Linearity Theorem
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f8aab825-e78f-41ce-a5fd-3a528d9b95e3 · inbound
HyperQuant: A Rate-Distortion-Optimal Quantization Pipeline for Large Language and Diffusion Models Pushing the Limits of Large Language Model Quantization via the Linearity Theorem
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.