Pith. sign in

Paper Citation Record · LEDGER

ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2206.01861.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2206.01861 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:37:51.939462Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:16:16.046751Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5d495910-c8b9-4bfa-b361-25c92364875a · inbound

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale cites this paper.

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 171

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:35:36.185243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-13T13:35:35.972596Z digest=sha256:ac13697a09dc47d697fffa3fb3edb8362d148f15c41fa7ac3bfd32ca48b270f5

Observation 03b4f879-9892-4bcb-abf0-9017dbf482d2 · inbound

GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers cites this paper.

GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T17:18:35.274980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T17:18:35.153078Z digest=sha256:b097246f4058b369981eed7a27c7a31cb2feb4b6288d1fc6b85c9f2fbfdefe9d

Observation 3b7bbd66-67a3-49f0-b4ba-f25606e2f5ac · inbound

Accelerating Large Language Model Decoding with Speculative Sampling cites this paper.

Accelerating Large Language Model Decoding with Speculative Sampling ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:29:36.307343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T07:29:36.205026Z digest=sha256:f9b3b123031553a107e17337406ba85175c1f48ff3583be03cb8de0efd5e95fe

Observation f6589bb8-472a-4d55-b1fb-2a71b309c258 · inbound

QLoRA: Efficient Finetuning of Quantized LLMs cites this paper.

QLoRA: Efficient Finetuning of Quantized LLMs ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:29:53.790343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T13:29:53.345251Z digest=sha256:28f067d2b4c1406226fd265b0c07b21a449a262bedcb7f69a3e834d92e75e04f

Observation 419f3a01-d77c-4b52-9378-02df1fff1d21 · inbound

AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration cites this paper.

AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-24T08:29:11.224966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-24T08:27:35.798991Z digest=sha256:ce7f1f2963f0c3bda6de3fbf503c221f457f14cb5e32e3391763a99a4d083e74

Observation 29d83d7a-d3a3-4d48-b1df-85a29bb200fd · inbound

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models cites this paper.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T18:00:50.460651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:cc014baf04f4bee2d1a6a6d8ebe6abeff88f3d13a7b1294e50c3828c2aa31f21

Observation 405b02e5-3090-4efd-b4dc-2c5bd61b737b · inbound

BitMoD: Bit-serial Mixture-of-Datatype LLM Acceleration cites this paper.

BitMoD: Bit-serial Mixture-of-Datatype LLM Acceleration ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T18:18:05.575588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:18:05.575588Z digest=sha256:e81fbdc0bb0b93a9f5fbfcca92c1389ef6f2777581c6494dfed4bf30cf7331a4

Observation 9015d2e5-5189-4571-abb4-2871b0e365e5 · inbound

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem cites this paper.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.258416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.258416Z digest=sha256:dbf58e4045f7ad29eca6c256ac9e581bfa8c265b5c4776b9b6bcff7f7f0af89b

Observation 4cc0df57-4fc6-4af5-9030-8a7ae42ce0c6 · inbound

4bit-Quantization in Vector-Embedding for RAG cites this paper.

4bit-Quantization in Vector-Embedding for RAG ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T19:12:58.837909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:12:58.837909Z digest=sha256:a6a00515adf70c7f53533250b6eef70648765a8911a3e087128848ed8eb8a658

Observation d1c187c5-2f9a-4c8e-b164-77642734c8de · inbound

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting cites this paper.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.988135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.988135Z digest=sha256:31cbc11eda66cc5a53e10f72aef986cf21fa1fbbb565a123b1c09588a53e3b07

Observation a138f024-ee03-4bc0-97c7-7b1d9e2e4d83 · inbound

PCDVQ: Enhancing Vector Quantization for Large Language Models via Polar Coordinate Decoupling cites this paper.

PCDVQ: Enhancing Vector Quantization for Large Language Models via Polar Coordinate Decoupling ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:42:41.045802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:42:41.045802Z digest=sha256:219fda5a80d049f4bf20de5733f07e7ee0db6215cc3bd6c9d02553cff7e616d8

Observation 934d023b-00d6-45e5-a64e-f1d9d38e76e3 · inbound

Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study cites this paper.

Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T13:22:47.848505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:22:47.848505Z digest=sha256:0cad2fc17ab0d3c339e1d96d5ce34cc30557c1daa28211ea085e1b2f485c9e25

Observation a3f54d90-0c07-462f-b783-f990656e6967 · inbound

Diagnostic-Driven Layer-Wise Compensation for Post-Training Quantization of Encoder-Decoder ASR Models cites this paper.

Diagnostic-Driven Layer-Wise Compensation for Post-Training Quantization of Encoder-Decoder ASR Models ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T17:31:08.335975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T17:30:43.808662Z digest=sha256:f1487fb992fb926a90c0377e8a3cc39c927bd6b228cf9494dac4dabf34486e28

Observation 38c47fdd-5b79-4043-b192-f7df2b5a3817 · inbound

A KL Lens on Quantization: Fast, Forward-Only Sensitivity for Mixed-Precision SSM-Transformer Models cites this paper.

A KL Lens on Quantization: Fast, Forward-Only Sensitivity for Mixed-Precision SSM-Transformer Models ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:25:29.242563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T14:24:53.745784Z digest=sha256:3e682a883cd7dc3f260c19c5ac7015a3b4d03177c1f32ae6bb7f925116182fb9

Observation b3d0d9e1-e4d3-4c83-882e-1e7635f2e040 · inbound

Motion-Compensated Weight Compression cites this paper.

Motion-Compensated Weight Compression ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:04:40.440211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T12:58:27.637822Z digest=sha256:d4a9cc79851ea001a246ff49033ddeb16c76753bad1851792b25f73b1ccb526b

Observation 2687967a-64b3-4698-8ad1-f18a8546b249 · inbound

Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation cites this paper.

Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:16:16.048491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T15:37:34.129339Z digest=sha256:0dc7d04260e726402d5fded8e94ee61077c089f26e347fa92a762f351e1cc694

Observation 6276d88e-1bee-4694-b241-3aaf305918bf · inbound

CubicQuant: Parametric Non-Uniform Codebooks for High-Throughput LLM Inference with 1-8-Bit Weights cites this paper.

CubicQuant: Parametric Non-Uniform Codebooks for High-Throughput LLM Inference with 1-8-Bit Weights ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:51.939462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:51.939462Z digest=sha256:7236d4e054e3b5acc7e342f3144cdad7ed892d76d47c959ad4ae7e86bd60c597