Pith. sign in

Paper Citation Record · LEDGER

ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2206.01861.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2206.01861 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:42:41.045802Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:16:16.046751Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5d495910-c8b9-4bfa-b361-25c92364875a · inbound

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale cites this paper.

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 171

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:35:36.185243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T13:35:35.972596Z digest=sha256:eea2254a6ecd0309af00190db23398ab65be0496c3d625cd4b6cd9175b96e508

Observation 03b4f879-9892-4bcb-abf0-9017dbf482d2 · inbound

GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers cites this paper.

GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T17:18:35.274980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:18:35.153078Z digest=sha256:31285cbcbbb842522bb6fea10a7b93a855cbdeac231a8e7e73e7fb159bb22c3c

Observation 3b7bbd66-67a3-49f0-b4ba-f25606e2f5ac · inbound

Accelerating Large Language Model Decoding with Speculative Sampling cites this paper.

Accelerating Large Language Model Decoding with Speculative Sampling ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:29:36.307343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T07:29:36.205026Z digest=sha256:f9716ef309d265d53da3092a93352fb427848d81cc0080cb8fbf12d4388da1c0

Observation f6589bb8-472a-4d55-b1fb-2a71b309c258 · inbound

QLoRA: Efficient Finetuning of Quantized LLMs cites this paper.

QLoRA: Efficient Finetuning of Quantized LLMs ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:29:53.790343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T13:29:53.345251Z digest=sha256:9bd272ced386fe3edda00f72ca833d973417fac53093cd69c6916306a2f703cb

Observation 419f3a01-d77c-4b52-9378-02df1fff1d21 · inbound

AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration cites this paper.

AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-24T08:29:11.224966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-24T08:27:35.798991Z digest=sha256:a14cdbd2b8afce6faf2b34657565b9c3d9c95a301aedc4395a1f2e111c5ff177

Observation 29d83d7a-d3a3-4d48-b1df-85a29bb200fd · inbound

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models cites this paper.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T18:00:50.460651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:49ffbef25093c297c5aacd72ae87cab3ef780f1b5cbb089508ea3faaa7b5d91c

Observation a138f024-ee03-4bc0-97c7-7b1d9e2e4d83 · inbound

PCDVQ: Enhancing Vector Quantization for Large Language Models via Polar Coordinate Decoupling cites this paper.

PCDVQ: Enhancing Vector Quantization for Large Language Models via Polar Coordinate Decoupling ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:42:41.045802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:42:41.045802Z digest=sha256:4f3a5c758cf0d07c492d86dde6adfe997c7dff402bcf683383396535b362fcc1

Observation 934d023b-00d6-45e5-a64e-f1d9d38e76e3 · inbound

Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study cites this paper.

Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T13:22:47.848505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:22:47.848505Z digest=sha256:a69f5318c0ed12db1a02c5ea80374cbf94d814a0c8d8f0266e912791f94474a7

Observation a3f54d90-0c07-462f-b783-f990656e6967 · inbound

Diagnostic-Driven Layer-Wise Compensation for Post-Training Quantization of Encoder-Decoder ASR Models cites this paper.

Diagnostic-Driven Layer-Wise Compensation for Post-Training Quantization of Encoder-Decoder ASR Models ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T17:31:08.335975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T17:30:43.808662Z digest=sha256:cc8e07609e5b9b1a12e41da860749f7897d26b980c00cac7388781333d3ad183

Observation 38c47fdd-5b79-4043-b192-f7df2b5a3817 · inbound

A KL Lens on Quantization: Fast, Forward-Only Sensitivity for Mixed-Precision SSM-Transformer Models cites this paper.

A KL Lens on Quantization: Fast, Forward-Only Sensitivity for Mixed-Precision SSM-Transformer Models ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:25:29.242563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T14:24:53.745784Z digest=sha256:5a6c335f9dead55342222c15809e9c21706878606d9f2498087fc072d1b69d99

Observation b3d0d9e1-e4d3-4c83-882e-1e7635f2e040 · inbound

Motion-Compensated Weight Compression cites this paper.

Motion-Compensated Weight Compression ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:04:40.440211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T12:58:27.637822Z digest=sha256:6aa485e8d94999e0de3b81318fa6dbdbd74c03dc19653875164fd1c65112851a

Observation 2687967a-64b3-4698-8ad1-f18a8546b249 · inbound

Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation cites this paper.

Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:16:16.048491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T15:37:34.129339Z digest=sha256:29ce79b2af68c7e778f8350eda94a564b3d517258a01d74e91699ff0b51419e7