Pith. sign in

Paper Citation Record · LEDGER

ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2303.08302.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.08302 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:57:41.320727Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T00:56:24.897034Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f26d289c-09a8-4f5c-a6f8-ce0ad0956639 · inbound

FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance cites this paper.

FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:56:45.442680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T19:56:45.348973Z digest=sha256:54cfc4cebda133b48ec893a69a87de4309c79ebe8c248ad2e5951302e81b275a

Observation f06b8139-c80f-4fba-aeb4-50b171d3cb12 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 206

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:39:33.241090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:1e7e30c4bb121fc03630001c10d8e3e892251fdeff05aa1dfe702348ef4b98f2

Observation 5ed274c0-c79e-4d6a-8a69-55a310c6f575 · inbound

Enabling On-Device Medical AI Assistants via Input-Driven Saliency Adaptation cites this paper.

Enabling On-Device Medical AI Assistants via Input-Driven Saliency Adaptation ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:57:41.320727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:57:41.320727Z digest=sha256:16db0910a2988a6597d2c3d1b92ef7789e09791d5ba00c73cac1b718a206e3fb

Observation 25797477-e26a-421e-9ef0-e44c03e8bedc · inbound

BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity cites this paper.

BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:00.467026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:20:00.467026Z digest=sha256:3123a30d44a28dcf068f658830ff0fac4d91a0c14a83b4066fe147c927b1ffb4

Observation d49f8fb0-a2f7-488c-890d-ab937df70595 · inbound

Towards Inclusive NLP: Assessing Compressed Multilingual Transformers across Diverse Language Benchmarks cites this paper.

Towards Inclusive NLP: Assessing Compressed Multilingual Transformers across Diverse Language Benchmarks ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T14:11:13.303967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:11:13.303967Z digest=sha256:1559a6b615ae9ef00ae49eae59c65e63098ba79a13ecd14889e332acb1b210b3

Observation ef2bf2c0-6452-4e33-bb6e-4c5c73e30e2f · inbound

Enhancing Model Privacy in Federated Learning with Random Masking and Quantization cites this paper.

Enhancing Model Privacy in Federated Learning with Random Masking and Quantization ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T16:10:53.398370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:10:53.398370Z digest=sha256:7ad7a3383fae754ffed86663b7c2e47fe90ce66f6aefffd32b436594b31e779e

Observation 8566c60f-50ec-4410-94e6-6776a6cd901a · inbound

You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations cites this paper.

You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:50:30.273755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-21T18:46:04.926179Z digest=sha256:9b7304a5efa653256d6ae145b2183529f00a243838f26664ef1ef6e2ecaa8a78

Observation 9777d95e-76ab-4c39-b98f-ae42929a6f48 · inbound

You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations cites this paper.

You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-03T23:21:42.110167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:21:42.110167Z digest=sha256:4df6707dfdb1d6a7b93b32448059a8f6940af97a9b2f62ae9846d9ac4b951907

Observation 310bac23-4950-49a3-a4ed-0b616205c63b · inbound

MCAP: Deployment-Time Layer Profiling for Memory-Constrained LLM Inference cites this paper.

MCAP: Deployment-Time Layer Profiling for Memory-Constrained LLM Inference ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:54:48.415115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T00:54:33.897112Z digest=sha256:8de54f6e2175d1f7f784f675adbcfd28f8a8d316efccf5b95b7f79a7a84a1552

Observation f05640c3-6400-4fc2-b46b-aacd0f9b81f3 · inbound

DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices cites this paper.

DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:36:19.950190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T03:36:12.915133Z digest=sha256:456732089ce7b7189df404b86b9564f5995aeb721376b77c7389ffc56a959f55

Observation 03866696-993d-496c-93ff-04e990ad6929 · inbound

DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices cites this paper.

DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:32:30.196922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T07:29:14.545746Z digest=sha256:c8f52b5a4e1c518f32baafded64b160bf37c0efe8d3cd3fb9c56484defa95f90

Observation 4246816d-2079-4874-b669-88e413c3a234 · inbound

DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices cites this paper.

DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T07:59:50.166136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-21T07:57:49.746594Z digest=sha256:fe8cbfe00976dbc6e6de0b55217b79858003f27c44529dae67168ef2d89607b5

Observation bc7bccb3-9cea-42ea-8cce-3c2bf492175e · inbound

Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models cites this paper.

Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:23:03.633191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T05:20:45.264341Z digest=sha256:51a0b2f9dd2586814691467c8fcddac85e4779576e857547b2e46b6eb58bfece

Observation 9e9c1558-3a9b-4696-8b17-2112027836db · inbound

ProjQ: Project-and-Quantize for Adapter-Aware LLM Compression cites this paper.

ProjQ: Project-and-Quantize for Adapter-Aware LLM Compression ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:12:34.512202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T19:10:58.845006Z digest=sha256:c36939f3b72d9135b84d30c4c690ee355f78b5de226a60ca135157808429af23

Observation 07652548-2994-4d7f-9f53-23a64a978848 · inbound

GPTQ-intrinsic LoRA: A Near-optimal Algorithm for Low-precision Quantization with Low-rank Adaptation cites this paper.

GPTQ-intrinsic LoRA: A Near-optimal Algorithm for Low-precision Quantization with Low-rank Adaptation ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:06:14.396687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T17:28:14.160341Z digest=sha256:9736c2766b20fa7bed5db58f841ce5c4f98148003f933def364e6b38afa8cea0

Observation c504d84a-0b2d-4614-a750-898e16cfb00e · inbound

TwinQuant: Learnable Subspace Decomposition for 4-Bit LLM Quantization cites this paper.

TwinQuant: Learnable Subspace Decomposition for 4-Bit LLM Quantization ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 100

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T00:56:24.899516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T13:07:53.001390Z digest=sha256:13ab329672b5366076745fceea464624286fdb36e68d1b5a87ea6f7987b57c18

Observation 66defc3c-1c96-409c-be9e-3130e323c914 · inbound

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems cites this paper.

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T03:55:41.125011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:55:41.125011Z digest=sha256:c1600a5f74ad99822a6843882879a15558022fb2a70cfc4b96e710cb54468821

Observation 5b0249f8-bee6-47b8-9021-6f42da1652e7 · inbound

GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference cites this paper.

GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T03:16:46.948303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:16:46.948303Z digest=sha256:343d5317a26ab2e285fbfb3c0d1aa86907f212b2ae78c82aebd9e8bedbdb0b99