Pith. sign in

Paper Citation Record · LEDGER

WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2402.12065.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.12065 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T20:34:01.368235Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7761c196-453c-4558-84f1-8e046a244107 · inbound

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models cites this paper.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.368235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.368235Z digest=sha256:d9f7af4f4083e895a965be2a1a22318c585b84f35f74e759c2d7df90b33d4a75

Observation 16fc40a5-c601-48e2-9250-23b0f845db06 · inbound

MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization cites this paper.

MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-09T19:12:56.315307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:12:56.315307Z digest=sha256:03d9e05f5706f7f84614a55a261621ade613b5da11538e0edbff8e159c22a4c8

Observation 9a149383-d99e-4dbf-a631-492662798daf · inbound

PolarQuant: Quantizing KV Caches with Polar Transformation cites this paper.

PolarQuant: Quantizing KV Caches with Polar Transformation WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T13:26:55.458763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:26:55.458763Z digest=sha256:83a07e627388d8c9a552e62250285bb4c60fbec4b3d3cbf80ca28dbba93420e9

Observation b4b0e042-85be-4feb-a2be-b6c1509ba365 · inbound

Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents cites this paper.

Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-08T14:13:38.457867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:13:38.457867Z digest=sha256:32c6b09e88b0c2c954442771ec59ff3cc018e03a77968c6f6f534471ea6b5e20

Observation 02cb7c13-5fe1-446d-8bab-00c05bb8e8dd · inbound

TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate cites this paper.

TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-20T08:09:22.393487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T08:09:22.226608Z digest=sha256:c0bdd3ebbfe9981ab5177046a290761f104a012a37c40fc7ad66819a5cb1dd7a

Observation c81bd392-b7a7-49c4-974a-27b5949d74d5 · inbound

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs cites this paper.

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:27.199163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:27.199163Z digest=sha256:41c70cfcb0b52ac56a79c374dded254d3a83309c4242d32157940d0a68f73d7c

Observation 0e775c07-916f-4896-ae0c-6090247f3f33 · inbound

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression cites this paper.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:19.784103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:19.784103Z digest=sha256:6ad9babc56aa2260df5e5ec1c417198a4b84b2352036ebcb6d8f7e431eb49ea3

Observation 07300448-297a-413c-a67d-91aef5f612f8 · inbound

PCDVQ: Enhancing Vector Quantization for Large Language Models via Polar Coordinate Decoupling cites this paper.

PCDVQ: Enhancing Vector Quantization for Large Language Models via Polar Coordinate Decoupling WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:42:41.053905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:42:41.053905Z digest=sha256:b3affe32dc7937311ebafb6b9a6a90454647fc8c5b4560be9f3d51393a9460fa

Observation 50122e54-e669-4aa5-b2b0-b6048149c522 · inbound

OpenGCRAM: An Open-Source Gain Cell Compiler Enabling Design-Space Exploration for AI Workloads cites this paper.

OpenGCRAM: An Open-Source Gain Cell Compiler Enabling Design-Space Exploration for AI Workloads WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:30:10.324505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:30:10.324505Z digest=sha256:d38233acca12aaa41e376621ae7ea68086f3b628f79e6eac0b9681b3a859ca81

Observation 595fb424-d5e9-4355-97ff-ad0f54ba2c4b · inbound

Which Heads Matter for Reasoning? RL-Guided KV Cache Compression cites this paper.

Which Heads Matter for Reasoning? RL-Guided KV Cache Compression WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T10:50:10.957901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:50:10.957901Z digest=sha256:1ef71d9954702b18d7bf3aae88c4c52d6228417f242afa83ae479ab668e8629d

Observation 41c1dd36-b19d-45ca-b8da-e2768d717baf · inbound

Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference cites this paper.

Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T04:37:29.030814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:37:29.030814Z digest=sha256:c53560734d8fef8539813e3208ec8e3df5a188b0d80d342e0566e8f27db323ca

Observation ec06f206-a996-4b6d-81e3-714dd2922dd0 · inbound

Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization cites this paper.

Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-09T23:04:17.325023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T23:01:56.712670Z digest=sha256:89b099c95cfda41d4804b3ec87e5930cad4a05e2102b02070e42ab51a992622c

Observation 3fcf9aaf-ed29-432e-943b-9c150db5dcd2 · inbound

OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization cites this paper.

OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:30:44.202534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T18:23:14.935801Z digest=sha256:0095c8328c4f21df90e283d1e2895fe4d43199785c3da0ae40d68337e6187f90

Observation 2c33c427-3e9b-42fb-ba8d-d75426045cb6 · inbound

OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization cites this paper.

OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:01:18.181638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T02:59:00.997742Z digest=sha256:dd1af003d9729c557dcf237645743c48190d3ab95b24aeb956ee2a57e3e210a6

Observation 1a3f28e5-04e7-48ed-9f91-e573cb8ab6c3 · inbound

Search Your Block Floating Point Scales! cites this paper.

Search Your Block Floating Point Scales! WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

Reference 168

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:24.100491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T05:52:26.558984Z digest=sha256:5ca123cbda687327222fc7364046747afd538a54f70796d7f873a4d784a86c8b

Observation 7ecb6e09-3a5c-405a-939e-25e31b5be9a4 · inbound

OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization cites this paper.

OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:18:16.543015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:16:07.797702Z digest=sha256:5d68e5653e200c1dd14284bfb47042008d3e97654a7b440f7313e9057d83886c

Observation 28a54ca8-4bad-4e1d-b588-1e7a1e20bc1f · inbound

HACK++: Towards More Effective Head-Aware Key-Value Compression for Efficient Visual Autoregressive Modeling cites this paper.

HACK++: Towards More Effective Head-Aware Key-Value Compression for Efficient Visual Autoregressive Modeling WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:27:24.336242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T19:46:43.514413Z digest=sha256:aa6fda01dc2c2b7470fd892253d14c7c201a7de9f2bdf20a59e673b94d661d46

Observation 85915596-fc29-44b7-9cc9-dcaf28810ac3 · inbound

PhoenixRepair: Rethinking Repair Strategy Exploration in Software Agents cites this paper.

PhoenixRepair: Rethinking Repair Strategy Exploration in Software Agents WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-01T14:12:00.997267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:12:00.997267Z digest=sha256:c437a2981c33cf80bdb1fd221f34be2bfcdcbb39d45d2203ccedb4728366b75d

Observation 491d83c7-113f-450b-98cc-b1b3f55df3f2 · inbound

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs cites this paper.

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:55.630826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:43:55.630826Z digest=sha256:c64a93227efbf90980e35cba55754f759bad367c8ff22b094f9a91af6be9f158

Observation dfa44a68-0b6e-43d8-bb37-b93d3f95749f · inbound

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs cites this paper.

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T00:11:49.059203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:11:49.059203Z digest=sha256:3cc0f9ba1ceb42099bfc07ab298c5cc2f4064f6e375e4f62f1c32d6d84b9b546