Pith. sign in

Paper Citation Record · LEDGER

KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2405.03917.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.03917 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:52:48.819926Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T01:36:44.144655Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9578d8f9-5a7b-4e17-a3bd-3178a851a30d · inbound

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression cites this paper.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.819926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.819926Z digest=sha256:973ccac970119b8dbad7e285bda220e5ce42121f7c47357c784460eea6883397

Observation 94352534-5c7c-4e04-bc9b-602f81faec4d · inbound

PolarQuant: Quantizing KV Caches with Polar Transformation cites this paper.

PolarQuant: Quantizing KV Caches with Polar Transformation KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T13:26:55.470592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:26:55.470592Z digest=sha256:7f2856d48a9b1a129d5e96bb843c938993f9c3e9b0ee1bb5b02e9d05b98e9dc6

Observation 0c04dabc-3b5d-43dc-a42e-79054b65ac2a · inbound

TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate cites this paper.

TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-20T08:09:22.411208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T08:09:22.226608Z digest=sha256:561f30fda25ebe1b3c243c1938e1158e69527bf949b886be2aa9e9c6def44988

Observation 0aef9934-a407-4c49-8162-428eb3976373 · inbound

CaliDrop: KV Cache Compression with Calibration cites this paper.

CaliDrop: KV Cache Compression with Calibration KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.821815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.821815Z digest=sha256:e853b93ae675ab95b57133b0a3e810cb4625006fa390fc72bfc120f0c5475eec

Observation 5150f09a-df9f-40a6-a18e-a57a59854d2c · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization

Reference 148

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:44.146556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:c05fdf5df75dcdbc05843a13c7333d29eb5af6e574536a4602fc3f93af218f5e