Pith. sign in

Paper Citation Record · LEDGER

KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 91 inbound Pith citation observations for arXiv:2401.18079.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.18079 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 91 of 91 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 91 of 91 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:01:20.089202Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

9
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 49358396-5a24-42bf-a26e-8110a97c2a08 · inbound

ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models cites this paper.

ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T13:49:33.837276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T13:49:33.747672Z digest=sha256:63b207c2327f792112810c4ed45b1c0d6a5bb1c3633d46fdc189f4fa18bbe337

Observation f0ce2c05-fd48-4cd7-8700-7480987707f1 · inbound

SGLang: Efficient Execution of Structured Language Model Programs cites this paper.

SGLang: Efficient Execution of Structured Language Model Programs KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:20:01.106533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:93d951ae1a6320a2c21a404c585401207d16a82ba1928ad5dc92a02f758313a4

Observation 1d83d1f5-4133-4531-ba96-1b93e15fbeb0 · inbound

KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache cites this paper.

KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:53:12.320304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T08:53:12.253243Z digest=sha256:29a79ee3ed7a26577850e37e937c1c8c5d320e0925e0ed7bd6c0a1d4566b3f96

Observation 8d622d52-b7d1-44e9-8a56-00fc999812ca · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 219

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:39:33.280738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:7c4cec07d0c6f050772cf443c77d723885663c00cb4a5f82a72dd328e30f1ad1

Observation b6d19c84-6bc0-4145-8471-c17e03e40217 · inbound

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model cites this paper.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:36:26.480168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:0220304a5a78a9a4f5200c36572c64b2768e99d954283265c328051cbc6d300b

Observation b8ca3e16-9a01-4171-917a-ef22f8c98e96 · inbound

FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision cites this paper.

FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T19:45:36.452789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:45:36.337956Z digest=sha256:93a317f0169f881ba675e499cfa00badb2ee8b81fcfdb902e49f9d13e735ea51

Observation e31e43a2-f16c-464b-a096-57ffd31fb592 · inbound

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format cites this paper.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.037498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.037498Z digest=sha256:f2693c9f477c1bb074d1594cdb10864921e267faabc3d4d1d59d90ea209e6bff

Observation 6070763a-9f62-4b58-96cd-0a12f6c51a35 · inbound

CLOVER: Cross-Layer Orthogonal Vectors Pruning and Fine-Tuning cites this paper.

CLOVER: Cross-Layer Orthogonal Vectors Pruning and Fine-Tuning KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-12T12:13:55.213100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:13:55.213100Z digest=sha256:98c46fb6c4f0790249b6f822a47e8fe96c63cd790335ab50e6b1800e69086896

Observation 2960792d-8304-48a2-8529-904342be526d · inbound

APEX: An Extensible and Dynamism-Aware Simulator for Automated Parallel Execution in LLM Serving cites this paper.

APEX: An Extensible and Dynamism-Aware Simulator for Automated Parallel Execution in LLM Serving KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:23.547196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:23.547196Z digest=sha256:8eb98695f462360e724daf3479979a873a5dc186ff55b2176cb93c7f9d2b7228

Observation 4cf2ff58-02f5-4c4c-91b4-2c10cc73130d · inbound

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache cites this paper.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.577744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.577744Z digest=sha256:d53e0862151f57420e2c2d9c94ae6d54b0ac94db662801cb593023c51d8bce60

Observation 9b52e2c3-d4b1-4afa-9cdf-4999171ba5d8 · inbound

Scaling New Frontiers: Insights into Large Recommendation Models cites this paper.

Scaling New Frontiers: Insights into Large Recommendation Models KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T05:10:44.300444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:10:44.300444Z digest=sha256:c319f68d656fc4f2d6fd70ea71345a5a83f8bd58118261834e0ed09225b0c035

Observation c18c02e2-eb94-4759-aa0c-59bbb2876a97 · inbound

DiffKV: Differentiated Memory Management for Large Language Models with Parallel KV Compaction cites this paper.

DiffKV: Differentiated Memory Management for Large Language Models with Parallel KV Compaction KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:12.078529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:12.078529Z digest=sha256:9e1deb976added72abc6a76e90b4f01cf2f18ffbf5f08eba2757dfd24ceb0e7e

Observation f1f659a8-005c-48d9-b736-fb859617fdb4 · inbound

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference cites this paper.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T20:18:18.267468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:18:18.267468Z digest=sha256:acd1ed4f967162121f6e28d25a72c2782286916fbbc9f968d711445c3295be4d

Observation ad012494-5137-4507-bdda-9643a54b41c4 · inbound

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries cites this paper.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.947554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.947554Z digest=sha256:3510d05c4c8c4b21ae9410c398077578e6ed4a6bb7c0d675da50ff064555c198

Observation c0a4b25d-8ab6-4d48-a875-382b2a14a53b · inbound

CSR:Achieving 1 Bit Key-Value Cache via Sparse Representation cites this paper.

CSR:Achieving 1 Bit Key-Value Cache via Sparse Representation KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T14:43:53.736238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:43:53.736238Z digest=sha256:5a5fef0342ed647bd62ec338e494d3bdd8395d66860275ce13e050dbf3c5f110

Observation 3535e96e-a105-4684-b7c5-8b957cf300b8 · inbound

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression cites this paper.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.686945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.686945Z digest=sha256:35aea86f152344d725335c47e0043ab74be8edaffcddb211f8868f69ad4f4ad0

Observation 6aab0856-dddc-441b-bce0-76553e2988e1 · inbound

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals cites this paper.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.838579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.838579Z digest=sha256:5e97d7fd0d9c7ca7247ccd151ccd4efa1d7662b0e10e43b479564ebb1eb02d68

Observation dcde423f-c079-4156-aff8-96a24c3c45aa · inbound

DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs cites this paper.

DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T11:57:17.920808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:57:17.920808Z digest=sha256:f0c780d46bd8ba9d8e21b60e990286dbcbcff33133e2afc06ee49776194e919f

Observation 66bd7130-778b-4b9b-b5f9-bb1cb30f5782 · inbound

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing cites this paper.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.162332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.162332Z digest=sha256:4f81cf3b3baa74c881a7db3db79636f5c9427f36f0946057a7d07368ca8d2c97

Observation cd871bcf-fddd-495a-a514-daf8f60c2f4b · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:47.289879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:47.289879Z digest=sha256:7eb37ce54f25322af8fc8b5c541929ebb4bdd50ac9a3773f00985bf7898574e0

Observation e256fefb-690f-4ec1-9063-af310016def5 · inbound

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference cites this paper.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.120356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.120356Z digest=sha256:3cc6a45ecda24be25310c9f1fbddda461f8edcacc0b6ffd7f5bb608126744b10

Observation bc330e45-6dd7-4f4d-b307-5b56cc54e84d · inbound

Qrazor: Reliable and Effortless 4-bit LLM Quantization by Significant Data Razoring cites this paper.

Qrazor: Reliable and Effortless 4-bit LLM Quantization by Significant Data Razoring KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:57.440237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:57.440237Z digest=sha256:bd984e190540ec27b9070b7f8fe56c53059e675ac3e90930a20eecd4284fd1cf

Observation 75a3079e-0cc1-4050-8109-c0aa5f96f701 · inbound

KVDirect: Distributed Disaggregated LLM Inference cites this paper.

KVDirect: Distributed Disaggregated LLM Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:14.581366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:14.581366Z digest=sha256:7d3bad19dd872c2d5075e924328ddc67802a848114c68c2509b04050666d2da2

Observation e98dbfdb-7b16-4efd-8aca-1176e80b66db · inbound

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models cites this paper.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:02.636154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:02.636154Z digest=sha256:59534476df1b24b4d485e43561f10c4e1362148a159ce78a3d6b7e194fd67e00

Observation f0bef686-892a-42fb-bace-ec0f75292a9d · inbound

RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations cites this paper.

RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 1978

Resolution
unresolved
no resolver link, observed 2026-08-10T14:54:22.656654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:54:22.656654Z digest=sha256:affee5f3fa7dddfad7d6f711c3fbc06f9ead7f3b13e8c7b1d1314ede4d0d4821

Observation 63704b6d-d0da-4d2f-9966-e36b374e5998 · inbound

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models cites this paper.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.238595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.238595Z digest=sha256:d9de7a03588af9805f4e8442f0129910d5ef91e8b6d9f398bb52a4a92d07883c

Observation 0e95d2d7-7b5f-448a-98e9-7cc2b1147221 · inbound

PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration cites this paper.

PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T18:46:51.063278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:46:51.063278Z digest=sha256:1b7a53defbf5be5a49678ccd2fc346da753d63510985d1c619cc7a9d90c44155

Observation 0d86c255-4a62-48c3-adae-804db6dd5993 · inbound

PolarQuant: Quantizing KV Caches with Polar Transformation cites this paper.

PolarQuant: Quantizing KV Caches with Polar Transformation KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T13:26:55.377329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:26:55.377329Z digest=sha256:35c1c08fef393a736311ad509562ce5ff0ce217e796a5ba145ae63c8eefe37a4

Observation 457f745f-79a5-4965-b2cb-17a66955aeb7 · inbound

CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective cites this paper.

CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T00:45:56.852280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T00:45:56.852280Z digest=sha256:10bfb3650782879edd16c8e981658682fb736977e93539054f0003f74d1fd2bf

Observation 9fe47305-c1fb-4cdd-a013-2f6bdd13060f · inbound

Systematic Outliers in Large Language Models cites this paper.

Systematic Outliers in Large Language Models KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T15:37:37.470136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:37:37.470136Z digest=sha256:bffa91dab654fa319ec36264fcb8f8c1b83039b7f3ffde26b9890e1207840422

Observation ecda896b-6e75-443b-8092-fc93d81d32bc · inbound

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models cites this paper.

EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:53.780997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:48:53.780997Z digest=sha256:9d653b499c22f0d784a2b68126be931fbc7dd721646ece5c0e943c4b705ab5a9

Observation 2dbfab0e-01dd-4f32-8e8e-538cd0110b33 · inbound

Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents cites this paper.

Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T14:13:38.095954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:13:38.095954Z digest=sha256:2f7a6027ab23f1dfe1d93d2fffd59fe20e4f9803b0e272c52d5ce6b9291081ab

Observation 238d4a53-ec19-4a44-95f3-f91b2843baa0 · inbound

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU cites this paper.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.838100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.838100Z digest=sha256:f5673e8160debc570fea5d3c66d0f920fa7d9fa4b777736d850ba25b4ef9bf68

Observation e846879e-f25e-4d65-aa78-a5774b046e17 · inbound

Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation cites this paper.

Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T05:40:21.484068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:40:21.484068Z digest=sha256:ad3447f8f058813df4862df50fac84a443e2b55f909ec004a8bc526fa74c2fce

Observation 7b6d9095-11a4-4ee2-aac3-bd77a67b6b77 · inbound

TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate cites this paper.

TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-20T08:09:22.397439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T08:09:22.226608Z digest=sha256:48da4b25d2b21a5f4782587946cdf9bb562d51af94949e472098504a8da2e21a

Observation e648adde-0231-4344-914a-9972409ac979 · inbound

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design cites this paper.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.089202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.089202Z digest=sha256:fe3a4385aa213b551ed92a264d926defb58f7045788f840df86b073e9e203c79

Observation 1bc47a00-7d92-4512-be9c-e6aea921184b · inbound

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression cites this paper.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.257055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.257055Z digest=sha256:708b2ee8366c9f2b3fffb06d81cddecb19579cbeaa86a31b240695340f9bb952

Observation 070ec353-da31-46bf-9097-ec604a14edf7 · inbound

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications cites this paper.

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T22:27:53.540334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:27:53.540334Z digest=sha256:05062f25a89bafa77c13049cf6d688f9f73c7de725d7ad309d69b2cc2ea25ee2

Observation a833d507-3149-4d46-84b6-22eb6fd5b50f · inbound

FastCar: Cache Attentive Replay for Fast Auto-Regressive Video Generation on the Edge cites this paper.

FastCar: Cache Attentive Replay for Fast Auto-Regressive Video Generation on the Edge KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T20:54:59.293772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:54:59.293772Z digest=sha256:726383e41c20fe17113e6ad6bc549a082460855c5e2dcc95c3376c7f37bb713c

Observation 626ad239-cb82-4223-92e2-612a9d4ef34a · inbound

Titanus: Enabling KV Cache Pruning and Quantization On-the-Fly for LLM Acceleration cites this paper.

Titanus: Enabling KV Cache Pruning and Quantization On-the-Fly for LLM Acceleration KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:43:24.481761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:43:24.481761Z digest=sha256:a6caec96ba377715c467c5cddc9755f5583186dd1dd4941b77dee999c9ac3407

Observation 061c7b6a-8acf-416f-8cd8-aeb6b27e50cf · inbound

Win Fast or Lose Slow: Balancing Speed and Accuracy in Latency-Sensitive Decisions of LLMs cites this paper.

Win Fast or Lose Slow: Balancing Speed and Accuracy in Latency-Sensitive Decisions of LLMs KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:45.310548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:45.310548Z digest=sha256:856bc5550e7731b511db315ae9ae76d3a5db62e78b354972169aaa762cadc164

Observation b846c7a3-8c26-4820-9f30-fd34b7596dbf · inbound

Hardware-Efficient Attention for Fast Decoding cites this paper.

Hardware-Efficient Attention for Fast Decoding KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:30.954664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:30.954664Z digest=sha256:2b3a140733689f9d574190f2f981641d68d73c34daa1006b400d289f3f32cb48

Observation 19c23bce-5875-44f6-92fb-259f5d75ee64 · inbound

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference cites this paper.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:59.316579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:25:59.316579Z digest=sha256:6f03701f5851f998954fa19e937ea34ca4ab51caf53d7ecfc751032a5051fe1b

Observation dec87567-bce7-4b01-9448-54d41b907068 · inbound

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering cites this paper.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.862820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.862820Z digest=sha256:23c0ff24556c6b066098bac7801712e716be878da9242be1e02e83a901f38c2a

Observation 683195ac-343c-4526-8a60-3fb774b83e18 · inbound

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library cites this paper.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:32.875169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:32.875169Z digest=sha256:e8d082cb7e25f17d3725780489239b1b7b21431c2de74e39d90998352e2ca24b

Observation f8f7e5ad-7493-4c4b-bf6d-c01e6c2334aa · inbound

Beyond Homogeneous Attention: Memory-Efficient LLMs via Fourier-Approximated KV Cache cites this paper.

Beyond Homogeneous Attention: Memory-Efficient LLMs via Fourier-Approximated KV Cache KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T01:10:38.223015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:10:38.223015Z digest=sha256:df590cfad50fbd4c2c62e832fe4114dfae2b5b0cef8491bbd8b694ecd6048fe9

Observation 40c2c40e-455f-493a-aec0-348b59a97461 · inbound

MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference cites this paper.

MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:20:41.660846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:20:41.660846Z digest=sha256:7f643572a013cd09395d1dd319fffbb02783dabc28ad018b9d3a61548629cee3

Observation 2baeb153-49e3-4b37-908b-b498371ecb17 · inbound

CommVQ: Commutative Vector Quantization for KV Cache Compression cites this paper.

CommVQ: Commutative Vector Quantization for KV Cache Compression KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 1984

Resolution
unresolved
no resolver link, observed 2026-08-15T18:49:24.074952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:49:24.074952Z digest=sha256:447ad88e77da2717ab21d6f0529d67ec1c280fe27b1a492cf2c1c57426f1f4ec

Observation 19c85dd1-eaf2-46ec-b9ce-a79b014b1e82 · inbound

Breaking the Boundaries of Long-Context LLM Inference: Adaptive KV Management on a Single Commodity GPU cites this paper.

Breaking the Boundaries of Long-Context LLM Inference: Adaptive KV Management on a Single Commodity GPU KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:00:11.077208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:00:11.077208Z digest=sha256:4df5908c36b9996cd1deebb4879ce1c8c07907fd141f47f6d90e62d3091f5591

Observation 53b5b577-79bb-4a38-8752-585dbd12bc91 · inbound

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs cites this paper.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:33.153280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:33.153280Z digest=sha256:104e76f0bcc99e57bbb7ed052a4a47d5344785d257711bf6f511303c6710bee0

Observation ffc9ce03-1f66-45f5-a56c-51d090804ebd · inbound

CaliDrop: KV Cache Compression with Calibration cites this paper.

CaliDrop: KV Cache Compression with Calibration KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.710581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.710581Z digest=sha256:e85423e63e52df93986b439472173a38c9a92a4e637fff4718f8cfebe6b7ece8

Observation b78e0d83-6f47-466f-8038-3a25652de897 · inbound

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration cites this paper.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:33.951900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:33.951900Z digest=sha256:f57379c605a8b02e6e6399c7ae1ababcae257a1a35c878dc4fa25025cfe8d91f

Observation 30df9cd3-da21-4b66-94cc-6e529e1f16c4 · inbound

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference cites this paper.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:10.035877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:10.035877Z digest=sha256:4d13bd7839a5be488adf72988031ea7bb26d357c31d344bc7d68f9198ccecefc

Observation 692dbdfc-6c73-4425-9967-499b6355cf2b · inbound

PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference cites this paper.

PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T10:16:13.074677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:16:13.074677Z digest=sha256:eea8fc9dd20bf97c1140090e7f01d3a2731363e141449650a02ab8cc45a7c7e0

Observation ed0ab0ff-6081-457f-8c89-c47ef52b9221 · inbound

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs cites this paper.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:54.973004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:54.973004Z digest=sha256:99d4b35d1b9fee79e943977ec0c33b7af0fe8fc240400bc4a59d0c41ba9801f1

Observation 983a32bc-107b-44c9-a3e0-41e1c4621837 · inbound

EpiCache: Episodic KV Cache Management for Long-Term Conversation on Resource-Constrained Environments cites this paper.

EpiCache: Episodic KV Cache Management for Long-Term Conversation on Resource-Constrained Environments KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T21:54:22.821265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-21T21:54:05.442769Z digest=sha256:02c9b2807c31cb1aa24148efe4e768fedc75653c6d49024200320653664f0608

Observation 15760c83-790d-4e7a-b65e-b4d4e8dcae70 · inbound

Token Sample Complexity of Attention cites this paper.

Token Sample Complexity of Attention KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 1963

Resolution
unresolved
no resolver link, observed 2026-08-03T17:10:08.496622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:10:08.496622Z digest=sha256:896474be84434d6c9fca56f200333d00d40404356dca76e336c64c7ac15dd157

Observation d8e2c0b5-458f-4baf-b766-d3d418fc5841 · inbound

SR-TTT Does Not Learn Retrieval: A Correction and Mechanistic Post-Mortem of Surprisal-Aware Residual Test-Time Training cites this paper.

SR-TTT Does Not Learn Retrieval: A Correction and Mechanistic Post-Mortem of Surprisal-Aware Residual Test-Time Training KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T20:43:37.256381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:43:37.256381Z digest=sha256:eb2522a817bff2774e383e2bc9ec0dacad1f2181f1e9ef0c1238adb82b62221b

Observation d5d4f3d5-68a8-4f87-9c2e-6f24aedf0a2d · inbound

Sequential KV Cache Compression via Probabilistic Language Tries: Beyond the Per-Vector Shannon Limit cites this paper.

Sequential KV Cache Compression via Probabilistic Language Tries: Beyond the Per-Vector Shannon Limit KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:11:05.119893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T16:46:14.866351Z digest=sha256:e8d707ee5356bbf499c345e9b9af09159d9804289747232ef5657257007dd392

Observation cde0251b-0d4b-4c31-bb72-819a7691390c · inbound

PolyKV: A Shared Asymmetrically-Compressed KV Cache Pool for Multi-Agent LLM Inference cites this paper.

PolyKV: A Shared Asymmetrically-Compressed KV Cache Pool for Multi-Agent LLM Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:51:34.356303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T03:55:32.585495Z digest=sha256:05151bd1028519e9c063e69a07bc167ea7bb0fe000b8327a889c20dd72cf794f

Observation c7fc36aa-0ad4-44f9-93ab-ccb34322ac62 · inbound

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization cites this paper.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:08.002040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:628bcd147b5d34e8788fcc2f9dbe4d93f6e7e4d9468f05bb3d3f938ba0e63122

Observation d57de22e-4e01-405f-8866-348250b34b03 · inbound

HeadQ: Model-Visible Distortion and Score-Space Correction for KV-Cache Quantization cites this paper.

HeadQ: Model-Visible Distortion and Score-Space Correction for KV-Cache Quantization KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:29.776641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-07T17:07:28.163304Z digest=sha256:376b5c38f8bd983c81ad5b84538ad270bca7f0224c1038e854560be6909b77ae

Observation 0cd80202-ee07-453b-b303-5406d11e96aa · inbound

HeadQ: Model-Visible Distortion and Score-Space Correction for KV-Cache Quantization cites this paper.

HeadQ: Model-Visible Distortion and Score-Space Correction for KV-Cache Quantization KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:20:42.515117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T18:33:19.999707Z digest=sha256:3514d6f6a03457841f94bb97dc7d31230f6230b66b5742e06b379de5095d342b

Observation 6413a979-1151-4d0a-b089-2bb3db893e42 · inbound

VORT: Adaptive Power-Law Memory for NLP Transformers cites this paper.

VORT: Adaptive Power-Law Memory for NLP Transformers KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:46:26.259266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T02:01:53.000938Z digest=sha256:a3ed35b0c3e05179cbcdd116ebd9b3bffdeb8a8721f96b1f353548d68d609f06

Observation e94c01d3-5644-4303-8353-e0656cc4ef32 · inbound

Protection Is (Nearly) All You Need: Structural Protection Dominates Scoring in Globally Capped KV Eviction cites this paper.

Protection Is (Nearly) All You Need: Structural Protection Dominates Scoring in Globally Capped KV Eviction KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:13:15.987749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:13:14.295659Z digest=sha256:1c75fd84efdeea1276bf1bb55652589da81c7c54338e83a23bdf393a1b65fc70

Observation f06e10bb-9277-48af-bb23-f84e9d816afa · inbound

SPHERICAL KV: Angle-Domain Attention and Rate-Distortion Retention for Efficient Long-Context Inference cites this paper.

SPHERICAL KV: Angle-Domain Attention and Rate-Distortion Retention for Efficient Long-Context Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:33:43.345093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T20:30:33.208386Z digest=sha256:2d7b3ed0c38d533cc435bc1c6d2f2561c5f7abe7cf6ab7d7b971256e5218fb5e

Observation cf1e2c40-e430-4817-aeb3-4d4d392feff0 · inbound

Runtime-Certified Bounded-Error Quantized Attention cites this paper.

Runtime-Certified Bounded-Error Quantized Attention KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:49:41.109858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T05:45:42.295527Z digest=sha256:88fc7eff50e6732f7de20150a74f99de575a694debcee6674f5e4ff44ee0adeb

Observation 4bfb4e11-c9c5-404c-906c-7a5c0666f1d5 · inbound

Adaptive Mass-Segmented KV Compression for Long-Context Reasoning cites this paper.

Adaptive Mass-Segmented KV Compression for Long-Context Reasoning KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:25:23.488548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-25T05:22:47.198185Z digest=sha256:9085362e94d5a25bedd31edc955de825d31b8ae0f8cc7d199d259af1394b4ede

Observation e1012cfb-6667-488f-bf37-162f4e42edaf · inbound

Influence-Inspired Spectral Rotations for Extreme Low-Bit LLM Quantization cites this paper.

Influence-Inspired Spectral Rotations for Extreme Low-Bit LLM Quantization KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:14:39.072739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T12:13:18.805668Z digest=sha256:8e4b8152f1acd147e1dabb0290919dd0d91a8d36e6924865f56b3ff4b8fe4291

Observation e4e72868-c583-4321-a84e-da0be5a57f9e · inbound

Do Value Vectors in Deep Layers Need Context from the Residual Stream? cites this paper.

Do Value Vectors in Deep Layers Need Context from the Residual Stream? KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:16:23.448886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T14:35:48.292081Z digest=sha256:f0c40fe563767e766c0e2c7999c1f20e16f5baf6b976ade423d4fd1d5a0e8e00

Observation 2c5b01e8-7ed2-4d15-b55c-7efbd6e447c3 · inbound

Do Value Vectors in Deep Layers Need Context from the Residual Stream? cites this paper.

Do Value Vectors in Deep Layers Need Context from the Residual Stream? KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T12:41:20.552739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T12:41:20.552739Z digest=sha256:1909383cc43ab2c87f7c3810e6999bdc07c62256176188c867cc070b557c006c

Observation 66f98172-5a50-4385-9606-d3916ec2c9c9 · inbound

STAR-KV: Low-Rank KV Cache Compression via Soft Thresholding for Adaptive Rank Control cites this paper.

STAR-KV: Low-Rank KV Cache Compression via Soft Thresholding for Adaptive Rank Control KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T22:57:26.297219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T18:34:20.645677Z digest=sha256:8a1e7be1b6f7560b9d21965189e0e1ed43f9f80c92660aaa197bfeafe4551c2f

Observation 0d00ab5b-bcb2-46ef-b136-993c3c1a8dbf · inbound

From Rigid to Dynamic: Entropy-Guided Adaptive Inference for Long-Context LLMs cites this paper.

From Rigid to Dynamic: Entropy-Guided Adaptive Inference for Long-Context LLMs KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-27T17:01:07.733642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T16:57:02.709967Z digest=sha256:28019adce8648327742631a1901b1cf07a5e22ddd704cd4b9a2155b9fd297242

Observation 19ab3f4d-16ff-4062-b6f0-4af4d02abc83 · inbound

Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation cites this paper.

Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:16:16.033365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T15:37:34.129339Z digest=sha256:ecd8b362fde57c4441153a13549129b7b792d01e621ecd04deaa90748e30bb7d

Observation 2bd10106-03ec-442b-b763-c59f9b220815 · inbound

GSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cache cites this paper.

GSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cache KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T15:57:06.412465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-02T15:55:40.177742Z digest=sha256:c58aebad4fb2af2c115c982d985a15d5c2395017f96793d36477f14ee0a8aff9

Observation 8e5f9032-d7cd-4df2-aa58-ef1b88649739 · inbound

Edge-Deployable LLM Fine-Tuning on a Single GPU for Telecom Network Troubleshooting cites this paper.

Edge-Deployable LLM Fine-Tuning on a Single GPU for Telecom Network Troubleshooting KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T17:31:15.972423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:31:15.972423Z digest=sha256:ed419bbb1f9cceba0d04bc7b80403c918c839abb7b62621f5a98b4b7459809d8

Observation 318bfd26-f9f5-4a53-a10a-d232fb84b2f9 · inbound

Fractal KV-Cache Archives: Lossless Symbolic Storage with In-Place Retrieval for Long-Context LLM Inference cites this paper.

Fractal KV-Cache Archives: Lossless Symbolic Storage with In-Place Retrieval for Long-Context LLM Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-09T19:16:27.748453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T19:13:15.765652Z digest=sha256:f90de18c74fefc39d7b80186a3f04ccaf9d4ad12cf2e894fbb8ae01c51b344f9

Observation 21137c33-95f5-443c-af0c-65450a67d571 · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:44.040808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:4b1ac5d86b2802d27efdba5bd4b2b43da177223331ec102caae953ed994c9cd4

Observation 079a02e3-8f49-4470-85f5-7c3e64a2be09 · inbound

Stateful Worlds, Stateless Elasticity: Exact-State Serving for Interactive World Models cites this paper.

Stateful Worlds, Stateless Elasticity: Exact-State Serving for Interactive World Models KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T12:07:13.700787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:07:13.700787Z digest=sha256:127b0f01c20a70a4f6695f5eff74424b7bc94a75ad74d336094519f03317f217

Observation 370400af-5b24-4d45-9b3e-a0fe5069a658 · inbound

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems cites this paper.

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T03:55:40.433775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:55:40.433775Z digest=sha256:1379776afd53633a09eca3bce5b350fa2853b7d4fc482f9596286b156a9fc4d3

Observation c399909b-00ee-449b-aa0e-cdb766908015 · inbound

Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers cites this paper.

Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T23:22:30.957456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:22:30.957456Z digest=sha256:3e0f8cd8253295df33cba571fddd2477c3e7f4ad945e8329d17dee56a4e35a38

Observation ad5553b7-0e5a-4377-b8c2-422f584943ce · inbound

High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration cites this paper.

High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T09:50:32.327649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:50:32.327649Z digest=sha256:e98f265b4dbd2d997ca01d8f2b8a5e30d65b72ecac42fad419220d77c3496cca

Observation 27aab2f4-0d4c-4593-9b3e-55914237ada8 · inbound

SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling cites this paper.

SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T09:51:03.359796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T09:51:03.359796Z digest=sha256:2ea9923798696ddbae2bd3e0f14055cd820f797445fc192db47ed52c9a640d18

Observation 6ac5d3aa-ed7d-4947-bbca-b783de3c136c · inbound

Where Facts Go Missing: A Layerwise Taxonomy and Per-Layer Attribution of Information Omission in Air-Gapped LLMAgent Pipelines cites this paper.

Where Facts Go Missing: A Layerwise Taxonomy and Per-Layer Attribution of Information Omission in Air-Gapped LLMAgent Pipelines KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T04:48:25.926385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:48:25.926385Z digest=sha256:af852e7878a7f6cf0e551263415a57efc3b3fcee3395c698969e0e2c4b1a181d

Observation 270ca459-9ac7-4564-81ca-0a1b40cfb920 · inbound

A Photonic-CXL Memory Appliance for Scalable KV Cache Management in LLM Inference cites this paper.

A Photonic-CXL Memory Appliance for Scalable KV Cache Management in LLM Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-30T11:12:05.608333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T11:12:05.608333Z digest=sha256:ea0f9c2e6d7a7c06ef840f7d35f6ebecdfb7e979d532bd2009ff97c2025e37df

Observation 4c113f63-5339-43b9-95e9-a9b797220e11 · inbound

WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization cites this paper.

WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T00:43:49.304692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:43:49.304692Z digest=sha256:d54c4b10f79e0580fbc3210fc82e9f31fc739bfbb19064c8f3934e0f29b1c387

Observation 60c500bb-4bef-4e45-af9e-d8f47a319e9c · inbound

SAKI: Score-Aware Low-Rank Key Indexing with Random-Matrix Noise Correction for KV Retrieval cites this paper.

SAKI: Score-Aware Low-Rank Key Indexing with Random-Matrix Noise Correction for KV Retrieval KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:35.719156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:35.719156Z digest=sha256:5372931d533c8277a0c229864332d430abf604806bd27f38a42bbe6b9d79448d

Observation 7d9282e2-7c49-456f-967f-e3304a2e1eac · inbound

SAKI: Score-Aware Low-Rank Key Indexing with Random-Matrix Noise Correction for KV Retrieval cites this paper.

SAKI: Score-Aware Low-Rank Key Indexing with Random-Matrix Noise Correction for KV Retrieval KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T00:54:42.716534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:54:42.716534Z digest=sha256:341e64335e884a2d2fc7b94c1e942bb1c7f0172212bc1768b68a871b3d9d4947

Observation b01ebdca-cafa-45f6-b606-579db731137f · inbound

Deltoris: Enabling Real-time VLA Inference in Embodied AI via Bit-level Sparsity and Speculative Inference cites this paper.

Deltoris: Enabling Real-time VLA Inference in Embodied AI via Bit-level Sparsity and Speculative Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T14:46:03.997449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:46:03.997449Z digest=sha256:ccd7d25d84b35fea02131f0364327b1c1d9b621503fa23188beee82628716337

Observation e887a3ac-789a-4f76-9b4f-c99892a87403 · inbound

Hidden Language Consistency Phenomena in Reasoning LLMs cites this paper.

Hidden Language Consistency Phenomena in Reasoning LLMs KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 226

Resolution
unresolved
no resolver link, observed 2026-08-14T04:40:10.093000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:40:10.093000Z digest=sha256:c6eefd8e8ba47a5aa953f6f38ae36a471cf847ee2104bf2388467b6454644b0d

Observation f8271b02-afab-4e8d-b0b4-d7e3fd4ce051 · inbound

When Local Variance Optimality Is Not Enough: RoPE-Aligned Q/K Rotations for Dynamic 4-Bit Quantisation cites this paper.

When Local Variance Optimality Is Not Enough: RoPE-Aligned Q/K Rotations for Dynamic 4-Bit Quantisation KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T12:54:26.504222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T12:54:26.504222Z digest=sha256:ccc87958d33030cf47c27e6bdbf39f44c3dfb400d604e4bb6709bd39061b9f69