Pith. sign in

Paper Citation Record · LEDGER

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression

As of 12 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 4 inbound Pith citation observations for arXiv:2412.12706.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12706 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:52:48.849977Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:22:44.472480Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved58
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 4aac6843-46ab-4b38-8ab9-5c78d5a22b2d · outbound

This paper cites GPT-4 Technical Report.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.603179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.603179Z digest=sha256:5cf1ae6f77a9dc40065caaa36d941f1ed8afa980fcf44a95fec645fb7feda277

Observation b4e2eee3-477a-4c79-abd7-3dfbddb30765 · outbound

This paper cites an unresolved cited work.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.609590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.609590Z digest=sha256:98b163e74997a1245a58853115613459551632667d504b4d5ad55fc554a44228

Observation f5ab3da4-7e46-4448-8936-1caf9816542f · outbound

This paper cites an unresolved cited work.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.613738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.613738Z digest=sha256:91f23eb2e4d7aa72ba2368dfa324a8bde70dbead1f801f1037ee812e610fdb57

Observation 8a250ea3-447b-4b62-bfc6-eae6475f56d6 · outbound

This paper cites an unresolved cited work.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.618338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.618338Z digest=sha256:e0097d794286bf0642745a6b9686ca2defe5bdba2a66a3868ae5f58db7bd27cc

Observation 3572522b-9059-4164-93e1-a936006ebe83 · outbound

This paper cites Reducing Transformer Key-Value Cache Size with Cross-Layer Attention.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.622944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.622944Z digest=sha256:f49deecf74daa972975b3191ca9ed10bd95e2fa8636c96b8ea7a77610aad54a4

Observation d966f8cd-083a-4ced-b1b4-aa87034af389 · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.627633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.627633Z digest=sha256:021de369fc6e98d20ef2115ae336b345febbe0af0e1f3fce448efa3a64c71b49

Observation 319992dc-705e-44b7-a0e3-c735e20097dd · outbound

This paper cites A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.632269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.632269Z digest=sha256:495f1f50969b18dc37dbf4b55f604800499f01b6171e9ae62d48c3aa79f056fc

Observation a3839389-59a8-44b1-a9d8-8c8105e0c4fa · outbound

This paper cites A Simple and Effective $L_2$ Norm-Based Strategy for KV Cache Compression.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression A Simple and Effective $L_2$ Norm-Based Strategy for KV Cache Compression

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.636562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.636562Z digest=sha256:ceadeae21e946f0432bb3d54f140ff4676db66ab600051a825e5fe3ab0e12c38

Observation 9e61393d-9a6e-47bf-ab4c-ebfa3a863255 · outbound

This paper cites QAQ: Quality Adaptive Quantization for LLM KV Cache.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression QAQ: Quality Adaptive Quantization for LLM KV Cache

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.640865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.640865Z digest=sha256:a2a2b096295b84463df23c5e372b69b8abcd7e6c8d523bd002c0bd4d64611ef2

Observation e0d00c62-bf97-4e3c-a472-6052a117034c · outbound

This paper cites The Llama 3 Herd of Models.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.645345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.645345Z digest=sha256:15c61f504048ddbb0b7a47bbbf8f48e24b645f8c0a920b9afc4d861738d58ef7

Observation f5054f91-5156-4c1e-871a-9067277221a8 · outbound

This paper cites Change Is the Only Constant: Dynamic LLM Slicing based on Layer Redundancy.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Change Is the Only Constant: Dynamic LLM Slicing based on Layer Redundancy

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.649611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.649611Z digest=sha256:ac2cb6fdddb0d46e6716c7e94035e7f51f3d0e43c45e6738bb50befb1fe8bc21

Observation 8a050f92-c402-4a6f-b993-cf53f3901440 · outbound

This paper cites Multi-News: a Large-Scale Multi-Document Summarization Dataset and Abstractive Hierarchical Model.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Multi-News: a Large-Scale Multi-Document Summarization Dataset and Abstractive Hierarchical Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.653555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.653555Z digest=sha256:bc13e360526b6d29c9225fe6588598b69dd55a6919a1a6f27c65ebb24d204b06

Observation dfbcca3e-0405-4041-bb9b-fc3170baf4af · outbound

This paper cites Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.657547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.657547Z digest=sha256:4d1d219e45f60e71fdf3b2e8c55283ca5735139d597e602272a626f4860347b8

Observation 8a5cee16-8ed1-42ca-927c-7fc156bd969e · outbound

This paper cites Challenges in Deploying Long-Context Transformers: A Theoretical Peak Performance Analysis.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Challenges in Deploying Long-Context Transformers: A Theoretical Peak Performance Analysis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.661704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.661704Z digest=sha256:8c06c2a336c28f8bb323604142b96b8b208accd2de2a867a49736c0f4c38d791

Observation bc93497a-f592-47ee-b778-755f4fe91740 · outbound

This paper cites an unresolved cited work.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.665984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.665984Z digest=sha256:659066c9df392b1bda0c91549c2bb317c1dcd9662c2300a072e0f4cf987a9ccd

Observation bd5c9e8b-2929-47ae-836c-42a1e62494ef · outbound

This paper cites an unresolved cited work.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:52:49.540046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T13:52:48.669780Z digest=sha256:dc5740e8266bf179aad8078ad6fff7294f5760fe84c845306e638da210baf0fc

Observation 0cceca8d-4697-44d0-b08b-360710715d99 · outbound

This paper cites SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.674253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.674253Z digest=sha256:bd57bc2222b48799d7d64576802eae7691307f943d2d352c510b962fa9751191

Observation 9cc8db87-9ffb-4a0b-aba1-06d773342aba · outbound

This paper cites an unresolved cited work.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.678662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.678662Z digest=sha256:fb2925efa0c72c0bcc5c35d24874c41d8680e3685a74ead3694769b5a4630719

Observation 85512ca1-b096-4e75-a4e7-4dcf9e2c5833 · outbound

This paper cites Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.682753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.682753Z digest=sha256:750a4e724988dda2d8e654bf0d0ff860e00285bb39d5620e69bdf25c6bc144ee

Observation 3535e96e-a105-4684-b7c5-8b957cf300b8 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.686945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.686945Z digest=sha256:ae69132a49ccea4c56fb80fb96dac52be7d14c7d0d7447e496738229b2337840

Observation 0203ab4c-f42e-4cb9-abcf-494cce6135c9 · outbound

This paper cites an unresolved cited work.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.691110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.691110Z digest=sha256:5649132c1f602adb761161e728b330fbab364cda6b9fa1aac9ff763af6657b47

Observation 326f13d4-abc0-4fe5-a144-9b06272784dc · outbound

This paper cites Efficient Attentions for Long Document Summarization.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Efficient Attentions for Long Document Summarization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.695396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.695396Z digest=sha256:9e168e77f999258baef2a1308e7e10cc077fa750ccd25b1a4767dfa67ca0e431

Observation a4a01b9b-c08d-4885-8d73-ea8cebbd24a3 · outbound

This paper cites Mistral 7B.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Mistral 7B

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.699699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.699699Z digest=sha256:e163a1562854f3daafb5eedaac5508b0438582e196fd1e578433c1353d745b7b

Observation 61658d2a-b389-49c8-9c35-c44c75e4e0f4 · outbound

This paper cites TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.704179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.704179Z digest=sha256:cfa1b2d64dd03975312b72f2816659d9b28eeca80fdb519d7b9f67f7374d2cd7

Observation 1d9b74a7-66b3-4087-bc06-5894db2567de · outbound

This paper cites an unresolved cited work.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.708262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.708262Z digest=sha256:f71b95f9f01d8e83e1ae7695e8d616029787288409c2d470f003b3da4007477a

Observation ae71be05-5aad-457e-b094-8edc2005b616 · outbound

This paper cites GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.712247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.712247Z digest=sha256:aa228e5916919f95527e957b90af63314fd91eb0518a87ffbe679a7f2e1f0d1e

Observation 39ef9cbc-8728-42a1-b94b-03cb1052f3a9 · outbound

This paper cites an unresolved cited work.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.716262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.716262Z digest=sha256:867115a4212afd032d8b5f19314a8f7eca14dd08ee401f9646414f03ebfef9ab

Observation 9d2dbc50-4cd7-4c6f-bffb-b28e3030b5d8 · outbound

This paper cites an unresolved cited work.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.720090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.720090Z digest=sha256:a0dd7dd3adfab9bdb48ce37fc65811019d6c52ea25649255285a8077640b9c4c

Observation 074a42db-77c8-4044-b3d8-48237ff5f0cb · outbound

This paper cites SnapKV: LLM Knows What You are Looking for Before Generation.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression SnapKV: LLM Knows What You are Looking for Before Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.723644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.723644Z digest=sha256:bccd57f88ce78782a01a80bd4160c896dbfb5e7e465746161d697ee5043650eb

Observation 47f656c3-4bbf-4ead-b092-3adb7a000df4 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.727805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.727805Z digest=sha256:946ea90bc2ee1b9d15d0f55022c1603404b79f208d9bb9aaebf4d09134e161f4

Observation ca6baa67-4dcc-4e11-bed4-6f5eabcda06c · outbound

This paper cites RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.731748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.731748Z digest=sha256:46c9609dd245dfb64ffffd6d0d8ebcbd83b46b41d282331b1ee50a118193afdb

Observation acf45d95-2fe2-4f44-8501-d9de12c47191 · outbound

This paper cites an unresolved cited work.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.735743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.735743Z digest=sha256:626d9f920eb3f6718dbb4b0fea0b0a3253ab19da6c88a1ea83bdc3d63c04917a

Observation 22dc7c87-611a-4b1c-bad4-dab585df6687 · outbound

This paper cites an unresolved cited work.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:52:49.481036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T13:52:48.739410Z digest=sha256:cfc32e56c7831696430b91f33f1b3148b1cd4ace117fd226245a8757f56e60aa

Observation adb61239-7bbc-48ef-b624-cd985096f4f7 · outbound

This paper cites ShortGPT: Layers in Large Language Models are More Redundant Than You Expect.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression ShortGPT: Layers in Large Language Models are More Redundant Than You Expect

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.744290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.744290Z digest=sha256:5fedd3dd6165fad81c01b1acbb53cd500f29c32efeea894db0b927d86fce41ba

Observation d971f70a-682a-4856-bcfa-f608c3334d5f · outbound

This paper cites Transformers are Multi-State RNNs.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Transformers are Multi-State RNNs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.748341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.748341Z digest=sha256:44796913c683303e41fc0778363686d5224eb5305792478e80a4f239ceaebb6f

Observation d6191a40-59c5-49d3-bf19-55df5188ce17 · outbound

This paper cites an unresolved cited work.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:52:49.466179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T13:52:48.752530Z digest=sha256:bd3c7f39c72821fe193850325e2199a0c83c64c1211582c34ee10c12866dbabb

Observation e9dae6ba-9077-4563-aeac-939eae87cb71 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.756456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.756456Z digest=sha256:ddc5db0ab92c5befbf932ac00f6a521f077e81fee0635e0f070570abb2257e81

Observation 879af09d-5eb5-47d0-9bd8-6bd52689fa4d · outbound

This paper cites On the Efficacy of Eviction Policy for Key-Value Constrained Generative Language Model Inference.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression On the Efficacy of Eviction Policy for Key-Value Constrained Generative Language Model Inference

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.760345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.760345Z digest=sha256:76a02c37dfc0f108f21cf352d0d321946ec974d72010f4edc9f7b58a56721de3

Observation 41705c5d-61cf-4c5e-b7a3-fb74a5dd1d4e · outbound

This paper cites Code Llama: Open Foundation Models for Code.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Code Llama: Open Foundation Models for Code

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.764236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.764236Z digest=sha256:9483f1592924af37b20c218c9470a186a3595f62e3137c3e4496f4a670f55278

Observation 1570d5ce-cc5a-4149-969a-36a576e1ca9a · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Fast Transformer Decoding: One Write-Head is All You Need

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.768164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.768164Z digest=sha256:08da3135b7fcd38cc62b73d52618ffee5c89bc4c9b31a06554b737492001fc7b

Observation 86b5861d-d14e-4027-861a-4da00ddfff1b · outbound

This paper cites FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.772338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.772338Z digest=sha256:3551bb249f71fe0dc91d28de9cb172e39279520cc5930fd9a138cb75d02c3d42

Observation 608b902e-6534-470d-99d3-5c17230e5068 · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.776673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.776673Z digest=sha256:8d1061e24cbd2b8ad9e15de30555acfe4554a4dcb4b410c0557890b427cc9877

Observation f46d77db-056b-4eba-8efa-464c6fefa035 · outbound

This paper cites You Only Cache Once: Decoder-Decoder Architectures for Language Models.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.780704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.780704Z digest=sha256:e13c2be9ed3c4e89ad303e7c15da584a4dbd55c23b5cbbefe381f9e30c8de81d

Observation 433572ef-8c99-4af6-aea1-0aa3ca1a42a8 · outbound

This paper cites RazorAttention: Efficient KV Cache Compression Through Retrieval Heads.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression RazorAttention: Efficient KV Cache Compression Through Retrieval Heads

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.784828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.784828Z digest=sha256:562b5e04661c8609dc1b9a742e9a3f6be8b39c5bc4e49334bf4042fccb47dfba

Observation db3e8c39-8a61-41de-8197-23b10c592cb4 · outbound

This paper cites an unresolved cited work.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work

Reference 45

Resolution
verified exact
doi, observed 2026-08-11T13:52:48.887018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T13:52:48.790293Z digest=sha256:00120151fade6a812c9a527e3b9a1f4f48e5d58c8f93e61032a048e0efb94145

Observation 0ef2ad29-0ca0-4bc1-b811-e926fd7bf045 · outbound

This paper cites Layer-Condensed KV Cache for Efficient Inference of Large Language Models.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Layer-Condensed KV Cache for Efficient Inference of Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.794830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.794830Z digest=sha256:fbac526ecdcc03be147eae0e1f5c15cd536533941061a22827a6c244e5765290

Observation 81713329-28f6-4f7d-a9db-cca3c5548cfa · outbound

This paper cites DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.798947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.798947Z digest=sha256:76f5957780c681ca080c074f931c44e51d1b86ee82bb471aab77ad2ba68fc4c5

Observation a4524f31-26d4-437d-ae31-51fe80d91a8a · outbound

This paper cites an unresolved cited work.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.803181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.803181Z digest=sha256:fe92cacc1d10764992cb40604c40d6f2363337c0136943afbc9f8cf07fbd4112

Observation f2ec4fe1-747a-4a2b-8a12-5f7720a90062 · outbound

This paper cites PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.806959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.806959Z digest=sha256:0d837ac4052e04884550fcc719762d32c705831a3b751da49730392ee2b46254

Observation 66ba1ade-5ea3-41f5-95cf-c3d6a593eddc · outbound

This paper cites No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.811586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.811586Z digest=sha256:f2afe83060c6181242a7458db4cae0504a944ae7122188d292efccabf8109f59

Observation 2225f4a7-6c56-4669-8861-843b7c683579 · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.815703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.815703Z digest=sha256:484b2bcdb946e031b3888bdbcd84c5a87a4b4c025112b6f468585aaef8e2af01

Observation 9578d8f9-5a7b-4e17-a3bd-3178a851a30d · outbound

This paper cites KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.819926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.819926Z digest=sha256:45b713f127aa27bace599e29aade3c5e576520d02ef20709e83f5a451fdd6f65

Observation 2d08c666-29be-480f-b0ce-4b2dfa2b0909 · outbound

This paper cites an unresolved cited work.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:52:49.443532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T13:52:48.824338Z digest=sha256:a9620919f5c22bc98424c85115d760122f3d2dd8110bd797b3020e5d44ace6fc

Observation 88e68680-e31c-43d1-95fa-47e4dda70b40 · outbound

This paper cites an unresolved cited work.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.828430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.828430Z digest=sha256:4617b8727e1a881c64e9a99167d52d84044dacb6dca92ca2417d75ce7d418f9a

Observation d866bc3e-e5b0-485c-8a25-d6aef81ebdcd · outbound

This paper cites QMSum: A New Benchmark for Query-based Multi-domain Meeting Summarization.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression QMSum: A New Benchmark for Query-based Multi-domain Meeting Summarization

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.832331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.832331Z digest=sha256:e159a39885014098bb03e004620ccb6d2a9fb03cdbeb0da876887d34c009be6e

Observation 138c695b-ded6-41ab-b44a-1d52440a28cc · outbound

This paper cites DOCBENCH: A Benchmark for Evaluating LLM-based Document Reading Systems.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression DOCBENCH: A Benchmark for Evaluating LLM-based Document Reading Systems

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.836440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.836440Z digest=sha256:2cab193a26404712c3fa4dcb9d05e9b742b3d9cb424964f1e197e8164f853d06

Observation 379f3c06-11b5-4a3f-8aa5-1a0921f3d02b · outbound

This paper cites MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.840582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.840582Z digest=sha256:2aeb87dce3cfb847676d01b5002cabb715ca41b1bd3e4a92785f2ece6e9a034e

Observation 0b35799b-4dc9-45e2-905a-3e88ad523a71 · outbound

This paper cites online" 'onlinestring :=.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression online" 'onlinestring :=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.845359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.845359Z digest=sha256:c88e515e0b818556e124cee303b4b8b05346881722ba3fe1e79cd1ff2b408f4e

Observation 17ab0270-cadf-430d-8938-fcd03c62b8c6 · outbound

This paper cites write newline.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression write newline

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.849977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.849977Z digest=sha256:4867dbe04b4be934c7e8d89aa74414318946828b8c380409868b599a127758a3

Pith citing papers

Observation 9f28759a-2cf3-4dd8-8106-86b6acf5f8e9 · inbound

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference cites this paper.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:44.472480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:44.472480Z digest=sha256:35c624f2aaab505bb94c99b2b6866719260be131c9fd687c12b2ec7597b394f3

Observation 0a2bfef6-566d-478b-b29d-f25090467f8b · inbound

RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache cites this paper.

RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:06:33.927545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T01:16:15.869198Z digest=sha256:4dffa10db333c0d3f831bbf64fb696a1c626e17ed7efcb1aeb7f3837cf068f2a

Observation d1656047-a3a4-48e7-9acf-17d8fd237be7 · inbound

YouZhi: Towards High-Concurrency Financial LLMs via Adaptive GQA-to-MLA Transition cites this paper.

YouZhi: Towards High-Concurrency Financial LLMs via Adaptive GQA-to-MLA Transition More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:06:59.413384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-28T01:33:54.383507Z digest=sha256:c280fe1464a5f2b07a2a0e8d0710b09f35ee45b997eac4f29c05793ad495b693

Observation 8452ac4d-a63d-4cdd-9813-4130c14201b1 · inbound

DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression cites this paper.

DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression

Reference 145

Resolution
verified exact
local_arxiv, observed 2026-07-08T03:14:31.621551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-07-08T03:07:23.648382Z digest=sha256:71596117b1320065cd1fa181f1d2cd1fe6d3fc5596949af39b344b2a11f42839