Pith. sign in

Paper Citation Record · LEDGER

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries

As of 23 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 7 inbound Pith citation observations for arXiv:2412.08890.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.08890 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:32:39.990624Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T13:26:55.395701Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T18:40:03.257305Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6996c122-ee75-4e9d-a5c2-89085e9018f3 · outbound

This paper cites Gqa: Training generalized multi-query transformer models from multi-head check- points.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries Gqa: Training generalized multi-query transformer models from multi-head check- points

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.921656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.921656Z digest=sha256:fb48475bd7e976703513dd30956ae17d04d77aa25e26255af56fe3bf4eb1db34

Observation 7304570a-1d63-4f0b-a935-fde0456f0668 · outbound

This paper cites Longformer: The Long-Document Transformer.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries Longformer: The Long-Document Transformer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.928529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.928529Z digest=sha256:e06f56b2e4ac5a8fa5caa7738e5f307da2b0249f3339567bdbc1fde62bc88a7a

Observation 16760b27-39cb-4865-a6f4-e8cb8057b340 · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.931782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.931782Z digest=sha256:84f63d2fed4ced1d67d51ab47001c049cc2cf1b8e133b274b32c4172923a0eb9

Observation 2a14adf7-5e1a-4605-a1b1-e1e792b0d052 · outbound

This paper cites ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.944609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.944609Z digest=sha256:2729a4c7d6c999485561f6922c6a3fce90494296136532f3c2e9e39e7b8b928a

Observation ad012494-5137-4507-bdda-9643a54b41c4 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.947554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.947554Z digest=sha256:e20c983eabdf52be9e041ff81aeb189b2907e76dba5b7d58735c5ed185709073

Observation 44398d71-fd32-41e3-987c-c6fb764c9fef · outbound

This paper cites GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.949962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.949962Z digest=sha256:1c2dcf7bbbca20379c0231a65af4e7a01ef4f18dfe7cb039869092089f8d2184

Observation a6e0556e-6d30-467b-90d9-eb08435f1879 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries Adam: A Method for Stochastic Optimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.952337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.952337Z digest=sha256:828ded16c0e6b897db5444ecb14407e8e89db6b41d978ec9c25cdb22b044e98f

Observation 9a573994-344b-4f0a-9c48-6cd4df391568 · outbound

This paper cites SnapKV: LLM Knows What You are Looking for Before Generation.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries SnapKV: LLM Knows What You are Looking for Before Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.954480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.954480Z digest=sha256:922902691e004f7e29a0962a42c92e3e17689ab7d842638294f9e88286ff5c2d

Observation 3ff7a8d3-7e33-4302-a0d9-1ceac9c3347d · outbound

This paper cites Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.956930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.956930Z digest=sha256:c2f075e3db9810a9de608c1bdf65f6f82ec7eb4092f8202404a039f82e082b60

Observation 940f14fa-18b7-4b36-ba3f-0d9e6e88aca2 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.959695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.959695Z digest=sha256:f861f27f4075e29ae57d1723b81bdb86c37fdb656497de468af471f1ae582fb8

Observation 0325c185-17c6-478f-8446-98063aa181c8 · outbound

This paper cites k-Sparse Autoencoders.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries k-Sparse Autoencoders

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.962371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.962371Z digest=sha256:4b67b837f35315b429bbe5b57874687549b5f4cfe8051cb2ed1f2c4deaa2b7b3

Observation 04eba951-ee5d-4c55-9cf5-612fecbf5103 · outbound

This paper cites Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.965032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.965032Z digest=sha256:5b66427da53b6ec4a311edaef6135a82e1d9e7f8c6c9a973844e6e926aa2b845

Observation de5dc1e6-9709-4b4c-aa49-54da6f37c8c2 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries Fast Transformer Decoding: One Write-Head is All You Need

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.967735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.967735Z digest=sha256:3cfc33fdf90bc50aed4b8dd11b8f748b602a0812a895330d73dbb5ce4e90dead

Observation 9c290bd5-84e9-4430-ac12-2e7c84472eb9 · outbound

This paper cites You Only Cache Once: Decoder-Decoder Architectures for Language Models.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.973242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.973242Z digest=sha256:2cefe7a43cadea9989f896795c2f4817828992fdd04a2e75d4d5e3c7d3cffd5d

Observation 565271d7-35bb-44d4-841d-f0e363bb956c · outbound

This paper cites Attention is all you need.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries Attention is all you need

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:32:40.151637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T17:32:39.976451Z digest=sha256:2682940c597f1dc5abb9f4fd71bf253d29261e74dd0452f45b0e2c5100b58b05

Observation 7704faf6-a812-4e49-99c2-178b456a92d9 · outbound

This paper cites ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.982079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.982079Z digest=sha256:807cc0e911f8177b4cfce99a6023b7e6a975d9b3d3e047da567c6fa6b153850b

Observation f8b489a9-e388-4c65-9016-444ae4d34c70 · outbound

This paper cites WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.987580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.987580Z digest=sha256:1d4b8947889683b404eaccadf89079bfdc31fd18eb90f080dfb67a4ebef4f371

Observation b7a2ab87-7f4b-483f-8907-2113a51c0c9b · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries Training Verifiers to Solve Math Word Problems

Reference 2006

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.935049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.935049Z digest=sha256:864c1f2159becb999ad39ceaf7b14bdc24dc3612045a164864449b93ec2a5087

Observation cebab9c8-3e35-4a9e-8c79-c79e017138a2 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.979186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.979186Z digest=sha256:39ae0285a00c403989ba1939d255f7c1f854a8f828139f46dfc5b3c2e0638fc8

Observation 2fe16acb-c6c4-42fa-9764-69071f0ef38a · outbound

This paper cites Loki: Low-rank Keys for Efficient Sparse Attention.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries Loki: Low-rank Keys for Efficient Sparse Attention

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.970666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.970666Z digest=sha256:e71831b3b6694d2e1283c2367eacb43f61492befb5c2de6bdc2eb60228222b18

Observation adb11b81-00c7-4e26-8c6f-2ee114404806 · outbound

This paper cites APPENDIX A I MPLEMENTATION DETAILS Algorithm 1 illustrates a naive implementation of OMP for understanding.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries APPENDIX A I MPLEMENTATION DETAILS Algorithm 1 illustrates a naive implementation of OMP for understanding

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:32:40.143729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T17:32:39.990624Z digest=sha256:a688253b2225efa11cb5310902a9f7420111c861bd3d61da11702d8f6df7d5d7

Observation 7ca79eb1-e2b5-4e91-8f93-196b6704489d · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.938181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.938181Z digest=sha256:78301face9e9764224a2c1922b614fcdbd132250c8071d745ba6ffde83d6371d

Observation 79b0f42e-f2ec-4eca-8a4e-b02232bc8564 · outbound

This paper cites Effectively Compress KV Heads for LLM.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries Effectively Compress KV Heads for LLM

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.984824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.984824Z digest=sha256:a14cc12ebafee6b87322e05c5b59f284c178e819a9485009511f6701f0d90b94

Observation 633141e6-7103-42f7-8519-27a5ba51c752 · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.924932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.924932Z digest=sha256:36015e218c0991d3ca466e94f42a26c2d72663e4a0c2243e4284f03797eb2888

Observation dcea4b12-2e86-4870-a977-e26441bb3c7c · outbound

This paper cites A Simple and Effective $L_2$ Norm-Based Strategy for KV Cache Compression.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries A Simple and Effective $L_2$ Norm-Based Strategy for KV Cache Compression

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.941220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.941220Z digest=sha256:d61f7a46a4d274ffa2cc75e5e80bd416863f73b9ce2343f8dde494eebdc4226c

Pith citing papers

Observation e0d28393-71b5-46a6-a19d-bb518fb9b178 · inbound

PolarQuant: Quantizing KV Caches with Polar Transformation cites this paper.

PolarQuant: Quantizing KV Caches with Polar Transformation Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T13:26:55.395701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:26:55.395701Z digest=sha256:eaaac6dc7ae30b96bd50964eb12a83eb69721da4ce3e4adac2c2781986277605

Observation 624879cc-1d29-4d46-85a0-ca699718236c · inbound

TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate cites this paper.

TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-20T08:09:22.270351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T08:09:22.226608Z digest=sha256:595e44e45c32cb256bef76c92fe4ac011a78a42bc45ad3e58db36a3fdc5099dc

Observation 1916c9f4-cb6c-49b4-b9f7-9d002151684e · inbound

Cartridges: Lightweight and general-purpose long context representations via self-study cites this paper.

Cartridges: Lightweight and general-purpose long context representations via self-study Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:31.436604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:31.436604Z digest=sha256:b91f406dbea051e6cd19e75d6aa55f8325811d5c9ef141f76e9fdc74d5b4661f

Observation d94ec039-2f70-4e0b-ae55-9c8fedb922fe · inbound

ZAYA1-8B Technical Report cites this paper.

ZAYA1-8B Technical Report Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries

Reference 142

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:26:05.138038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-08T17:36:37.182196Z digest=sha256:acf03488c77f976f85e7c32266763b2117fd4c79a1fff4d31207889d36fe53db

Observation 28d8da42-9c4a-465a-9895-a050ec8cdefa · inbound

ZONOS2 Technical Report cites this paper.

ZONOS2 Technical Report Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries

Reference 167

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T18:40:03.258656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-25T22:37:15.072758Z digest=sha256:5e74b112523929e563d554aceaf46e63450d5dbcde9617e5b90887b2fce06c8c

Observation 81229a86-7276-4c50-9de8-532cb996de2a · inbound

ZONOS2 Technical Report cites this paper.

ZONOS2 Technical Report Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries

Reference 167

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T18:15:59.119424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-29T02:07:31.791835Z digest=sha256:06ee88053498b1f384cc1305134053395704404592b8a64692ad52da557686f0

Observation b3d8b8f7-8db9-4516-91fe-ce7f84b53e57 · inbound

ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution cites this paper.

ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries

Reference 152

Resolution
unresolved
no resolver link, observed 2026-08-01T09:52:04.379792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:52:04.379792Z digest=sha256:bc38ec2de3755d6c9216e8cf273474cc6b12c595ac13541ad6f7a909ac2d6113