Pith. sign in

Paper Citation Record · LEDGER

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms

As of 16 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2608.04074.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.04074 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:54:14.123344Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dffa9505-32f4-4a44-8b11-943e79995d5d · outbound

This paper cites gpt-oss-120b & gpt-oss-20b Model Card.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms gpt-oss-120b & gpt-oss-20b Model Card

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:13.974467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:13.974467Z digest=sha256:1be828d078be50267c0f71838825ad67ef0bdddab18f189ee247febecdd5165a

Observation f2463839-7179-43f1-8f21-facf893d71e5 · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.029950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.029950Z digest=sha256:b6ae3bc7bbe2ed8ffc422389e4cdc54a0a5094721199397b5d458e1cd3d75e63

Observation 2dc846ac-aa8f-4dca-b295-7140eb5ea9d3 · outbound

This paper cites Livecodebench: Holistic and contamination free evaluation of large language models for code.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Livecodebench: Holistic and contamination free evaluation of large language models for code

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:54:14.731144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T14:54:14.035404Z digest=sha256:17667b2cfcd45612ef51601533f10bec22c2ea334f38ba5df3fa37289c19ce99

Observation 7ebaf411-d1b1-49d8-94fa-34c1c5beb7ae · outbound

This paper cites LooGLE: Can Long-Context Language Models Understand Long Contexts?.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms LooGLE: Can Long-Context Language Models Understand Long Contexts?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.040515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.040515Z digest=sha256:d9004360a479499a6d04e01a064c7d2a4b507aa09675b25d9923c2c47a30ac5b

Observation 8e17b076-902a-442c-9477-b77ee178bbda · outbound

This paper cites CommVQ: Commutative Vector Quantization for KV Cache Compression.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms CommVQ: Commutative Vector Quantization for KV Cache Compression

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.045740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.045740Z digest=sha256:793cdb9b7b80565b5a69d156f7fd4870507a667f1ef46d5b5f9f8f7b84df093f

Observation 5884f49d-38b9-46de-87c7-b7891340b358 · outbound

This paper cites Jamba: A Hybrid Transformer-Mamba Language Model.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Jamba: A Hybrid Transformer-Mamba Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.050475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.050475Z digest=sha256:3c8ddc2b35bd388f982694593f8f5284eeeedd1a0412b4414550df2574e1bfe9

Observation 683bf42f-6190-4a5d-8f23-abce0a8533fd · outbound

This paper cites Let’s verify step by step.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Let’s verify step by step

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.055320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.055320Z digest=sha256:0953c2fdd8770253e2bcac722db37c81789c956babc14e229c8023b6c917427c

Observation 6448dbdd-70ab-4cc5-993a-64b0028f97d3 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.065127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.065127Z digest=sha256:af550d4e93f65f495ea213f0644ee1d96807c57d94b78e653d1b988ae242a7a0

Observation 00318c5f-ff3f-4a87-ab9e-71b2e0d5eee2 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.075430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.075430Z digest=sha256:fc2ff9cbfc27369ddf89b41f07f4b2f93da8808f66397e8b15c573a735076f28

Observation 5c26530b-5a78-49b2-b132-e37d20363a2a · outbound

This paper cites Magicdec: Breaking the latency-throughput tradeoff for long context generation with speculative decoding.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Magicdec: Breaking the latency-throughput tradeoff for long context generation with speculative decoding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:54:14.687491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T14:54:14.080404Z digest=sha256:c4ebf5bf00630e13e46ef2a49fc3200b349f5dba4c619d36542111794b49e05d

Observation 28187aa7-6b66-411d-acbf-b8e988321572 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Fast Transformer Decoding: One Write-Head is All You Need

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.086026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.086026Z digest=sha256:f4b0f14e0b4e438df43c0dc2201741889bd7efe3a5000d960a1aa3128b518238

Observation a05d0bf4-a162-4d4e-af95-648437d2875e · outbound

This paper cites RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.091576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.091576Z digest=sha256:bcf67867acba26a6dc0d0dc7609dcf813a1374a95d9461347971c9c20c2b1c02

Observation bfa47726-d9bf-4943-a713-36c852ab965a · outbound

This paper cites Efficient streaming language models with attention sinks.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Efficient streaming language models with attention sinks

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:54:14.671026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T14:54:14.108543Z digest=sha256:3728737e1cca65285e7e87f90db8c41a069f90596053c7b38621747f203164f9

Observation 2a34392d-b271-4c82-855c-cad61477afb9 · outbound

This paper cites Qwen3 Technical Report.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Qwen3 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.113222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.113222Z digest=sha256:a3b5a69af880cabfb4b434c5aefed5d0a4e5711ccef549b650efb22a38efaa6b

Observation 29d86e9c-3cb9-4aaf-ac2d-639523fdbbef · outbound

This paper cites OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.123344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.123344Z digest=sha256:eff2d045bf6fa61ccae19779acbf7ca3dd35ba74307cf89c68d1eba91cf0a47d

Observation c7a97561-1927-4eae-994c-3b2b5b9b3412 · outbound

This paper cites On the Properties of the Softmax Function with Application in Game Theory and Reinforcement Learning.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms On the Properties of the Softmax Function with Application in Game Theory and Reinforcement Learning

Reference 1949

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.009025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.009025Z digest=sha256:5e13727226e4cfd71f317236c12067e6b93e70b877e62061300ddd29fc664fdf

Observation ae330aa2-8bc2-43e9-aa41-b3782d315ff2 · outbound

This paper cites NQKV: A KV Cache Quantization Scheme Based on Normal Distribution Characteristics.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms NQKV: A KV Cache Quantization Scheme Based on Normal Distribution Characteristics

Reference 1971

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:54:14.618846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T14:54:13.992944Z digest=sha256:0ea57578c5eeec9a664d27da092f6c45856ab9370cff688f3f0a358fae5e771c

Observation 9845232d-b613-4b20-a425-e9ae727f475d · outbound

This paper cites TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate

Reference 1982

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.118227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.118227Z digest=sha256:e1218c0f4e039d02641623668afe41796eaaf75adb21e31741a089772b70da1d

Observation 1bacb15e-2bbe-45f4-b93c-5198d689821a · outbound

This paper cites AIME 2025: American invitational mathematics examination.https://maa.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms AIME 2025: American invitational mathematics examination.https://maa

Reference 1989

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:54:14.704663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T14:54:14.070261Z digest=sha256:2897772e04bb6c83d6666a7cc6c58784a0b77074f006c37094d8a4d10fd70470

Observation 4e55d921-80a1-4a16-afc3-37cce8ae555b · outbound

This paper cites doi: 10.1007/978-1-4615-3626-0.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms doi: 10.1007/978-1-4615-3626-0

Reference 1992

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.014851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.014851Z digest=sha256:65f43b7ebc8696f2e5fb158a8892e150a10fbd287d8498acf2ef8f7fade0c1d3

Observation 68dad4eb-3689-4e02-b479-8c4761700957 · outbound

This paper cites Gemma 3 Technical Report.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Gemma 3 Technical Report

Reference 1998

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.097282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.097282Z digest=sha256:70f677f4e1449294de29635aa7ab0d517e08ae4de49eb3ce6f10c1f43925edb8

Observation ca02879d-05da-4900-b1e0-7cf66f6e422c · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 1999

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.060218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.060218Z digest=sha256:89c49b1683ffa5659f7f805776a14b6a6a86e7bfb89661882722eb26526a4c08

Observation 008b43d0-b42b-4e71-bc8c-c7cd9ac68b61 · outbound

This paper cites The Llama 3 Herd of Models.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms The Llama 3 Herd of Models

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.019799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.019799Z digest=sha256:b4993ce81691d5b49dbeda3a29f5bd9bb3c1213d8142f9fc9525b28a9f3d9ba7

Observation ff032e55-9615-4709-88a9-71175f57a4f8 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Measuring Mathematical Problem Solving With the MATH Dataset

Reference 2002

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.024939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.024939Z digest=sha256:cff58591a66cc0b54c9f8705af1c6b7ea18de7c82ca2a796aa3d3c20bc145119

Observation 08fd2b8d-8ee7-44b9-8920-d8d4c513ef0b · outbound

This paper cites Kitty: Accurate and efficient 2-bit kv cache quantization with dynamic channel-wise precision boost.arXiv preprint arXiv:2511.18643,.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Kitty: Accurate and efficient 2-bit kv cache quantization with dynamic channel-wise precision boost.arXiv preprint arXiv:2511.18643,

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.103353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.103353Z digest=sha256:e79b14a16676e24b3d65ee7a79fad96fc4d81370d356bab0d07e1905e4056549

Observation f52db0ef-2c90-4882-9b9f-6a74d1d43053 · outbound

This paper cites Expected attention: Kv cache compression by estimating attention from future queries distribution.arXiv preprint arXiv:2510.00636,.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Expected attention: Kv cache compression by estimating attention from future queries distribution.arXiv preprint arXiv:2510.00636,

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.003797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.003797Z digest=sha256:3f7699710ccdc96ee4e8dacb0d3cf3bbb27cfe949a8e98b3b8044f9e30e9df0b

Observation 0c25ae6c-ede3-45bb-985f-35df6b90644b · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Evaluating Large Language Models Trained on Code

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:13.998330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:13.998330Z digest=sha256:80633a58d3422047b1070fc859f98b316caca50c5a255d40e08ff52426c7756e

Observation affaf05b-ca25-4a55-ab28-a150aa3d4916 · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:13.987272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:13.987272Z digest=sha256:d95fde03d3048f573d34eea5b2610bed6e72fe525ce59b3035fde0f46b9febe5

Observation 098c89fd-a716-4d80-b679-63aa5339b356 · outbound

This paper cites Gqa: Training generalized multi-query transformer models from multi-head checkpoints.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Gqa: Training generalized multi-query transformer models from multi-head checkpoints

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:13.981185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:13.981185Z digest=sha256:d3f75449db42fd5633e88c84634e7329b825f844639df022dd535cd637825b46

Pith citing papers

No inbound Pith citation observations are available.