Pith. sign in

Paper Citation Record · LEDGER

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms

As of 16 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2608.04074.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.04074 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:54:14.123344Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dffa9505-32f4-4a44-8b11-943e79995d5d · outbound

This paper cites gpt-oss-120b & gpt-oss-20b Model Card.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms gpt-oss-120b & gpt-oss-20b Model Card

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:13.974467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:13.974467Z digest=sha256:ae1959c52e95814aba8b250fa2820f4e1aa3aebe00c1731cc238835ae2fe6786

Observation f2463839-7179-43f1-8f21-facf893d71e5 · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.029950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.029950Z digest=sha256:2216d6f6c449e51a09d85f23e2cde45e4413442e639aaa6aba7543db117d0950

Observation 2dc846ac-aa8f-4dca-b295-7140eb5ea9d3 · outbound

This paper cites Livecodebench: Holistic and contamination free evaluation of large language models for code.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Livecodebench: Holistic and contamination free evaluation of large language models for code

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:54:14.731144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T14:54:14.035404Z digest=sha256:829afe5dc2a49666ff14b4b7290f733f8abb316bc32fcc50ed161d1100f7576a

Observation 7ebaf411-d1b1-49d8-94fa-34c1c5beb7ae · outbound

This paper cites LooGLE: Can Long-Context Language Models Understand Long Contexts?.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms LooGLE: Can Long-Context Language Models Understand Long Contexts?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.040515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.040515Z digest=sha256:db94d979ac31afa7dcc29e3d42b048c4021c66013a9d209df19558860d4d17b9

Observation 8e17b076-902a-442c-9477-b77ee178bbda · outbound

This paper cites CommVQ: Commutative Vector Quantization for KV Cache Compression.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms CommVQ: Commutative Vector Quantization for KV Cache Compression

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.045740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.045740Z digest=sha256:9c14086afb64374d8ae3f24b56a00a739308b31a3fed4c9790bb434e7c638cf0

Observation 5884f49d-38b9-46de-87c7-b7891340b358 · outbound

This paper cites Jamba: A Hybrid Transformer-Mamba Language Model.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Jamba: A Hybrid Transformer-Mamba Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.050475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.050475Z digest=sha256:b650a42152b27b8c93adc5b568a515ce537b6cbfe955c44a00d2e7e2c3b744c4

Observation 683bf42f-6190-4a5d-8f23-abce0a8533fd · outbound

This paper cites Let’s verify step by step.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Let’s verify step by step

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.055320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.055320Z digest=sha256:b89ab0304ab9fd8c9a07543f1d2704c341b4871c5047825bceab71e5b11f792a

Observation 6448dbdd-70ab-4cc5-993a-64b0028f97d3 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.065127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.065127Z digest=sha256:de4841e267d3ccc9fd06b08161da55bfa5c2e5f9bf2e7e4033e3dba90cefa8a2

Observation 00318c5f-ff3f-4a87-ab9e-71b2e0d5eee2 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.075430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.075430Z digest=sha256:42918013aad131bdaf283796be5fb03f66985f0a5460a6ce42078a20127a3b85

Observation 5c26530b-5a78-49b2-b132-e37d20363a2a · outbound

This paper cites Magicdec: Breaking the latency-throughput tradeoff for long context generation with speculative decoding.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Magicdec: Breaking the latency-throughput tradeoff for long context generation with speculative decoding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:54:14.687491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T14:54:14.080404Z digest=sha256:44ffe895af0d394c3a4fd808c7e72cd0c946e6d027e607c7cad9a68d8c7da23b

Observation 28187aa7-6b66-411d-acbf-b8e988321572 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Fast Transformer Decoding: One Write-Head is All You Need

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.086026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.086026Z digest=sha256:6fc9d761d39f0ee05ca28d5ae36806637a263d2fc6fe74f977e05a63e931840c

Observation a05d0bf4-a162-4d4e-af95-648437d2875e · outbound

This paper cites RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.091576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.091576Z digest=sha256:d7fcc8b97714367ddf23a99cc6a655a7b1caec7c5cc2992373856042475d2df0

Observation bfa47726-d9bf-4943-a713-36c852ab965a · outbound

This paper cites Efficient streaming language models with attention sinks.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Efficient streaming language models with attention sinks

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:54:14.671026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T14:54:14.108543Z digest=sha256:4bd8ea3a3a792358071a8b62cecb65f49ea638c04cb21e796dc8936acb9c3487

Observation 2a34392d-b271-4c82-855c-cad61477afb9 · outbound

This paper cites Qwen3 Technical Report.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Qwen3 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.113222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.113222Z digest=sha256:a2733e665f2e2dcc68f75b929158fa51667be0a598dbe0dee0e59aa86a046f53

Observation 29d86e9c-3cb9-4aaf-ac2d-639523fdbbef · outbound

This paper cites OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.123344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.123344Z digest=sha256:3c84e5cb63f355633a0984bd9ea6d6f525d19a37b7a1875665acb84bf1fbed68

Observation c7a97561-1927-4eae-994c-3b2b5b9b3412 · outbound

This paper cites On the Properties of the Softmax Function with Application in Game Theory and Reinforcement Learning.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms On the Properties of the Softmax Function with Application in Game Theory and Reinforcement Learning

Reference 1949

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.009025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.009025Z digest=sha256:5408ddec5079d253832b005115593f3b7a6cdfd654e8925587784951ad4c4f78

Observation ae330aa2-8bc2-43e9-aa41-b3782d315ff2 · outbound

This paper cites NQKV: A KV Cache Quantization Scheme Based on Normal Distribution Characteristics.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms NQKV: A KV Cache Quantization Scheme Based on Normal Distribution Characteristics

Reference 1971

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:54:14.618846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T14:54:13.992944Z digest=sha256:8d3d9ab58f9836ba2763162d156dc2796cc7c84a23511884f58b5168a8226ce0

Observation 9845232d-b613-4b20-a425-e9ae727f475d · outbound

This paper cites TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate

Reference 1982

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.118227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.118227Z digest=sha256:9a16f1f361147e2fb4d1e1b4e695dbc45347f63441f3127a0e304266dbda5980

Observation 1bacb15e-2bbe-45f4-b93c-5198d689821a · outbound

This paper cites AIME 2025: American invitational mathematics examination.https://maa.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms AIME 2025: American invitational mathematics examination.https://maa

Reference 1989

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:54:14.704663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T14:54:14.070261Z digest=sha256:786c3ac825d186d6676c808ea37fa01dd465a1fffeb0f31a6619e1be0db9faea

Observation 4e55d921-80a1-4a16-afc3-37cce8ae555b · outbound

This paper cites doi: 10.1007/978-1-4615-3626-0.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms doi: 10.1007/978-1-4615-3626-0

Reference 1992

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.014851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.014851Z digest=sha256:8464bab23fec5bacbbf8e6598add91142c01581a39b81e6a8c75bd541f98642d

Observation 68dad4eb-3689-4e02-b479-8c4761700957 · outbound

This paper cites Gemma 3 Technical Report.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Gemma 3 Technical Report

Reference 1998

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.097282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.097282Z digest=sha256:d05ef65df38bfadf34b8e6154a209d9a098c9c07b57049d3584a33cd8d171509

Observation ca02879d-05da-4900-b1e0-7cf66f6e422c · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 1999

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.060218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.060218Z digest=sha256:7bc8c12393e73ea0393284a766545e2efb167e0f972fecbdd38a5a192bcd3846

Observation 008b43d0-b42b-4e71-bc8c-c7cd9ac68b61 · outbound

This paper cites The Llama 3 Herd of Models.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms The Llama 3 Herd of Models

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.019799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.019799Z digest=sha256:eebd77b1d750c1feb1e4f131c82fbab0dbe227155b5076c307305b94be9c378c

Observation ff032e55-9615-4709-88a9-71175f57a4f8 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Measuring Mathematical Problem Solving With the MATH Dataset

Reference 2002

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.024939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.024939Z digest=sha256:a750884d42c82e7b933b47d06ea1a216bbc38603657c763c6e486b4d026cdaa0

Observation 08fd2b8d-8ee7-44b9-8920-d8d4c513ef0b · outbound

This paper cites Kitty: Accurate and efficient 2-bit kv cache quantization with dynamic channel-wise precision boost.arXiv preprint arXiv:2511.18643,.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Kitty: Accurate and efficient 2-bit kv cache quantization with dynamic channel-wise precision boost.arXiv preprint arXiv:2511.18643,

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.103353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.103353Z digest=sha256:f4f0f6ce4b396474081d7925074d1a5eadf73cc067be8fce9ac976ed1af34dd1

Observation f52db0ef-2c90-4882-9b9f-6a74d1d43053 · outbound

This paper cites Expected attention: Kv cache compression by estimating attention from future queries distribution.arXiv preprint arXiv:2510.00636,.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Expected attention: Kv cache compression by estimating attention from future queries distribution.arXiv preprint arXiv:2510.00636,

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.003797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.003797Z digest=sha256:8a7efcd86d217b3e81699cd9c18df9dab4b3528d61046f401a3d01923eed63ca

Observation 0c25ae6c-ede3-45bb-985f-35df6b90644b · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Evaluating Large Language Models Trained on Code

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:13.998330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:13.998330Z digest=sha256:a160400fe991b8ac66bf5cfe1aa5bfa30a529aa40184db4766654bcfc42308aa

Observation affaf05b-ca25-4a55-ab28-a150aa3d4916 · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:13.987272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:13.987272Z digest=sha256:f1af046ec9cf4df82f8ea341b6415ad8e907a25a233d16e8981e9c037c5c95c4

Observation 098c89fd-a716-4d80-b679-63aa5339b356 · outbound

This paper cites Gqa: Training generalized multi-query transformer models from multi-head checkpoints.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Gqa: Training generalized multi-query transformer models from multi-head checkpoints

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:13.981185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:13.981185Z digest=sha256:a6806a920560cfd9e9ac0af4a926471c2bb1d007b9c9673731200b983b0deb7c

Pith citing papers

No inbound Pith citation observations are available.