Pith. sign in

Paper Citation Record · LEDGER

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering

As of 16 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2506.04642.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04642 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:45:09.969345Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T00:47:31.653881Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-10T05:30:23.456663Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-10T05:30:23.456663Z

Outbound references

Observation 95ddf0a1-0c66-4313-9ee2-d46877579bad · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.814940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.814940Z digest=sha256:14ee54b711e0db055d469c9bbe2e91e796dafcd29f7eb70cb8b89201e79a70e8

Observation 269b8026-bb4a-49c8-85a8-2ccdac6e1ca7 · outbound

This paper cites an unresolved cited work.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.820633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.820633Z digest=sha256:e2e472dc1326205c429017e672f32969cb423071bdb36a049a0cbadb3f4a8082

Observation 484ef146-5291-4782-9dd0-fdc49ba24a4b · outbound

This paper cites Palu: Compressing KV-Cache with Low-Rank Projection.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Palu: Compressing KV-Cache with Low-Rank Projection

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.825323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.825323Z digest=sha256:6df5d9543048d904a3e4ec566f424cfc26de0855f099f6e8c18f9ebd0ed8c382

Observation 3ff40ad3-f55c-46db-926c-d51afbbd0020 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Training Verifiers to Solve Math Word Problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.832690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.832690Z digest=sha256:cb5f98614df614a4d3130a7db59e618b4f9d54f764d0a1dee84431337ddec0b0

Observation f11cb9aa-bcc5-4735-a342-3b24f0178fc2 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.837526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.837526Z digest=sha256:2914491ad648dd5db61ecba39a277d78cedae881dd7c343e8ff31e06698b2a41

Observation 352a4bf6-f68e-4646-98a1-65937f420473 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.842117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.842117Z digest=sha256:3dab806a8602656be190235c69ffd6f5fca0e3b13271ef92ac9de1bf35f69503

Observation b3e9de94-c814-45eb-be2e-14625b18e067 · outbound

This paper cites QAQ: Quality Adaptive Quantization for LLM KV Cache.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering QAQ: Quality Adaptive Quantization for LLM KV Cache

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.847376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.847376Z digest=sha256:20a3fd26b2ef21e1dc8a4bffbf587155557e2ffb42212441e5bfee10c58c5fe2

Observation 5bd22f1a-af28-459a-9ec7-6ad28740fc42 · outbound

This paper cites The Llama 3 Herd of Models.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.852555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.852555Z digest=sha256:bbf3e028af20ebc8fa8ac97b348afa850e14177d2819c72f96a43d99b8cb24d2

Observation ac17d14e-2ccb-4c22-adc6-cccc746d29ca · outbound

This paper cites FlashDecoding++: Faster Large Language Model Inference on GPUs.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.857737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.857737Z digest=sha256:752a0132963d1da2f9b510d96adc40e199f039047c41638dd27f3616119e5c75

Observation dec87567-bce7-4b01-9448-54d41b907068 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.862820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.862820Z digest=sha256:b933327ef6ae25e7ff72fcb69a179895faee28dbd61e01efe1b830426580572e

Observation 2d192574-f02b-4608-8b1f-b3895ab9bcc1 · outbound

This paper cites Mistral 7B.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Mistral 7B

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.868431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.868431Z digest=sha256:4f31d7ac4f0fcc482be9d1194be0087446d838d90f0d321c59cbaf81774e4dcb

Observation dac99d7f-cb4a-4bca-8a6d-92b8b0ae33c7 · outbound

This paper cites QCQA: Quality and Capacity-aware grouped Query Attention.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering QCQA: Quality and Capacity-aware grouped Query Attention

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.873091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.873091Z digest=sha256:fb640ae1c658fcaa1af4313e5e00b7f2548d039eabd94b416cb8ec9e107add2a

Observation d693ca11-8206-4880-b0da-41d86b5d3cf7 · outbound

This paper cites GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.879572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.879572Z digest=sha256:0e73d4c29e1ee8da7ac545829cd6b92a311db03b1912b6e4154c90c49e258802

Observation e40f040f-1e6d-45c2-ab28-b6db27918ed3 · outbound

This paper cites an unresolved cited work.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:45:10.757623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T10:45:09.884072Z digest=sha256:0c9ae3a427cc61d3e2bc2d4e9d1b277010026a39354ea948f7343bd3b909bcde

Observation 3b2606d6-5df9-419f-af99-9e6c5e097be2 · outbound

This paper cites an unresolved cited work.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:45:10.733674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T10:45:09.889818Z digest=sha256:6d12225f06dc4cf0c2c455c5803dac9ad42b203741bc218fe9524b043d383db1

Observation 74362e38-9bcc-4d39-8626-bfe1b68a55ea · outbound

This paper cites an unresolved cited work.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:45:10.715024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T10:45:09.896238Z digest=sha256:33a7009cd85d5be00414e998c90915c910f1e61fa00e78b3c4d8c10eba1fdd01

Observation 5f34c911-4420-46cd-a5ae-21abb5158006 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Fast Transformer Decoding: One Write-Head is All You Need

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.901527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.901527Z digest=sha256:a40b74582441ce0be04cc15501619add9a661f207194c65e38cb78fad0da21a4

Observation c9696f4f-1dba-44c6-8b5a-95c3d630a868 · outbound

This paper cites FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.907446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.907446Z digest=sha256:35788721b77913484a958bd3a175f63f3c1773da41b59e206790528904886c46

Observation 59a20d92-ce06-48fe-866a-4610cd1b7b11 · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.914455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.914455Z digest=sha256:5e91c8fe9257a510217e88f85a6b7f0c03653f354fd8dc85e8ff68d98d353e16

Observation d6cb2726-e83b-48eb-b906-833c6a995ec5 · outbound

This paper cites an unresolved cited work.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.920790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.920790Z digest=sha256:b3ce053c69e04fd1b17f90db0bfb19efe5ae3bbc8d3e5f2626e67867f575e221

Observation 1396b045-3cf4-47a9-a74b-d7326c79c876 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.926851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.926851Z digest=sha256:e2e99fe09f8991b166d54daa4f33f0b6f0991ee8dc28b0d822b3c83f6f236016

Observation a773ade7-6a5e-4d4b-bb47-81844a9ece5c · outbound

This paper cites an unresolved cited work.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:45:10.695637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T10:45:09.933063Z digest=sha256:d08e0eda27a3f1b52f9b50097213dc1228e894ea31f3b0c25e602b48ee83b22c

Observation 1a3c3d1d-0243-40c2-ab7c-11379f14c19f · outbound

This paper cites Attention Is All You Need.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Attention Is All You Need

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.938312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.938312Z digest=sha256:56d782fd645d90cd9c8c816d1f01c7ea32108f26bf511bc3498569328f490790

Observation bd5e42cf-18cb-4fea-9e28-bde12a9c36a8 · outbound

This paper cites Cohen, Ruslan Salakhutdinov, and Christopher D.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Cohen, Ruslan Salakhutdinov, and Christopher D

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.943240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.943240Z digest=sha256:697feb2c99bd0c8a1db12cb2fbe9acfc3457ef82059920d0866d82116e8b1f2b

Observation 0f93231f-9814-4dc5-85c4-ecb43a2b7e43 · outbound

This paper cites Effectively Compress KV Heads for LLM.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Effectively Compress KV Heads for LLM

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.947823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.947823Z digest=sha256:63e9a3df25d8ff565236b7eaee08b8f1aedb846ba23c70e12f5b7e5221dda656

Observation 2c7daa70-4d90-4fdb-8a50-23d4b97d8245 · outbound

This paper cites Accelerating Training of Transformer-Based Language Models with Progressive Layer Dropping.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Accelerating Training of Transformer-Based Language Models with Progressive Layer Dropping

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.952447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.952447Z digest=sha256:1511cd9ffd2e2ec863f40f794235bbde48712c91aa9b85ca2568b557e0638407

Observation ee8404e7-937d-4390-a1da-bb201b4e0a22 · outbound

This paper cites an unresolved cited work.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:45:10.649641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T10:45:09.957404Z digest=sha256:081f059dbd0d642f19ea564c24b402c5878833b794ee74b96188ae523ce932ee

Observation 8a744468-b86b-4f68-96bc-c8900601386b · outbound

This paper cites online" 'onlinestring :=.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering online" 'onlinestring :=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.964123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.964123Z digest=sha256:8792a0fdd2c4da89fe3169614ff5ea0a60089bcfea03c35c09deea53808c68ab

Observation 1c9ebf99-05a5-4efc-9c46-b326a0197ddf · outbound

This paper cites write newline.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering write newline

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.969345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.969345Z digest=sha256:b4728bed4bb93859eb04f4e979eeb5b224150d8a58486e28f617b2eab4426577

Pith citing papers

Observation 2d0a32dc-d528-42fe-83d6-2749609e6e03 · inbound

S$^4$R: Selective Sampling, Subspaces, and Sparse Reconstruction for Compressed Long-Context KV Caching cites this paper.

S$^4$R: Selective Sampling, Subspaces, and Sparse Reconstruction for Compressed Long-Context KV Caching TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T00:47:32.072537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-05T00:47:31.653881Z digest=sha256:914e62d40b090ac84d6cbc3c86fb3bb1a11fffca6f6c66348e6a36f24d7c9aaf