Pith. sign in

Paper Citation Record · LEDGER

Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2405.12981.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.12981 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:22:46.999943Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c00b63c8-c9a6-4544-978d-b68ed14a0472 · inbound

LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation cites this paper.

LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-23T18:33:19.446287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-23T18:31:35.391674Z digest=sha256:829c4eb06370b44328d828a3c3036b1d6ab884206ac3a42426fa37891f34bfe6

Observation ec66c560-a6b2-4182-a552-453dd9e2a2e0 · inbound

PoM: Efficient Image and Video Generation with the Polynomial Mixer cites this paper.

PoM: Efficient Image and Video Generation with the Polynomial Mixer Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:46.999943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:46.999943Z digest=sha256:9d5be972f960b3d80fad2f6d00c6c5286a2fd009f2df183fecc6e60d573f2f1f

Observation baa5f6ec-1d50-4f0b-9471-2fe7120524d9 · inbound

Hymba: A Hybrid-head Architecture for Small Language Models cites this paper.

Hymba: A Hybrid-head Architecture for Small Language Models Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.344744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.344744Z digest=sha256:bed026e3bef6e98a426c137e418949e6d84d60be668c2fa3d1d51ea67c7e5139

Observation 24e3e211-2a95-4ad7-ae8a-add9f701ff2f · inbound

CLOVER: Cross-Layer Orthogonal Vectors Pruning and Fine-Tuning cites this paper.

CLOVER: Cross-Layer Orthogonal Vectors Pruning and Fine-Tuning Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-12T12:13:55.149273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:13:55.149273Z digest=sha256:b933cd1797f61ea536cc203534e8e58cd9ade71eb2a1ef981e81986ffdbef244

Observation 592c3a29-2acc-4656-848d-50800257f37c · inbound

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache cites this paper.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.553067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.553067Z digest=sha256:fd880c5f10a24dd208dbaae0f0ba3b81794dc0597e8aec619dcc5c85198479db

Observation 87ff2fb3-099b-47c4-8c1a-56e45b31dc06 · inbound

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity cites this paper.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:01.984234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:01.984234Z digest=sha256:5e0561a23ace693dc751f7baa85ef7006802232515268142f692e30a3bf536b2

Observation a0785f33-eacf-4f05-ab0a-3c3e6e476661 · inbound

UniForm: A Reuse Attention Mechanism Optimized for Efficient Vision Transformers on Edge Devices cites this paper.

UniForm: A Reuse Attention Mechanism Optimized for Efficient Vision Transformers on Edge Devices Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T23:37:44.871522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:37:44.871522Z digest=sha256:ab47d6e1f4fd49ab94fd331315d4d817a58b8e19a7ab6a3fa38e2c5b9e8ba95f

Observation 3572522b-9059-4164-93e1-a936006ebe83 · inbound

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression cites this paper.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.622944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.622944Z digest=sha256:460c29925a64e781485d5ced92cdd1ec7c0fd947c3c89e9770acd3d46a286e44

Observation a5c07440-d1ed-432c-a602-7b2ea45b4c59 · inbound

DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs cites this paper.

DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T11:57:17.878098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:57:17.878098Z digest=sha256:8f5b4639933e7c051fec8971d75cbdd9a77d2f7f62cc58a11f3139d6a480d556

Observation 357356ce-20ee-44fa-b58b-cab024cf6828 · inbound

A Silver Bullet or a Compromise for Full Attention? A Comprehensive Study of Gist Token-based Context Compression cites this paper.

A Silver Bullet or a Compromise for Full Attention? A Comprehensive Study of Gist Token-based Context Compression Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:29.385894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:29.385894Z digest=sha256:256aa075f03459a336c74791e3acba34a68d8f4143d054a227c924a03a1ceff4

Observation e6d12d0e-1a79-4b0e-8f87-f82d9fe68267 · inbound

Multi-matrix Factorization Attention cites this paper.

Multi-matrix Factorization Attention Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.777756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.777756Z digest=sha256:a5d6e38b6a090c8d8080761e687b5d69d420ab899f40ad54e3a1f25aeb987cbf

Observation 2cc63f34-a742-404f-b05d-86d8dcaa63a1 · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 205

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:48.814000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:48.814000Z digest=sha256:46e9821148841dbd055950ced844fdcc496e813aecc4080b7f4d1e3e124a07f1

Observation c0982b95-36e1-4cb1-8f07-7655903b600f · inbound

Foundations of Large Language Models cites this paper.

Foundations of Large Language Models Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.581590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.581590Z digest=sha256:ae698168d755f76f96b14c12c590a2e4a502a05299092f8ded2563c03d95dad9

Observation 2a8512ca-d088-4f2a-b192-7d6577b5871b · inbound

Sigma: Differential Rescaling of Query, Key and Value for Efficient Language Models cites this paper.

Sigma: Differential Rescaling of Query, Key and Value for Efficient Language Models Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T15:50:21.846487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:50:21.846487Z digest=sha256:581a8eefc2ae05acd7a5560d5a537bef222e1474d58e29dcc215da9a0b33ce59

Observation 60eeb619-40af-4ecb-9862-e62d815b8235 · inbound

Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression cites this paper.

Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:17:31.245179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-23T04:15:36.906263Z digest=sha256:4dc73046f319720786766bde0a94fcf104e02ae5141edf132980b52ae6671721

Observation 1db8d200-bea1-4583-a4f8-07dca902121d · inbound

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach cites this paper.

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:39:41.169681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-12T15:39:40.845703Z digest=sha256:9409fefa01325bbd0418abc1e814e3e011d87ae18a7df76deff7b21b2fb13245

Observation d4937893-d0a2-45b9-8c82-11ad5be1b371 · inbound

Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents cites this paper.

Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T14:13:37.965936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:13:37.965936Z digest=sha256:c64e713d2156a95e92aebbb40057f2dbbfb54111d594cfdbe263a53bcd589102

Observation 30d4b816-ebab-44fb-b757-ee0633875c68 · inbound

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache cites this paper.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.903083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.903083Z digest=sha256:0b4309de63db900d4a389695744618cd6644e2ab0b2a5ad31f937351f4c4b045

Observation 9bbe0a22-fc85-4ac7-81a7-82c15fcbdcf7 · inbound

MUDDFormer: Breaking Residual Bottlenecks in Transformers via Multiway Dynamic Dense Connections cites this paper.

MUDDFormer: Breaking Residual Bottlenecks in Transformers via Multiway Dynamic Dense Connections Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T22:34:07.961504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:34:07.961504Z digest=sha256:931054d8d5b16e345f7fbfbcb241a99963bdde8316c6eb5ea19f0d0990f982dc

Observation cd97bfa7-81dd-4579-b112-b259b2a85479 · inbound

Hardware-Efficient Attention for Fast Decoding cites this paper.

Hardware-Efficient Attention for Fast Decoding Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:29.855424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:29.855424Z digest=sha256:2b36d024a39ed9a01d519465a907da89ec69c5cb7f285b4dfcc59ec90e321095

Observation b6a8c350-03bf-426b-891a-a5d23d81a1fd · inbound

KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding cites this paper.

KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:20:25.790122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:20:25.790122Z digest=sha256:24106a65cce89dac61c888e3c2e38d4035daffb1bc4bed284968d1769da668d5

Observation 345bb7a4-d061-4786-9478-af7175703411 · inbound

CaliDrop: KV Cache Compression with Calibration cites this paper.

CaliDrop: KV Cache Compression with Calibration Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.661526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.661526Z digest=sha256:d4c7781d923cdf0c57636fcc86830bed6be080bd3a1ee8505f2a43d619a78fe2

Observation 6d5022ab-4df1-4ddb-b02e-098ef4e89e67 · inbound

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference cites this paper.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:09.661574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:09.661574Z digest=sha256:039c4ef224296f1fd207d669df61ad91e96a86f2424f790639d809bda036aaec

Observation 0afe267d-674f-4f95-8d75-5f91593e0c76 · inbound

Stochastic KV Routing: Enabling Adaptive Depth-Wise Cache Sharing cites this paper.

Stochastic KV Routing: Enabling Adaptive Depth-Wise Cache Sharing Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:58:12.051857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T19:56:48.015363Z digest=sha256:288edabe3a38c7840f4ea5d6ed05f6c5523ad4c693ea39c7673c623e34327a91

Observation 75f6d43b-16fc-4490-bdf2-28c20b0612d6 · inbound

Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models cites this paper.

Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:45:57.887306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T02:44:36.697450Z digest=sha256:1e2950d2cb3de8882fe05408f8af4c07d592a6bc6b3014b547e9fda7b7c0a831

Observation 3673d03d-8509-48e2-be5b-673b600ca091 · inbound

Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models cites this paper.

Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:03:50.629580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T23:01:17.957634Z digest=sha256:ac9e223824a6864f3457b1213058012729c115bd4b49ee66789182946924485d

Observation 3a38211c-62ab-4de3-a429-0a1f48999f38 · inbound

Do Value Vectors in Deep Layers Need Context from the Residual Stream? cites this paper.

Do Value Vectors in Deep Layers Need Context from the Residual Stream? Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:16:23.469169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T14:35:48.292081Z digest=sha256:cf9a04fe8c37636ad7aac230d2473526cc7fc011692dd6fe2cee5894ab484624

Observation 306faa8c-84e0-49b9-808f-58f778b2b301 · inbound

Do Value Vectors in Deep Layers Need Context from the Residual Stream? cites this paper.

Do Value Vectors in Deep Layers Need Context from the Residual Stream? Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T12:41:19.361685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T12:41:19.361685Z digest=sha256:62ed56d14071c5f432079b7c1833c49e4ca06fa814c4d3a234310f67e2277b8e

Observation 94bdc96c-9f3f-447c-be23-74ce6c3bbf63 · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:44.309590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:271e026e15a1cb9cce3370989988f6b2af74f3e2ec0d794d10e22688fcbfd708

Observation 347be885-1d25-4ac8-91f9-9c10bb0815f5 · inbound

SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling cites this paper.

SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T09:51:03.342951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T09:51:03.342951Z digest=sha256:a5a7a8a6edb8f75fc916b255f781613496279db85f197e8b8d7e83122a70011e