Pith. sign in

Paper Citation Record · LEDGER

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity

As of 12 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2412.02252.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.02252 v2

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:46:02.187649Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e527d8e1-16c0-4cf5-8462-f8d1a65b1883 · outbound

This paper cites GPT-4 Technical Report.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:01.946174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:01.946174Z digest=sha256:3efb1006e696dc273f8cd3cd368a11e5eb9197ab993f21bb73d5358cf779355c

Observation 1dffe21b-8859-475d-94c4-5c36025c6c9a · outbound

This paper cites GQA : Training generalized multi-query transformer models from multi-head checkpoints.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity GQA : Training generalized multi-query transformer models from multi-head checkpoints

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:01.951909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:01.951909Z digest=sha256:c6275c5a675980c8cb2e337e3d7ff2adfc27283ba315ed7a457f0369d5875b6b

Observation af388563-0fd1-4914-b83f-c64befbb3d96 · outbound

This paper cites L -eval: Instituting standardized evaluation for long context language models.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity L -eval: Instituting standardized evaluation for long context language models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:01.957532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:01.957532Z digest=sha256:16e79cf3d62b8e52863b7805e974e81296c1a18626a27fde403966108a917153

Observation 94bc9141-ae7c-4e15-8733-7cca1e4f12b0 · outbound

This paper cites L ong B ench: A bilingual, multitask benchmark for long context understanding.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity L ong B ench: A bilingual, multitask benchmark for long context understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:01.962711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:01.962711Z digest=sha256:6619033eb088c17ffc05d65ba560bbc25d792e55109a7a1821630734c31096e9

Observation a6855b94-fefa-4307-8602-bd8b7a9b0790 · outbound

This paper cites Codeplan: Repository-level coding using llms and planning.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Codeplan: Repository-level coding using llms and planning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:46:02.969363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T23:46:01.967701Z digest=sha256:f00ff6bdd8e3078be220ef17d5df95707214177805a5807eddd8c8fb4f2de5ec

Observation b3fd7210-7bd1-4343-a18b-21d1373c918e · outbound

This paper cites Longformer: The Long-Document Transformer.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Longformer: The Long-Document Transformer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:01.972899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:01.972899Z digest=sha256:0e7dec1ab54d182ed692089a9f1bb2574dbf656f9bded573fa0721c44079e309

Observation 8604c5cb-b478-4d7a-9a74-16ea1eda6e5c · outbound

This paper cites Leveraging redundancy in attention with Reuse Transformers.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Leveraging redundancy in attention with Reuse Transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:01.978654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:01.978654Z digest=sha256:3ae3d80399cafd82d521eb09da1c5a9163704d1e1e284b8ce543d281f424507a

Observation 87ff2fb3-099b-47c4-8c1a-56e45b31dc06 · outbound

This paper cites Reducing Transformer Key-Value Cache Size with Cross-Layer Attention.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:01.984234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:01.984234Z digest=sha256:2c2f2b0f2e36cef4c37ad954f0a0643f84aaedeac6009024fe68088c0dc7ace6

Observation 98db3a0e-e6b3-4768-b801-d984748f683a · outbound

This paper cites an unresolved cited work.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:01.989136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:01.989136Z digest=sha256:7a4a431361b327dad9406cb09be8575975b932131be8d65289d1bb8e0ca790cd

Observation 648e1687-64f3-42c1-9dd0-29fe3dfb6c8e · outbound

This paper cites The Llama 3 Herd of Models.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:01.993734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:01.993734Z digest=sha256:c27e62b4147ca365d4b39a3665e9283f7e618849f46da0d5bf1d42942ca487c4

Observation d6cf978e-0eb9-44c4-9592-8fb91282da87 · outbound

This paper cites LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:01.998574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:01.998574Z digest=sha256:8c7a4fcfc63683dd98842bb173a555d1f606bbd3d647c76f0e4a072e82a69a4f

Observation e9bd01b4-c9c0-4188-846e-c488d9910d89 · outbound

This paper cites Challenges in Deploying Long-Context Transformers: A Theoretical Peak Performance Analysis.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Challenges in Deploying Long-Context Transformers: A Theoretical Peak Performance Analysis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.003540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.003540Z digest=sha256:841432233faaf4493b89bfe64793dcb62c41300771d78cca0b1c7aa435a5c0c5

Observation 97576a13-ca55-44f5-8fa1-338fa89e2dd6 · outbound

This paper cites How to train long-context language models (effectively), 2024.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity How to train long-context language models (effectively), 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.008707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.008707Z digest=sha256:b1be0710d1f2f3ef0883c57f1058016bb7eac941657eb586181e289963bd0493

Observation 6d91c8be-ec2f-49e7-a0dc-09b627cdd68e · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.014246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.014246Z digest=sha256:1c60f1bbbf5f956a29cac9404a4ae1b56711a30ff248e47305b32f99030e655a

Observation c8416f7e-c70d-404e-8e23-107bbe694135 · outbound

This paper cites LM -infinite: Zero-shot extreme length generalization for large language models.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity LM -infinite: Zero-shot extreme length generalization for large language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.019172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.019172Z digest=sha256:83e59111be8406e5a687c9ea54340983635310e8c190f21b07235a39fbe6e3f2

Observation 9ba4f3af-5ff8-430e-ba6d-531a423b4017 · outbound

This paper cites DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.023498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.023498Z digest=sha256:d6f36955e8c560f7c9496aa97c4a4f50d100f4459992ff62fe0a28abebbb6b70

Observation 726bed92-0016-47fd-bf64-5646b9ccce34 · outbound

This paper cites MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.028093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.028093Z digest=sha256:4edfb576c5700faeffa9ea480b23af5868c32e66e8e5e49e7d5b9235bb22c321

Observation 8abda245-0878-4b14-a047-106a40160817 · outbound

This paper cites L ong LLML ingua: Accelerating and enhancing LLM s in long context scenarios via prompt compression.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity L ong LLML ingua: Accelerating and enhancing LLM s in long context scenarios via prompt compression

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.032467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.032467Z digest=sha256:08adf0386187c561e67f0c4a30618e6940e6dc8a7e60b31c3c2d5b45b49a800d

Observation 2c712d5a-b9bc-4d6b-910d-28e376a5e614 · outbound

This paper cites Compressing context to enhance inference efficiency of large language models.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Compressing context to enhance inference efficiency of large language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.037136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.037136Z digest=sha256:e923f84f53867042eb5de54fc2db1f71a50d637ce8c7158d8aa477c30df25d3f

Observation d53162a3-2441-42cc-b969-463bd21dfebe · outbound

This paper cites SnapKV: LLM Knows What You are Looking for Before Generation.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity SnapKV: LLM Knows What You are Looking for Before Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.041282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.041282Z digest=sha256:cc9c146d591299da2f1a399d65aca2a0ec6880976a8c330e71b509e286d2de6c

Observation d3c1ed64-65ff-48c9-b87f-ccf31b814ee6 · outbound

This paper cites Awq: Activation-aware weight quantization for on-device llm compression and acceleration.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Awq: Activation-aware weight quantization for on-device llm compression and acceleration

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:46:02.945100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T23:46:02.046056Z digest=sha256:80976f5f623bf9aa5c959662b464a9b7d4388741a857f7a0cca34d922341f8f1

Observation f33968a4-7f52-4a65-8c53-7783935bb369 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.050262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.050262Z digest=sha256:5b2b0d48b5ebc145b5e2effcdf2c6eb11c99170e726bf3bd8889b372de0f5580

Observation ab5fa997-8eea-4f97-b365-6962333cce6e · outbound

This paper cites MiniCache: KV Cache Compression in Depth Dimension for Large Language Models.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.056231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.056231Z digest=sha256:a16ad628858287def76bcd0f868e52359fbd5cbfa3af8ea90abf973d0892aed3

Observation 93b7599b-0ce8-460a-9224-ba235a9dcbb1 · outbound

This paper cites RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.061116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.061116Z digest=sha256:d71feb9a7c0a99f8f7a4947f1b2e483cb31298ce14304030c4d5eea9977d9c2a

Observation b59d75df-2bc9-45be-ba7a-312c438fabb4 · outbound

This paper cites QLLM : Accurate and efficient low-bitwidth quantization for large language models.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity QLLM : Accurate and efficient low-bitwidth quantization for large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:46:02.928658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T23:46:02.066259Z digest=sha256:66ed2251c380b6e3a8e2eafee37db82537239f6beda37a6958c3638785d94c25

Observation c89effdd-193b-4566-b17e-66c90c46a7be · outbound

This paper cites and Liu, B.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity and Liu, B

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:46:02.914397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T23:46:02.071944Z digest=sha256:c85f81ab3180a550c64a0a12d3e60998f75006d8cdb9119acc48c84ccb9077e3

Observation 0b4de720-5d2c-4470-9e81-52b027644318 · outbound

This paper cites The jensen-shannon divergence.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity The jensen-shannon divergence

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.077749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.077749Z digest=sha256:5cff75a0a08d5acb5985ed5246a1d2e73bc73256fcc34fd6d360c8b82bbff976

Observation 3619e8c9-1702-46f2-b142-ebb2aedf3158 · outbound

This paper cites V., Qiu, L., and Zhang, D.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity V., Qiu, L., and Zhang, D

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.082175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.082175Z digest=sha256:0a382fe06d544150168dc7f35c5ddc72e841bd6743153eefcdf1450eb1dbfb36

Observation 81d264c3-548c-4428-b888-beaf6af0b3d0 · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Pytorch: An imperative style, high-performance deep learning library

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.086882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.086882Z digest=sha256:1e00351b6fd495dd8011c70a85e640952e855061f0f07faea32a4ebaa5b60b9d

Observation b050983d-e7b2-410d-9863-633e7927f744 · outbound

This paper cites Efficiently scaling transformer inference.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Efficiently scaling transformer inference

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.091264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.091264Z digest=sha256:4a3882c0e17318a93256e90894cc0df5664aad1fc360dd9f5ad910798bc140da

Observation 3e497bba-cdb3-4c0f-91b4-4d73d0e00615 · outbound

This paper cites Zero: memory optimizations toward training trillion parameter models.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Zero: memory optimizations toward training trillion parameter models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.095682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.095682Z digest=sha256:b0bb3ea1afc4b4ac687cfd16667dc1b6b26e9637c7bd830cc1b6d49e0811af62

Observation 45c9968c-51dc-4ab1-86fc-214c971c4686 · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.100260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.100260Z digest=sha256:0094d65f8b91460f3bc0c8842cd8b8b85fab4105ec99581b9fe3f7e8d2a202fd

Observation 35ded138-c09d-4515-9700-f63b0c235c4a · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.104917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.104917Z digest=sha256:b94726faa2b9833e1ecf6eab22f7f42bf4f04bc23487097e523ec675f09ba01a

Observation 158d2ced-a283-4de7-8ea4-2364c600a9f8 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Fast Transformer Decoding: One Write-Head is All You Need

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.109644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.109644Z digest=sha256:8d0b623eb7e342a8039aee40aa306ba38fc374242a804ec8ff7de0acd4fb5117

Observation f5459b65-f3c2-4d98-bd24-5c01e8d3d8cf · outbound

This paper cites Flexgen: high-throughput generative inference of large language models with a single gpu.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Flexgen: high-throughput generative inference of large language models with a single gpu

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.114792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.114792Z digest=sha256:8aac544e3a061fb99dd55274f4d7a938027ed1180be77e6541a95d8252afec98

Observation d672a08f-02e0-405a-9bc9-9f60643d96df · outbound

This paper cites Dolma: an open corpus of three trillion tokens for language model pretraining research.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Dolma: an open corpus of three trillion tokens for language model pretraining research

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.119255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.119255Z digest=sha256:37624e3e33dfe8da13718bdda14421d9406066b674622675fdd4f7d6322e718d

Observation ecaf3672-7514-4d70-8c26-d46b392151e2 · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.123980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.123980Z digest=sha256:2880ceb6546db7a317f363c1250d41f8265e8a69e171522888d470fe970b3735

Observation 39cb8619-c1c9-4898-9287-a94f10a8ec0a · outbound

This paper cites QUEST : Query-aware sparsity for efficient long-context LLM inference.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity QUEST : Query-aware sparsity for efficient long-context LLM inference

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:46:02.852007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T23:46:02.129198Z digest=sha256:e06383787ab83a990a8444546f625dc7df750a3679f677fad6f8eabebacf4100

Observation 8bd4e694-b592-4ecf-84b5-5cef6d10d38e · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Gemini: A Family of Highly Capable Multimodal Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.133622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.133622Z digest=sha256:e91370bd6b90434db905489ded52219c674d9a74b61895fc66d1254cd9958df3

Observation 9abecf7b-6a21-4e70-9c1e-f0a740fd5e5f · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity LLaMA: Open and Efficient Foundation Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.139083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.139083Z digest=sha256:4bfc66aecf5f00ec8b197f3ffcd1ed1389808cff827d93e274b021a3ade28f87

Observation eb66d2cf-2166-4468-b377-69189d136794 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.143383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.143383Z digest=sha256:8a32627c3903b8261a80deb70c04a68c82206f19fb9b6899c67105b87eb3668b

Observation 61f7bfbc-d96c-4484-86fb-599962c76105 · outbound

This paper cites N., Kaiser, L.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity N., Kaiser, L

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.147598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.147598Z digest=sha256:181c112f2adf273649f4215a63d3931e3ab418b62c899f13abc569bce9dd2e5d

Observation 6599008c-d467-400c-8a03-864bac047884 · outbound

This paper cites Transformers: State-of-the-art natural language processing.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Transformers: State-of-the-art natural language processing

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.151899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.151899Z digest=sha256:ef705949a6fad67ee03f19c58e7f21ddf763de4dabd66d059b7d0ebf8f42c99a

Observation 637c17ae-8609-4060-bd5e-ee9e6e7a49c8 · outbound

This paper cites and Tu, K.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity and Tu, K

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.156015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.156015Z digest=sha256:8935ce5b598a67b0159ad27de463977a8969b60c0aae61da29f1bac6da128f12

Observation 066b066f-a3f8-4d08-913b-dcc30296097c · outbound

This paper cites Efficient streaming language models with attention sinks.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Efficient streaming language models with attention sinks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.160318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.160318Z digest=sha256:43d31993f4601619d86c1830ad0299973b8fdc3e5fdc25501f3ae09f46e7c24d

Observation bee1dc97-650a-4ec2-b3ee-320486b3b22d · outbound

This paper cites Sharing attention weights for fast transformer.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Sharing attention weights for fast transformer

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.164871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.164871Z digest=sha256:4aafc2532315167b2e4d52981b0d9851b358ca04d14eaf94011c478e0ae77eae

Observation 7d73e4f0-05ec-4fff-8adc-4ef91e667dfa · outbound

This paper cites A., Oguz, B., Khabsa, M., Fang, H., Mehdad, Y., Narang, S., Malik, K., Fan, A., Bhosale, S., Edunov, S., Lewis, M., Wang, S., and Ma, H.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity A., Oguz, B., Khabsa, M., Fang, H., Mehdad, Y., Narang, S., Malik, K., Fan, A., Bhosale, S., Edunov, S., Lewis, M., Wang, S., and Ma, H

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.169436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.169436Z digest=sha256:d239a490b38a00251c4b75829490bf67aa60dcbe578dc354b4b15f5225ab5c24

Observation 3547b98f-4cdc-4e26-9900-7e289f56daf2 · outbound

This paper cites B ench: Extending long context evaluation beyond 100 K tokens.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity B ench: Extending long context evaluation beyond 100 K tokens

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.174036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.174036Z digest=sha256:674478b5fac7d36af451fed75cc385562b04477531fd2300914bf12fbd4d0993

Observation 3dd0b3a7-e200-4e85-8a80-b2346180c1a2 · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.178568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.178568Z digest=sha256:6474b40d4adc345ad7751dfafb75560d5369599806ce3cb50d02544917e6c73d

Observation 72a49ccc-cccd-4788-a785-97374f2dc53e · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity H2o: Heavy-hitter oracle for efficient generative inference of large language models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:46:02.798121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T23:46:02.183128Z digest=sha256:ab767819677cfacb7b12c07b50806cbe20ce150ce9deacf61635a414d4010673

Observation 1359955e-27e9-4492-9b9e-c2e319126b9c · outbound

This paper cites write newline.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity write newline

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.187649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.187649Z digest=sha256:08bf364f2e7cab12a21e2c89c6059aeae45438221dd0d4007d3bac76b7c2a74d

Pith citing papers

No inbound Pith citation observations are available.