Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T23:46:02.187649Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2412.02252.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T23:46:02.187649Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e527d8e1-16c0-4cf5-8462-f8d1a65b1883 · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dffe21b-8859-475d-94c4-5c36025c6c9a · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity GQA : Training generalized multi-query transformer models from multi-head checkpoints
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af388563-0fd1-4914-b83f-c64befbb3d96 · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity L -eval: Instituting standardized evaluation for long context language models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94bc9141-ae7c-4e15-8733-7cca1e4f12b0 · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity L ong B ench: A bilingual, multitask benchmark for long context understanding
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6855b94-fefa-4307-8602-bd8b7a9b0790 · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Codeplan: Repository-level coding using llms and planning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b3fd7210-7bd1-4343-a18b-21d1373c918e · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Longformer: The Long-Document Transformer
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8604c5cb-b478-4d7a-9a74-16ea1eda6e5c · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Leveraging redundancy in attention with Reuse Transformers
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87ff2fb3-099b-47c4-8c1a-56e45b31dc06 · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98db3a0e-e6b3-4768-b801-d984748f683a · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 648e1687-64f3-42c1-9dd0-29fe3dfb6c8e · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity The Llama 3 Herd of Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6cf978e-0eb9-44c4-9592-8fb91282da87 · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9bd01b4-c9c0-4188-846e-c488d9910d89 · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Challenges in Deploying Long-Context Transformers: A Theoretical Peak Performance Analysis
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97576a13-ca55-44f5-8fa1-338fa89e2dd6 · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity How to train long-context language models (effectively), 2024
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d91c8be-ec2f-49e7-a0dc-09b627cdd68e · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8416f7e-c70d-404e-8e23-107bbe694135 · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity LM -infinite: Zero-shot extreme length generalization for large language models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ba4f3af-5ff8-430e-ba6d-531a423b4017 · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 726bed92-0016-47fd-bf64-5646b9ccce34 · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8abda245-0878-4b14-a047-106a40160817 · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity L ong LLML ingua: Accelerating and enhancing LLM s in long context scenarios via prompt compression
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c712d5a-b9bc-4d6b-910d-28e376a5e614 · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Compressing context to enhance inference efficiency of large language models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d53162a3-2441-42cc-b969-463bd21dfebe · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity SnapKV: LLM Knows What You are Looking for Before Generation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3c1ed64-65ff-48c9-b87f-ccf31b814ee6 · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Awq: Activation-aware weight quantization for on-device llm compression and acceleration
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f33968a4-7f52-4a65-8c53-7783935bb369 · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab5fa997-8eea-4f97-b365-6962333cce6e · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93b7599b-0ce8-460a-9224-ba235a9dcbb1 · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b59d75df-2bc9-45be-ba7a-312c438fabb4 · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity QLLM : Accurate and efficient low-bitwidth quantization for large language models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c89effdd-193b-4566-b17e-66c90c46a7be · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity and Liu, B
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0b4de720-5d2c-4470-9e81-52b027644318 · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity The jensen-shannon divergence
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3619e8c9-1702-46f2-b142-ebb2aedf3158 · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity V., Qiu, L., and Zhang, D
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81d264c3-548c-4428-b888-beaf6af0b3d0 · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Pytorch: An imperative style, high-performance deep learning library
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b050983d-e7b2-410d-9863-633e7927f744 · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Efficiently scaling transformer inference
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e497bba-cdb3-4c0f-91b4-4d73d0e00615 · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Zero: memory optimizations toward training trillion parameter models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45c9968c-51dc-4ab1-86fc-214c971c4686 · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35ded138-c09d-4515-9700-f63b0c235c4a · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 158d2ced-a283-4de7-8ea4-2364c600a9f8 · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Fast Transformer Decoding: One Write-Head is All You Need
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5459b65-f3c2-4d98-bd24-5c01e8d3d8cf · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Flexgen: high-throughput generative inference of large language models with a single gpu
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d672a08f-02e0-405a-9bc9-9f60643d96df · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Dolma: an open corpus of three trillion tokens for language model pretraining research
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecaf3672-7514-4d70-8c26-d46b392151e2 · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity RoFormer: Enhanced Transformer with Rotary Position Embedding
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39cb8619-c1c9-4898-9287-a94f10a8ec0a · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity QUEST : Query-aware sparsity for efficient long-context LLM inference
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8bd4e694-b592-4ecf-84b5-5cef6d10d38e · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Gemini: A Family of Highly Capable Multimodal Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9abecf7b-6a21-4e70-9c1e-f0a740fd5e5f · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity LLaMA: Open and Efficient Foundation Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb66d2cf-2166-4468-b377-69189d136794 · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61f7bfbc-d96c-4484-86fb-599962c76105 · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity N., Kaiser, L
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6599008c-d467-400c-8a03-864bac047884 · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Transformers: State-of-the-art natural language processing
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 637c17ae-8609-4060-bd5e-ee9e6e7a49c8 · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity and Tu, K
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 066b066f-a3f8-4d08-913b-dcc30296097c · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Efficient streaming language models with attention sinks
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bee1dc97-650a-4ec2-b3ee-320486b3b22d · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Sharing attention weights for fast transformer
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d73e4f0-05ec-4fff-8adc-4ef91e667dfa · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity A., Oguz, B., Khabsa, M., Fang, H., Mehdad, Y., Narang, S., Malik, K., Fan, A., Bhosale, S., Edunov, S., Lewis, M., Wang, S., and Ma, H
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3547b98f-4cdc-4e26-9900-7e289f56daf2 · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity B ench: Extending long context evaluation beyond 100 K tokens
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dd0b3a7-e200-4e85-8a80-b2346180c1a2 · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72a49ccc-cccd-4788-a785-97374f2dc53e · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity H2o: Heavy-hitter oracle for efficient generative inference of large language models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1359955e-27e9-4492-9b9e-c2e319126b9c · outbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity write newline
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.