Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T14:06:38.022879Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2507.19823.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T14:06:38.022879Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ee6a794a-a478-4d69-b76b-a1f91e63179e · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs 2 In the approximate attention score computation, the computation is performed group-wise
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0414ced1-bf2c-4593-b73d-89948ed9b6be · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs needles" on the task performance. We take “The best thing to do in Paris is buy a fresh croissant and lounge by the Seine at twilight
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 96117a3d-8e7e-4b5d-ae64-ca5412b25625 · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs GPT-4 Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdebe3d6-067d-43c2-837a-d187f24ae57a · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Keyformer: Kv cache reduction through key tokens selection for efficient generative inference
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 606dceea-355c-4ae0-bf69-50c93b47d171 · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Qwen Technical Report
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecfac449-9f5c-462e-a6f0-a0adf0e10d7e · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs LongBench: A Bilingual, Multitask Benchmark for Long Context Under- standing
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8130c8ad-6e46-4174-b1ca-29ef63735ca7 · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Longformer: The Long-Document Transformer
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 389002f8-0351-49b5-8c60-8f50422b2ee5 · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1a40cc01-238f-4c1e-9b65-cfe96b455c10 · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adfc8e26-5816-431e-8d23-388877f32cbe · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs The Llama 3 Herd of Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 950360d6-18a2-4927-9922-c48f3179ece1 · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cca51106-df38-4722-a13a-e74117eb2f5a · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4e7b22e9-2adb-4006-a7d2-5f6170baf825 · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Mixtral of Experts
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fc44e7c-d87c-46d0-83a0-a26cfd1a7dce · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 615aef19-ece2-4278-b6c8-72c8a1b22de4 · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Kamradt.Llmtest_needleinahaystack: Doing simple retrieval from llm models at vari- ous context lengths to measure accuracy
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c4c7c256-85e4-49b5-bfb6-54a9f96df94b · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Beyond Single-Turn: A Survey on Multi-Turn Interactions with Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 540ef15f-0219-466f-b0bd-cb6a60457dd6 · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs FocusLLM: Precise Understanding of Long Context by Dynamic Condensing
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b1537645-a823-4aaf-aae9-821058aae735 · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f1c78db-f234-4b0d-b048-757ca3fc22ec · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs KIVI: a tuning- free asymmetric 2bit quantization for KV cache
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 288d86ab-58a8-495f-a228-5b9ee5ec114d · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs LIFT: Improving Long Context Understanding Through Long Input Fine-Tuning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation db751f23-2830-4e43-b0c2-665b4dc94514 · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Sentence-T5: Scalable Sentence Encoders from Pre-trained Text-to-Text Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5083c6d-33ab-49cf-825d-b81d20220c1c · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Transformers are Multi-State RNNs
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cf1c9358-62f9-4a64-bdc7-b1ac7d5880f6 · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Scikit-learn: Machine learning in Python
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 93f5acea-eed0-44a1-a7f2-d3b38187eb1b · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Web-scale k-means clustering
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation acf8426f-15ab-4ec1-bc9f-acf90e0f47f7 · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Adafactor: Adaptive learning rates with sublinear memory cost
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d8d2ff7a-a944-4ac6-88d1-a65a44d82955 · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs QUEST: Query-Aware Sparsity for Efficient Long-Context LLM Inference
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9b79bf44-ecb7-4a4f-b923-8ee1d1742da8 · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs AsymKV: Enabling 1-Bit Quantization of KV Cache with Layer- Wise Asymmetric Quantization Configurations
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4a804256-821c-4af7-941a-011ea97cf392 · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Gemini: A Family of Highly Capable Multimodal Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b2a5842-f14a-4cf0-827c-9eb9dfbe3fb6 · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11c6413d-c09a-4c77-9d51-63c2b6c1ab0a · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Attention is all you need
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 98f37da9-ea15-4352-a0a8-bdd8e46ac2f7 · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Fast transformers with clustered attention
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0f0aea8d-7a86-400a-859c-ac0a51eaf6ae · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs SqueezeAttention: 2D Management of KV-Cache in LLM Inference via Layer-wise Optimal Budget
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1c1e24bb-8c80-462b-a0a0-6e64696e4b37 · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a73e0ec4-c5a0-47a7-8d9a-0825d23228ff · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Efficient Streaming Language Models with Attention Sinks
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 58d002b9-2bf8-478d-9a61-9193fcab74e8 · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs vTensor: Flexible Virtual Tensor Management for Efficient LLM Serving
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e20b226-1612-4e36-adc7-a2a165d53f6b · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs ThinK: Thinner Key Cache by Query-Driven Pruning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c31314c2-8d44-479e-893e-6badf21f3034 · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs A Survey on Multi-Turn Interaction Capabilities of Large Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fe2a94b-3a14-4169-a780-0272201307f5 · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Long Context Compression with Activation Beacon
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2ae781a-3c43-4df4-9adf-74184f70c325 · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs OPT: Open Pre-trained Transformer Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31fd5540-f7d2-41fc-9bd2-b3a1764fdc57 · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs KV cache is 1 bit per channel: Efficient large language model inference with coupled quantization
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7cb6c8ec-a7bd-4690-8ba2-937a9b53679a · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs Chain of agents: Large language models collaborating on long-context tasks
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 03539e17-272c-4263-9570-932b78338fe3 · outbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs H2O: Heavy-hitter oracle for efficient generative inference of large language models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.