Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T13:52:48.849977Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 4 inbound Pith citation observations for arXiv:2412.12706.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T13:52:48.849977Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:22:44.472480Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
59 of 59 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
Observation 4aac6843-46ab-4b38-8ab9-5c78d5a22b2d · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4e2eee3-477a-4c79-abd7-3dfbddb30765 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5ab3da4-7e46-4448-8936-1caf9816542f · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a250ea3-447b-4b62-bfc6-eae6475f56d6 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3572522b-9059-4164-93e1-a936006ebe83 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d966f8cd-083a-4ced-b1b4-aa87034af389 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 319992dc-705e-44b7-a0e3-c735e20097dd · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3839389-59a8-44b1-a9d8-8c8105e0c4fa · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression A Simple and Effective $L_2$ Norm-Based Strategy for KV Cache Compression
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e61393d-9a6e-47bf-ab4c-ebfa3a863255 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression QAQ: Quality Adaptive Quantization for LLM KV Cache
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0d00c62-bf97-4e3c-a472-6052a117034c · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression The Llama 3 Herd of Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5054f91-5156-4c1e-871a-9067277221a8 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Change Is the Only Constant: Dynamic LLM Slicing based on Layer Redundancy
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a050f92-c402-4a6f-b993-cf53f3901440 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Multi-News: a Large-Scale Multi-Document Summarization Dataset and Abstractive Hierarchical Model
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfbcca3e-0405-4041-bb9b-fc3170baf4af · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a5cee16-8ed1-42ca-927c-7fc156bd969e · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Challenges in Deploying Long-Context Transformers: A Theoretical Peak Performance Analysis
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc93497a-f592-47ee-b778-755f4fe91740 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd5c9e8b-2929-47ae-836c-42a1e62494ef · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0cceca8d-4697-44d0-b08b-360710715d99 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cc8db87-9ffb-4a0b-aba1-06d773342aba · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85512ca1-b096-4e75-a4e7-4dcf9e2c5833 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3535e96e-a105-4684-b7c5-8b957cf300b8 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0203ab4c-f42e-4cb9-abcf-494cce6135c9 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 326f13d4-abc0-4fe5-a144-9b06272784dc · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Efficient Attentions for Long Document Summarization
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4a01b9b-c08d-4885-8d73-ea8cebbd24a3 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Mistral 7B
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61658d2a-b389-49c8-9c35-c44c75e4e0f4 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d9b74a7-66b3-4087-bc06-5894db2567de · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae71be05-5aad-457e-b094-8edc2005b616 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39ef9cbc-8728-42a1-b94b-03cb1052f3a9 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d2dbc50-4cd7-4c6f-bffb-b28e3030b5d8 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 074a42db-77c8-4044-b3d8-48237ff5f0cb · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression SnapKV: LLM Knows What You are Looking for Before Generation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47f656c3-4bbf-4ead-b092-3adb7a000df4 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca6baa67-4dcc-4e11-bed4-6f5eabcda06c · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acf45d95-2fe2-4f44-8501-d9de12c47191 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22dc7c87-611a-4b1c-bad4-dab585df6687 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation adb61239-7bbc-48ef-b624-cd985096f4f7 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression ShortGPT: Layers in Large Language Models are More Redundant Than You Expect
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d971f70a-682a-4856-bcfa-f608c3334d5f · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Transformers are Multi-State RNNs
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6191a40-59c5-49d3-bf19-55df5188ce17 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e9dae6ba-9077-4563-aeac-939eae87cb71 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 879af09d-5eb5-47d0-9bd8-6bd52689fa4d · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression On the Efficacy of Eviction Policy for Key-Value Constrained Generative Language Model Inference
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41705c5d-61cf-4c5e-b7a3-fb74a5dd1d4e · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Code Llama: Open Foundation Models for Code
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1570d5ce-cc5a-4149-969a-36a576e1ca9a · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Fast Transformer Decoding: One Write-Head is All You Need
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86b5861d-d14e-4027-861a-4da00ddfff1b · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 608b902e-6534-470d-99d3-5c17230e5068 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f46d77db-056b-4eba-8efa-464c6fefa035 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 433572ef-8c99-4af6-aea1-0aa3ca1a42a8 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression RazorAttention: Efficient KV Cache Compression Through Retrieval Heads
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db3e8c39-8a61-41de-8197-23b10c592cb4 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0ef2ad29-0ca0-4bc1-b811-e926fd7bf045 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Layer-Condensed KV Cache for Efficient Inference of Large Language Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81713329-28f6-4f7d-a9db-cca3c5548cfa · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4524f31-26d4-437d-ae31-51fe80d91a8a · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2ec4fe1-747a-4a2b-8a12-5f7720a90062 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66ba1ade-5ea3-41f5-95cf-c3d6a593eddc · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2225f4a7-6c56-4669-8861-843b7c683579 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9578d8f9-5a7b-4e17-a3bd-3178a851a30d · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d08c666-29be-480f-b0ce-4b2dfa2b0909 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 88e68680-e31c-43d1-95fa-47e4dda70b40 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression Unresolved cited work
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d866bc3e-e5b0-485c-8a25-d6aef81ebdcd · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression QMSum: A New Benchmark for Query-based Multi-domain Meeting Summarization
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 138c695b-ded6-41ab-b44a-1d52440a28cc · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression DOCBENCH: A Benchmark for Evaluating LLM-based Document Reading Systems
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 379f3c06-11b5-4a3f-8aa5-1a0921f3d02b · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b35799b-4dc9-45e2-905a-3e88ad523a71 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression online" 'onlinestring :=
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17ab0270-cadf-430d-8938-fcd03c62b8c6 · outbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression write newline
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f28759a-2cf3-4dd8-8106-86b6acf5f8e9 · inbound
Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a2bfef6-566d-478b-b29d-f25090467f8b · inbound
RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d1656047-a3a4-48e7-9acf-17d8fd237be7 · inbound
YouZhi: Towards High-Concurrency Financial LLMs via Adaptive GQA-to-MLA Transition More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8452ac4d-a63d-4cdd-9813-4130c14201b1 · inbound
DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression
Reference 145
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.