Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T20:34:01.382301Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 6 inbound Pith citation observations for arXiv:2501.19392.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T20:34:01.382301Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-25T04:51:04.068354Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-25T04:55:24.203465Z
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9ad5d31d-d34d-4027-8f5b-c90498d5144b · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 251e4d38-cb7c-48ca-8464-a9042a9ff396 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af7d1eb7-9eae-4398-a32d-2dc231825daa · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Understanding intermediate layers using linear classifier probes
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d464322-1ec4-4d62-add5-38a920dc34e3 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ab4ce44-4aa3-45d0-8dbf-615a79164ca4 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fae5ebc-01a5-44ce-8995-db483bb3304d · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models M., Gebru, T., McMillan-Major, A., and Shmitchell, S
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9e72fe70-0f02-408e-ad53-19ad07a7a089 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Palu: Compressing KV-Cache with Low-Rank Projection
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db045350-6224-45c4-a0ee-2771b8660ea2 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models PrefixQuant: Eliminating Outliers by Prefixed Tokens for Large Language Models Quantization
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed214d9c-c945-4b69-b3b3-c80e9c68c2e9 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7230ed0d-e4f3-4763-8368-4242e1735827 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models QAQ: Quality Adaptive Quantization for LLM KV Cache
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ed2fe7f-cb97-493d-871f-444192003f39 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a74aea1-516b-4ebb-8de6-e9e59c0f7091 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models The Llama 3 Herd of Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d792f3f2-6b06-45a3-a052-5d75ce161ae3 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Towards Measuring the Representation of Subjective Global Opinions in Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43925cb0-b1e0-4fa5-9897-df956ab7e938 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Extreme Compression of Large Language Models via Additive Quantization
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a374d1a-a87f-403a-b8ae-0f658c2b9086 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c7041fe-7771-41b7-8d3c-ab6a53ac0f1e · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2276e2c-5d4c-46cc-81dd-afd48dae68ad · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 95a1ce10-7f8c-4dca-8d26-f4291b42853f · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Fast matrix multiplications for lookup table-quantized llms
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b8b6f30e-69ab-4186-804d-ceb69798c826 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63704b6d-d0da-4d2f-9966-e36b374e5998 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82fe702c-a446-44c3-8773-9407f9878293 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Stochastic distributed learning with gradient quantization and double-variance reduction
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fa84ad4-54f3-4f07-8dc1-78f7ea39e1a7 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Optimum-quanto: A pytorch quantization backend for optimum
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d80ec495-591c-4b2b-98d8-d3be9eca364f · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddf7a799-c6fc-4f2c-95bb-f12f480a8f7f · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models H., Gonzalez, J., Zhang, H., and Stoica, I
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 363d5d4e-f8ef-4346-a9f0-3e4ee7b35342 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models A Survey on Large Language Model Acceleration based on KV Cache Management
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bf02429-cf6d-4568-9e3c-7552eecc6d51 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models SnapKV: LLM Knows What You are Looking for Before Generation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef36e9a9-22d8-4947-a948-1d7fb234a758 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models SCBench: A KV Cache-Centric Analysis of Long-Context Methods
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4d0dc2c-42d4-46ac-9b48-cf2b1d6fb4cd · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8100fd37-f4e5-42b6-8a23-f126bcce6ec9 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3c491b0-3ab9-45ff-a0f2-11633560713e · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models IntactKV: Improving Large Language Model Quantization by Keeping Pivot Tokens Intact
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70c22dd8-2852-4c04-93f9-ee6de6058fd1 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be5b2b39-17d7-4e12-baa0-d2ac0ad6d211 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0f3d9b0-5619-4ccf-a814-6d412e9edcfc · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Pushing the Limits of Large Language Model Quantization via the Linearity Theorem
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec4bf091-6217-4858-9150-f8330403e6bc · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Pointer Sentinel Mixture Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d05ec6d2-877e-4df3-a98f-56f099406ef1 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models PyTorch : An imperative style, high-performance deep learning library
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d5100856-34e6-483f-8383-9210abf52b3d · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Your Transformer is Secretly Linear
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f446cae9-71a6-4b8b-99eb-2b2ff23675c6 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5ecc08a-e6ac-4eec-84ef-bda8409605dc · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Societal Biases in Language Generation: Progress and Challenges
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a73bffc1-c291-4e1b-b4cc-510722dcd859 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models RoFormer: Enhanced Transformer with Rotary Position Embedding
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9b5555b-cc54-48a1-87a2-84fa96ffc1e3 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Qwen2.5: A party of foundation models, September 2024
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd55ea0e-1543-4e06-87e7-f2ab5767f297 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models LLaMA: Open and Efficient Foundation Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9926d273-81d8-46b8-9d2d-605d88f75458 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c3e592f-d489-49e3-825a-8c282473eae6 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f89d223d-34de-4a17-bc3c-b7f9018f125b · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9cefe2c7-1865-4dcd-b5e4-bdf3cb2700a8 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models GPTVQ: The Blessing of Dimensionality for LLM Quantization
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1a1c0d2-a727-469b-95c0-61804011d507 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Attention is all you need
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c20bd02e-0937-454f-a476-5e06d12ab2bd · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Unresolved cited work
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35ce49bf-5326-4b1f-8bab-7cc0fa564058 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Ethical and social risks of harm from Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27820dda-1aaa-41e4-b023-1620dc7e8ae4 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f5ec474a-361b-46dd-98f6-707024ab02b4 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models HuggingFace's Transformers: State-of-the-art Natural Language Processing
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2645100c-35ab-44b7-a89d-6b2d880f895f · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Efficient Streaming Language Models with Attention Sinks
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 685d429d-b916-4c2d-b21e-3e38f7f79ca4 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Qwen2 Technical Report
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e36af4c4-b7d4-4940-8b8d-a614cfbfadfe · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e75cced3-0e74-4d9b-a114-f262531871d2 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7761c196-453c-4558-84f1-8e046a244107 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09da41e6-ac57-436c-8e5e-4482df3496cc · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6291d10-d012-4508-af3a-b78590459ccf · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models FDC: Fast KV Dimensionality Compression for Efficient LLM Inference
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b458603b-e590-4201-bbfc-0d58d785b0d5 · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models H2o: Heavy-hitter oracle for efficient generative inference of large language models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b11442c7-e928-4a77-9eee-1621558020bd · outbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models Red teaming ChatGPT via Jailbreaking: Bias, Robustness, Reliability and Toxicity
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c036e79-b643-41f2-aef1-56bf2295a7f7 · inbound
KV Cache Offloading for Context-Intensive Tasks Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1c7f9c93-c852-478e-a569-06591cefc921 · inbound
KV Cache Offloading for Context-Intensive Tasks Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7779274f-0729-4706-a7d2-54b82190afe9 · inbound
KV Cache Offloading for Context-Intensive Tasks Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 34b564f1-c187-439d-8388-da7bf1f7fea5 · inbound
KV Cache Offloading for Context-Intensive Tasks Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 37f7136f-56d8-40a3-b40a-157028fe8598 · inbound
RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 074dd5fb-7213-4fb4-8bac-8cce632be527 · inbound
A Simple Plug-in for Improving Eviction-Based KV Cache Compression Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.