Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 35 inbound Pith citation observations for arXiv:2407.18003.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T18:49:24.114689Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 5219d9de-e731-44a0-90d5-02d903ff308e · inbound
LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d5b3f2d9-d1dd-4bdf-a1ed-cc660c9f4577 · inbound
PoM: Efficient Image and Video Generation with the Polynomial Mixer Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85aaef17-8091-466c-b1e5-1ef66ac6ee03 · inbound
Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8c53615-97bd-437a-8068-f3f255f0c4e4 · inbound
A Survey on Large Language Model Acceleration based on KV Cache Management Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b92d5962-84f4-4bbe-a72c-02ea9c1698b1 · inbound
Top-Theta Attention: Sparsifying Transformers by Compensated Thresholding Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9d6c010-c1f2-4d1a-ad8e-4885cbdcba68 · inbound
Judge a Book by its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a6650d7d-366e-4a5c-96bd-68c48e7e7dc7 · inbound
Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 156
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a27890fd-b34e-4647-ad95-b4702abfd465 · inbound
Effective and Efficient Schema-aware Information Extraction Using On-Device Large Language Models Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4897a087-dd91-41dd-a307-597c18fc83a6 · inbound
Semantic Scheduling for LLM Inference Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 788f684c-688f-4c76-a0d2-c0d1d8010cba · inbound
MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93a778ed-c57c-4e7a-ad1c-0000ebe9a475 · inbound
CommVQ: Commutative Vector Quantization for KV Cache Compression Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 461a5b00-b6bf-454a-900d-d65b113bee91 · inbound
Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34f513a7-45ee-4117-9607-8b2ff82b929e · inbound
SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f37ec339-0b23-4a03-bf5b-9a5ccd011380 · inbound
DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2932188b-feb0-4505-b425-00706e65caa3 · inbound
DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a39092e4-3a0b-44e8-ade1-6d67846053ab · inbound
EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94b1b7e0-2608-4c71-b194-af0574f455d9 · inbound
Voice-based AI Agents: Filling the Economic Gaps in Digital Health Delivery Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eafe9d99-b0e2-4280-91ad-76196c72b741 · inbound
A Distributed Learned Hash Table Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 491957c2-5a6c-460d-9d6c-9cc99aeca3ed · inbound
Adaptive KV-Cache Compression without Manually Setting Budget Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d25a45b9-ab71-462e-8b35-d99d6b246da4 · inbound
PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64453fe3-4c28-4191-8b70-bc5a7136d4cb · inbound
EvolKV: Evolutionary KV Cache Compression for LLM Inference Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2eb1b009-0208-437c-85a3-eaf280a9a3db · inbound
OjaKV: Context-Aware Online Low-Rank KV Cache Compression Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b21af5fd-a71a-4e57-a297-453a0d1f8cd8 · inbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ef827935-40c0-4108-87f3-cef628ce9116 · inbound
EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f78f4f86-0f53-4092-a17d-fcd360ec2729 · inbound
Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 842f3199-2da6-41b7-be1e-3e9db7153cfb · inbound
TriAttention: Efficient Long Reasoning with Trigonometric KV Compression Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5da7ee58-3e80-436a-8159-41aac99f4f6c · inbound
How Much Cache Does Reasoning Need? Depth-Cache Tradeoffs in KV-Compressed Transformers Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8dce6e59-b5f3-4860-ba32-eb1112d25229 · inbound
Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2b58e1d1-177a-4766-91d7-f114ac55f69b · inbound
On the (In-)Security of the Shuffling Defense in the Transformer Secure Inference Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2d31888f-da9b-495c-a3fa-7b03b3942266 · inbound
Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 168
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7804a3b1-2b9e-4312-924b-2cf77aa08c07 · inbound
Agents Should Replace Narrow Predictive AI as the Orchestrator in 6G AI-RAN Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 444c5df5-e6c5-478e-b03c-62bc57f60cfa · inbound
FlowNar: Scalable Streaming Narration for Long-Form Videos Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 95bdb966-0c58-4681-837b-d792a80376a8 · inbound
FastTPS: An Optimized Method for LLM Token Phase for AI accelerators Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c6b5074-bdbc-4779-8613-04ebeaf226a6 · inbound
CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f9d27ab-a6ec-4cb4-b4e3-a59308de6cc1 · inbound
KVDiagnosis: A Diagnostic Benchmark for KV-Cache Compression in Long-Context Language Models Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.