Pith. sign in

Paper Citation Record · LEDGER

Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 35 inbound Pith citation observations for arXiv:2407.18003.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.18003 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:49:24.114689Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5219d9de-e731-44a0-90d5-02d903ff308e · inbound

LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation cites this paper.

LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-23T18:33:19.467156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-23T18:31:35.391674Z digest=sha256:2e38a95fde418ae0b9c299483144d954cc8f75054a75c44e806e5abf6ad9f82c

Observation d5b3f2d9-d1dd-4bdf-a1ed-cc660c9f4577 · inbound

PoM: Efficient Image and Video Generation with the Polynomial Mixer cites this paper.

PoM: Efficient Image and Video Generation with the Polynomial Mixer Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:47.127086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:47.127086Z digest=sha256:50efddfb38c5a4bc5233a1754b141527970920d36ecfff12a2bd8dc290ae6924

Observation 85aaef17-8091-466c-b1e5-1ef66ac6ee03 · inbound

Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification cites this paper.

Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:00.325854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:00.325854Z digest=sha256:30ecd45de4915fc93802078277472a60085d3fc789e34b079b787cc47daefec3

Observation b8c53615-97bd-437a-8068-f3f255f0c4e4 · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:46.742080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:46.742080Z digest=sha256:54d8d9058ab350703d74eda18caa2e6a365c8e2bf07c936acb71d6525b977fd1

Observation b92d5962-84f4-4bbe-a72c-02ea9c1698b1 · inbound

Top-Theta Attention: Sparsifying Transformers by Compensated Thresholding cites this paper.

Top-Theta Attention: Sparsifying Transformers by Compensated Thresholding Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T05:29:16.660204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:29:16.660204Z digest=sha256:37a402cac76f57f07be919e302a23e8829ea83a5759fa5887188e8484d2356b3

Observation d9d6c010-c1f2-4d1a-ad8e-4885cbdcba68 · inbound

Judge a Book by its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription cites this paper.

Judge a Book by its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:25:19.420768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-23T02:23:22.682357Z digest=sha256:00e00f445b65cd4619e4e40b1c56e8c2f3ffa4e3bfa1f222cbebe17ca02d1345

Observation a6650d7d-366e-4a5c-96bd-68c48e7e7dc7 · inbound

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models cites this paper.

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 156

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:29:56.798694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-14T01:29:56.480020Z digest=sha256:9145c73b2d64ff1ad11e0dde9de911468b33f87bb2912205f2610bb25669c5ce

Observation a27890fd-b34e-4647-ad95-b4702abfd465 · inbound

Effective and Efficient Schema-aware Information Extraction Using On-Device Large Language Models cites this paper.

Effective and Efficient Schema-aware Information Extraction Using On-Device Large Language Models Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:15.745107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:15.745107Z digest=sha256:722251adf7fd4d304c658c259b28c2186e92add1172fde924a18431ea17054a3

Observation 4897a087-dd91-41dd-a307-597c18fc83a6 · inbound

Semantic Scheduling for LLM Inference cites this paper.

Semantic Scheduling for LLM Inference Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T01:09:21.238031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:09:21.238031Z digest=sha256:5426bc6d0ad8f769b19f3f4c22ec8635f80878e39baf0c26d652fe1bf19dcee8

Observation 788f684c-688f-4c76-a0d2-c0d1d8010cba · inbound

MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference cites this paper.

MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:20:42.997865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:20:42.997865Z digest=sha256:79447d69614632b9f5d5ba25d0ef23cede6a033a9e76ca573c8fa472e2626b7a

Observation 93a778ed-c57c-4e7a-ad1c-0000ebe9a475 · inbound

CommVQ: Commutative Vector Quantization for KV Cache Compression cites this paper.

CommVQ: Commutative Vector Quantization for KV Cache Compression Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T18:49:24.114689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:49:24.114689Z digest=sha256:f3d887a65f9e9f0dfbd8e46190e1970ff258caeb1fd50c33055f36ccf55cfbb9

Observation 461a5b00-b6bf-454a-900d-d65b113bee91 · inbound

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models cites this paper.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.768568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.768568Z digest=sha256:8d40fe1335d8b1acd434eeba655e4b75e97948f8705580717cbd26f68b05860b

Observation 34f513a7-45ee-4117-9607-8b2ff82b929e · inbound

SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers cites this paper.

SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:01.272177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:10:01.272177Z digest=sha256:940ceaf0c314977c0b2a49bfa448d457c43598383b0e08357c3a3ae9047730cb

Observation f37ec339-0b23-4a03-bf5b-9a5ccd011380 · inbound

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models cites this paper.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.255004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.255004Z digest=sha256:8a871374bb795b3cf65b20343c9932f9314b1833b4282611a37c07a28d357819

Observation 2932188b-feb0-4505-b425-00706e65caa3 · inbound

DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression cites this paper.

DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:05:37.229366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:05:37.229366Z digest=sha256:e9bb7b71679797bec824c5aaafd2de810c42756f792f76bdad811e7d4010129a

Observation a39092e4-3a0b-44e8-ade1-6d67846053ab · inbound

EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent cites this paper.

EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:38:43.886353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:38:43.886353Z digest=sha256:38bd0dafe82c094a87ab00429a917f253cc84ff5c879b07efc1e34a22e71d303

Observation 94b1b7e0-2608-4c71-b194-af0574f455d9 · inbound

Voice-based AI Agents: Filling the Economic Gaps in Digital Health Delivery cites this paper.

Voice-based AI Agents: Filling the Economic Gaps in Digital Health Delivery Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:40.748817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:40.748817Z digest=sha256:e0755ff6f6f862cb23494348a8ab928e052c32ddb5040a3b31ec62938e7c8f9e

Observation eafe9d99-b0e2-4280-91ad-76196c72b741 · inbound

A Distributed Learned Hash Table cites this paper.

A Distributed Learned Hash Table Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T18:46:59.019444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:46:59.019444Z digest=sha256:4193adde5fb838e63fac55d891c1927830d06f39befb16e3da7c55a9d52a3fd7

Observation 491957c2-5a6c-460d-9d6c-9cc99aeca3ed · inbound

Adaptive KV-Cache Compression without Manually Setting Budget cites this paper.

Adaptive KV-Cache Compression without Manually Setting Budget Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T11:10:56.981798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:10:56.981798Z digest=sha256:7246263a0da9d4f5cc6148eba7542ab6fa292b2473d3fcdc5ed1102274b104dc

Observation d25a45b9-ab71-462e-8b35-d99d6b246da4 · inbound

PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference cites this paper.

PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T10:16:13.096313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:16:13.096313Z digest=sha256:05548f948825eae1e2e25b437e6bd1a008cdd1798e0d2aa3a74739639eb661fd

Observation 64453fe3-4c28-4191-8b70-bc5a7136d4cb · inbound

EvolKV: Evolutionary KV Cache Compression for LLM Inference cites this paper.

EvolKV: Evolutionary KV Cache Compression for LLM Inference Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T20:53:30.115688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T20:53:30.115688Z digest=sha256:f4aaa7146eb3a474f5e253f20c096bccf9d2592314125ebbbf3d6667b507f56c

Observation 2eb1b009-0208-437c-85a3-eaf280a9a3db · inbound

OjaKV: Context-Aware Online Low-Rank KV Cache Compression cites this paper.

OjaKV: Context-Aware Online Low-Rank KV Cache Compression Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:26:24.702813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T13:26:02.980973Z digest=sha256:596eba535f9c210e0ed95533d61840c69c9d62904c8dbed04019274edde24370

Observation b21af5fd-a71a-4e57-a297-453a0d1f8cd8 · inbound

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining cites this paper.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:00:40.765255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:2cd09148306270e87128fb7381f71e55f81433842ffe0f779fa2d38b54d870ad

Observation ef827935-40c0-4108-87f3-cef628ce9116 · inbound

EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction cites this paper.

EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T01:09:36.990898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T01:09:07.983785Z digest=sha256:410cb71deb7f49e146150a52679373bf10b14c3aa48a8c53f3f8a2eaef519567

Observation f78f4f86-0f53-4092-a17d-fcd360ec2729 · inbound

Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward cites this paper.

Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:20:47.770772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T19:59:19.379119Z digest=sha256:6072bee090ad27a27b86638decb3ef24200d6ac90acb9f713a0dc6e821ad06f0

Observation 842f3199-2da6-41b7-be1e-3e9db7153cfb · inbound

TriAttention: Efficient Long Reasoning with Trigonometric KV Compression cites this paper.

TriAttention: Efficient Long Reasoning with Trigonometric KV Compression Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:20:49.203487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T19:55:45.739280Z digest=sha256:3687108dbe35fec8eaa05b523e34e769219ed8ef547441cbd05c1183b89c5fdb

Observation 5da7ee58-3e80-436a-8159-41aac99f4f6c · inbound

How Much Cache Does Reasoning Need? Depth-Cache Tradeoffs in KV-Compressed Transformers cites this paper.

How Much Cache Does Reasoning Need? Depth-Cache Tradeoffs in KV-Compressed Transformers Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:28:39.191396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T05:17:52.344313Z digest=sha256:4f1f0f0e296f9af568ae5d45e17d7a4a943c96e6ddccc379a27d59477a0d642e

Observation 8dce6e59-b5f3-4860-ba32-eb1112d25229 · inbound

Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective cites this paper.

Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-09T01:54:34.032488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-07T16:41:23.234607Z digest=sha256:ae6155650cfd823095f0a042370a1688f3411ab6ff2e6689da127d894432e6b3

Observation 2b58e1d1-177a-4766-91d7-f114ac55f69b · inbound

On the (In-)Security of the Shuffling Defense in the Transformer Secure Inference cites this paper.

On the (In-)Security of the Shuffling Defense in the Transformer Secure Inference Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:36:06.555949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-08T17:24:04.123827Z digest=sha256:d638eb62d2bd1ea51ea1c46140d05cc2f53142fc5fa2cb0ed8ef0470b30e3d1a

Observation 2d31888f-da9b-495c-a3fa-7b03b3942266 · inbound

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost cites this paper.

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 168

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:06:09.876717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-08T10:19:08.451445Z digest=sha256:46b81bf5a0eb677a3f0ed36c7112fb898ee72a9cec5150c044adb3e8584ed9eb

Observation 7804a3b1-2b9e-4312-924b-2cf77aa08c07 · inbound

Agents Should Replace Narrow Predictive AI as the Orchestrator in 6G AI-RAN cites this paper.

Agents Should Replace Narrow Predictive AI as the Orchestrator in 6G AI-RAN Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:17:06.969873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T02:12:21.409833Z digest=sha256:df4acbd1e2e178180fd246446d05d94d40b2c0f0e5cbe71d235989a56e5bd3d5

Observation 444c5df5-e6c5-478e-b03c-62bc57f60cfa · inbound

FlowNar: Scalable Streaming Narration for Long-Form Videos cites this paper.

FlowNar: Scalable Streaming Narration for Long-Form Videos Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:12:34.742615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T19:08:36.655886Z digest=sha256:c2ba74eeed40c4a6689d8615378013fdc8f1456cea2c525001ca512b3fae1ee4

Observation 95bdb966-0c58-4681-837b-d792a80376a8 · inbound

FastTPS: An Optimized Method for LLM Token Phase for AI accelerators cites this paper.

FastTPS: An Optimized Method for LLM Token Phase for AI accelerators Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-14T06:16:09.070414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:16:09.070414Z digest=sha256:4f949e00713e808e76de8e2efaa1463f37a9a73d5902672e9a5bbfed67c7c58e

Observation 6c6b5074-bdbc-4779-8613-04ebeaf226a6 · inbound

CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents cites this paper.

CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:51.051073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:50:51.051073Z digest=sha256:4cd58af2bbd1642d1c8d779318c1c81233866a6027f7452fb2d238ac27ebdfee

Observation 7f9d27ab-a6ec-4cb4-b4e3-a59308de6cc1 · inbound

KVDiagnosis: A Diagnostic Benchmark for KV-Cache Compression in Long-Context Language Models cites this paper.

KVDiagnosis: A Diagnostic Benchmark for KV-Cache Compression in Long-Context Language Models Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T17:52:00.794462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:52:00.794462Z digest=sha256:027bd88b1a7e37d398279ef284d7f3739d0d8bda92eb70918ac5dbc8aaccc806