Pith. sign in

Paper Citation Record · LEDGER

Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 35 inbound Pith citation observations for arXiv:2407.18003.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.18003 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:49:24.114689Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5219d9de-e731-44a0-90d5-02d903ff308e · inbound

LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation cites this paper.

LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-23T18:33:19.467156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T18:31:35.391674Z digest=sha256:34922b5c5203068bfbed7b732e7e802b9cb90b39ebc3bcd048f23f106b669b63

Observation d5b3f2d9-d1dd-4bdf-a1ed-cc660c9f4577 · inbound

PoM: Efficient Image and Video Generation with the Polynomial Mixer cites this paper.

PoM: Efficient Image and Video Generation with the Polynomial Mixer Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:47.127086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:47.127086Z digest=sha256:972c8e7a768e17f3c7eb01b63d67c21e0164e6b9c9c4252dea8667c2093b2080

Observation 85aaef17-8091-466c-b1e5-1ef66ac6ee03 · inbound

Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification cites this paper.

Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:00.325854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:00.325854Z digest=sha256:1c10d8dfd24e374432ddfeb54d163e44e32cff8baa8efd819a39ce879a8d4939

Observation b8c53615-97bd-437a-8068-f3f255f0c4e4 · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:46.742080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:46.742080Z digest=sha256:73b4a1e957072a32b2c648fb42f83c5bc639c558e4e9286b929cda280c14a954

Observation b92d5962-84f4-4bbe-a72c-02ea9c1698b1 · inbound

Top-Theta Attention: Sparsifying Transformers by Compensated Thresholding cites this paper.

Top-Theta Attention: Sparsifying Transformers by Compensated Thresholding Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T05:29:16.660204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:29:16.660204Z digest=sha256:431fdbcb19ca4901f49681b28632e763f8515434fbf8c5396daefa6151509c6e

Observation d9d6c010-c1f2-4d1a-ad8e-4885cbdcba68 · inbound

Judge a Book by its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription cites this paper.

Judge a Book by its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:25:19.420768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-23T02:23:22.682357Z digest=sha256:a1995772b6bc6b095548c24525a4902b15960061e2240800f659146f345a9676

Observation a6650d7d-366e-4a5c-96bd-68c48e7e7dc7 · inbound

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models cites this paper.

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 156

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:29:56.798694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-14T01:29:56.480020Z digest=sha256:ad4e675042c4c060ab91e52e353bcaa97e2249b251548c872b5e72c8c2423c2c

Observation a27890fd-b34e-4647-ad95-b4702abfd465 · inbound

Effective and Efficient Schema-aware Information Extraction Using On-Device Large Language Models cites this paper.

Effective and Efficient Schema-aware Information Extraction Using On-Device Large Language Models Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:15.745107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:15.745107Z digest=sha256:7ebf40ffcd21c4cd3f63134b47939fb4b6f1e87d61c230bc62ad97d22b9669ce

Observation 4897a087-dd91-41dd-a307-597c18fc83a6 · inbound

Semantic Scheduling for LLM Inference cites this paper.

Semantic Scheduling for LLM Inference Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T01:09:21.238031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:09:21.238031Z digest=sha256:06ca7aec352e1a68f41d9d39457a8d521401df2fb63bd98ba27c9a15bac15755

Observation 788f684c-688f-4c76-a0d2-c0d1d8010cba · inbound

MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference cites this paper.

MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:20:42.997865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:20:42.997865Z digest=sha256:b4f5bb4b1fc21fcda561d3732770878b9618aad7ce0f595b9f3ebae919621e88

Observation 93a778ed-c57c-4e7a-ad1c-0000ebe9a475 · inbound

CommVQ: Commutative Vector Quantization for KV Cache Compression cites this paper.

CommVQ: Commutative Vector Quantization for KV Cache Compression Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T18:49:24.114689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:49:24.114689Z digest=sha256:7851fcbf26fd3f3472cc327a3b4843a5d4d378cee600224e3d4d9b54ed0f57f8

Observation 461a5b00-b6bf-454a-900d-d65b113bee91 · inbound

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models cites this paper.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.768568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.768568Z digest=sha256:0e1b9dd1f1411b77b20e794d728b2987115000a81f7960cbde419f61cf4151d6

Observation 34f513a7-45ee-4117-9607-8b2ff82b929e · inbound

SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers cites this paper.

SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:01.272177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:10:01.272177Z digest=sha256:918db91214930e08a195b2d3ea9728bb3604638e0b0ac9cf0556d9d2bee35fc8

Observation f37ec339-0b23-4a03-bf5b-9a5ccd011380 · inbound

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models cites this paper.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.255004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.255004Z digest=sha256:4d43e9670c2b11a780794a453cb516d4db5d1d39f820d42caafc4746120d35c0

Observation 2932188b-feb0-4505-b425-00706e65caa3 · inbound

DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression cites this paper.

DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:05:37.229366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:05:37.229366Z digest=sha256:825ddf10f46b1b803d893f7d894414f84e11c559f0bc45e458da5bbcb96c8b0b

Observation a39092e4-3a0b-44e8-ade1-6d67846053ab · inbound

EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent cites this paper.

EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:38:43.886353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:38:43.886353Z digest=sha256:bb859b072be2c28de3c1830dcc221c998886b93ba1330e7002e13dfbddb67506

Observation 94b1b7e0-2608-4c71-b194-af0574f455d9 · inbound

Voice-based AI Agents: Filling the Economic Gaps in Digital Health Delivery cites this paper.

Voice-based AI Agents: Filling the Economic Gaps in Digital Health Delivery Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:40.748817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:40.748817Z digest=sha256:af5a1df87bb66c6c2000fb8495c5c674119a3fe04a5d46e77e96dbf99c6c6f85

Observation eafe9d99-b0e2-4280-91ad-76196c72b741 · inbound

A Distributed Learned Hash Table cites this paper.

A Distributed Learned Hash Table Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T18:46:59.019444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:46:59.019444Z digest=sha256:2d5e38fb840234e26181ab371b58add3dd618c90ce55ffeae264e59607465f10

Observation 491957c2-5a6c-460d-9d6c-9cc99aeca3ed · inbound

Adaptive KV-Cache Compression without Manually Setting Budget cites this paper.

Adaptive KV-Cache Compression without Manually Setting Budget Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T11:10:56.981798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:10:56.981798Z digest=sha256:c365e827c8dc8ab7376fc82827de7c206407aa515bbcc80660e0ffb8c3d06c2e

Observation d25a45b9-ab71-462e-8b35-d99d6b246da4 · inbound

PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference cites this paper.

PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T10:16:13.096313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:16:13.096313Z digest=sha256:a556db702a1416b532f89c2ddc67f6d3fe3e3ddd916282713fe39491c70d4910

Observation 64453fe3-4c28-4191-8b70-bc5a7136d4cb · inbound

EvolKV: Evolutionary KV Cache Compression for LLM Inference cites this paper.

EvolKV: Evolutionary KV Cache Compression for LLM Inference Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T20:53:30.115688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T20:53:30.115688Z digest=sha256:cbc641b0b0f39f3ef122af46b8d4cafb8bba856a44773260349d1c3567b66498

Observation 2eb1b009-0208-437c-85a3-eaf280a9a3db · inbound

OjaKV: Context-Aware Online Low-Rank KV Cache Compression cites this paper.

OjaKV: Context-Aware Online Low-Rank KV Cache Compression Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:26:24.702813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T13:26:02.980973Z digest=sha256:8823b2964904e6e61f66449b1416496831d65ff72d42fc2a4804246a8ac3f0a0

Observation b21af5fd-a71a-4e57-a297-453a0d1f8cd8 · inbound

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining cites this paper.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:00:40.765255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:ca48f0185c5e4422ad14b30a10c063b83e63fd63e124adc276960487436b3f56

Observation ef827935-40c0-4108-87f3-cef628ce9116 · inbound

EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction cites this paper.

EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T01:09:36.990898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T01:09:07.983785Z digest=sha256:eccb3d35ec392d7d16a0318624f11163c0f00c53f885f5be6724ee33bacd55e1

Observation f78f4f86-0f53-4092-a17d-fcd360ec2729 · inbound

Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward cites this paper.

Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:20:47.770772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T19:59:19.379119Z digest=sha256:0b87f803f6144ca6ef94f51847e99e6f03f4ce9cd3ab17735f94b0e75c40b2b3

Observation 842f3199-2da6-41b7-be1e-3e9db7153cfb · inbound

TriAttention: Efficient Long Reasoning with Trigonometric KV Compression cites this paper.

TriAttention: Efficient Long Reasoning with Trigonometric KV Compression Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:20:49.203487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T19:55:45.739280Z digest=sha256:e0d2e5335cb3a1500210cc398c77fe4d4faad4aceef7700da3cea2c09bf1bb7b

Observation 5da7ee58-3e80-436a-8159-41aac99f4f6c · inbound

How Much Cache Does Reasoning Need? Depth-Cache Tradeoffs in KV-Compressed Transformers cites this paper.

How Much Cache Does Reasoning Need? Depth-Cache Tradeoffs in KV-Compressed Transformers Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:28:39.191396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T05:17:52.344313Z digest=sha256:64c4aa687593b973c8c7678951227138f1a1d605f80e6419c2344a720d2cabf3

Observation 8dce6e59-b5f3-4860-ba32-eb1112d25229 · inbound

Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective cites this paper.

Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-09T01:54:34.032488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-07T16:41:23.234607Z digest=sha256:1ca47e8904e3f91ab5ca89f3361e2c307b150d67de689f90cba5429a6ab932c6

Observation 2b58e1d1-177a-4766-91d7-f114ac55f69b · inbound

On the (In-)Security of the Shuffling Defense in the Transformer Secure Inference cites this paper.

On the (In-)Security of the Shuffling Defense in the Transformer Secure Inference Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:36:06.555949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-08T17:24:04.123827Z digest=sha256:06ef5b20aac51fd8ef35ef382b05e58af22411761275a97956bb3170ffd140bc

Observation 2d31888f-da9b-495c-a3fa-7b03b3942266 · inbound

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost cites this paper.

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 168

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:06:09.876717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-08T10:19:08.451445Z digest=sha256:7360decfd8cd34b5b5514550de6e460eae68b46c31fd71a6e313e8b5bb81e7da

Observation 7804a3b1-2b9e-4312-924b-2cf77aa08c07 · inbound

Agents Should Replace Narrow Predictive AI as the Orchestrator in 6G AI-RAN cites this paper.

Agents Should Replace Narrow Predictive AI as the Orchestrator in 6G AI-RAN Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:17:06.969873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T02:12:21.409833Z digest=sha256:a3675e494d2614e8a6007b8db26cf35b1a24d81c1694308ed0356cfb3bc968b1

Observation 444c5df5-e6c5-478e-b03c-62bc57f60cfa · inbound

FlowNar: Scalable Streaming Narration for Long-Form Videos cites this paper.

FlowNar: Scalable Streaming Narration for Long-Form Videos Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:12:34.742615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T19:08:36.655886Z digest=sha256:20889342e64acfad862c71da299c5f6211920b36ed388e6935601b4f10f470a1

Observation 95bdb966-0c58-4681-837b-d792a80376a8 · inbound

FastTPS: An Optimized Method for LLM Token Phase for AI accelerators cites this paper.

FastTPS: An Optimized Method for LLM Token Phase for AI accelerators Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-14T06:16:09.070414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:16:09.070414Z digest=sha256:01f413249ebf50878d2314eb76777dc32e211172c1db2e839cd6a7648ae2133b

Observation 6c6b5074-bdbc-4779-8613-04ebeaf226a6 · inbound

CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents cites this paper.

CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:51.051073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:50:51.051073Z digest=sha256:2b718635ba4eea98a65a0af4a2a90446bfb102bd59110a77167b1125b6fc68e6

Observation 7f9d27ab-a6ec-4cb4-b4e3-a59308de6cc1 · inbound

KVDiagnosis: A Diagnostic Benchmark for KV-Cache Compression in Long-Context Language Models cites this paper.

KVDiagnosis: A Diagnostic Benchmark for KV-Cache Compression in Long-Context Language Models Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T17:52:00.794462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:52:00.794462Z digest=sha256:8768aa40771a69cfb2c4955a763d1b42bc06b1e33bfb4aaa315a39adbc5849d8