Pith. sign in

Paper Citation Record · LEDGER

Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 35 inbound Pith citation observations for arXiv:2407.18003.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.18003 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:49:24.114689Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5219d9de-e731-44a0-90d5-02d903ff308e · inbound

LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation cites this paper.

LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-23T18:33:19.467156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-23T18:31:35.391674Z digest=sha256:92cd1c439783dda2d769ce1e24794cfb760cdb073401490ea849d3862b8c93df

Observation d5b3f2d9-d1dd-4bdf-a1ed-cc660c9f4577 · inbound

PoM: Efficient Image and Video Generation with the Polynomial Mixer cites this paper.

PoM: Efficient Image and Video Generation with the Polynomial Mixer Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:47.127086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:47.127086Z digest=sha256:c152f824b461f37d8f1657f419da5600ac643733176d049ae61775c7185cb296

Observation 85aaef17-8091-466c-b1e5-1ef66ac6ee03 · inbound

Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification cites this paper.

Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:00.325854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:00.325854Z digest=sha256:f4b4466f604e0fd91c9b4bbecc28f3dcf619a520b744d0452b4f3674427c97b1

Observation b8c53615-97bd-437a-8068-f3f255f0c4e4 · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:46.742080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:46.742080Z digest=sha256:7e9cf0132215eb7a866e71e2b100dd028b2e7dc7a94f6591a316ff181384e8e8

Observation b92d5962-84f4-4bbe-a72c-02ea9c1698b1 · inbound

Top-Theta Attention: Sparsifying Transformers by Compensated Thresholding cites this paper.

Top-Theta Attention: Sparsifying Transformers by Compensated Thresholding Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T05:29:16.660204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:29:16.660204Z digest=sha256:8b528e80430eb85d7f83fd63301ff328d1445f00ac5ae35e5ebe4b730943fc97

Observation d9d6c010-c1f2-4d1a-ad8e-4885cbdcba68 · inbound

Judge a Book by its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription cites this paper.

Judge a Book by its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:25:19.420768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-23T02:23:22.682357Z digest=sha256:e5ae76274c3a01be7e06b5226cd1d0d4ebbfc63fc820591981d5c264b3eb17c5

Observation a6650d7d-366e-4a5c-96bd-68c48e7e7dc7 · inbound

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models cites this paper.

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 156

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:29:56.798694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-14T01:29:56.480020Z digest=sha256:5fbe867ec7965cc5b91d615baf650bdcf2c6648a6aef713a2fb64879f0c73d9c

Observation a27890fd-b34e-4647-ad95-b4702abfd465 · inbound

Effective and Efficient Schema-aware Information Extraction Using On-Device Large Language Models cites this paper.

Effective and Efficient Schema-aware Information Extraction Using On-Device Large Language Models Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:15.745107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:15.745107Z digest=sha256:c10976a4de3b78edfe1ed205bc09ab32b1be17ea2dff12386bf7876f17264375

Observation 4897a087-dd91-41dd-a307-597c18fc83a6 · inbound

Semantic Scheduling for LLM Inference cites this paper.

Semantic Scheduling for LLM Inference Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T01:09:21.238031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:09:21.238031Z digest=sha256:9de8faa2f9eda161d001ba75818bda9d69ede7d60308f18141623d20246c007c

Observation 788f684c-688f-4c76-a0d2-c0d1d8010cba · inbound

MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference cites this paper.

MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:20:42.997865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:20:42.997865Z digest=sha256:d261021506329db5cc1f5801d66416a85cb8221d2ebe70b1577210f89d835546

Observation 93a778ed-c57c-4e7a-ad1c-0000ebe9a475 · inbound

CommVQ: Commutative Vector Quantization for KV Cache Compression cites this paper.

CommVQ: Commutative Vector Quantization for KV Cache Compression Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T18:49:24.114689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:49:24.114689Z digest=sha256:3c688d8e3aaf8b1ffb7b0a166d10456aaa9764a5266e7574c6437cbacde54a31

Observation 461a5b00-b6bf-454a-900d-d65b113bee91 · inbound

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models cites this paper.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.768568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.768568Z digest=sha256:7671381789698065f55f4ccc4ae3775ba5964aee25c508bc102ac851268e3656

Observation 34f513a7-45ee-4117-9607-8b2ff82b929e · inbound

SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers cites this paper.

SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:01.272177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:10:01.272177Z digest=sha256:91622a98611230343ae5a01caf8989c865a162b41825f3c4454a322d6536d354

Observation f37ec339-0b23-4a03-bf5b-9a5ccd011380 · inbound

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models cites this paper.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.255004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.255004Z digest=sha256:7ed6d1719334a6cd7a766ef6441605d50a62e0ab6095ab3bf7abbccf36fb7c6b

Observation 2932188b-feb0-4505-b425-00706e65caa3 · inbound

DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression cites this paper.

DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:05:37.229366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:05:37.229366Z digest=sha256:6c9964fbd0c13e94222c22bf95ce7329635b87240c146363d20dd7088ba561fb

Observation a39092e4-3a0b-44e8-ade1-6d67846053ab · inbound

EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent cites this paper.

EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:38:43.886353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:38:43.886353Z digest=sha256:1dff52999b0f106f7adfe4486288f23e4c1b233e556a49661db5ad85690882f4

Observation 94b1b7e0-2608-4c71-b194-af0574f455d9 · inbound

Voice-based AI Agents: Filling the Economic Gaps in Digital Health Delivery cites this paper.

Voice-based AI Agents: Filling the Economic Gaps in Digital Health Delivery Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:40.748817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:40.748817Z digest=sha256:e08dc2d26e21c9dca63e9a4af1f6d6dee98af40af2a688f6439c1f08a3df6f0a

Observation eafe9d99-b0e2-4280-91ad-76196c72b741 · inbound

A Distributed Learned Hash Table cites this paper.

A Distributed Learned Hash Table Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T18:46:59.019444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:46:59.019444Z digest=sha256:a84313746193e12083fd32d821bb2e850612d4f2dadf06ae6a6f00f27bed9e54

Observation 491957c2-5a6c-460d-9d6c-9cc99aeca3ed · inbound

Adaptive KV-Cache Compression without Manually Setting Budget cites this paper.

Adaptive KV-Cache Compression without Manually Setting Budget Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T11:10:56.981798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:10:56.981798Z digest=sha256:c0364f7bd00aeaa9394e758405fec961ba1ef267179408b32f775ff124133722

Observation d25a45b9-ab71-462e-8b35-d99d6b246da4 · inbound

PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference cites this paper.

PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T10:16:13.096313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:16:13.096313Z digest=sha256:cd13862fe7eb6b8a563da9f454303698bbcc604ca65c67f1481d9caf723e2945

Observation 64453fe3-4c28-4191-8b70-bc5a7136d4cb · inbound

EvolKV: Evolutionary KV Cache Compression for LLM Inference cites this paper.

EvolKV: Evolutionary KV Cache Compression for LLM Inference Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T20:53:30.115688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T20:53:30.115688Z digest=sha256:0e0f6c1b732ab901a365b0758ca6437aea6030dd95020a60072de7e89607a2b6

Observation 2eb1b009-0208-437c-85a3-eaf280a9a3db · inbound

OjaKV: Context-Aware Online Low-Rank KV Cache Compression cites this paper.

OjaKV: Context-Aware Online Low-Rank KV Cache Compression Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:26:24.702813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T13:26:02.980973Z digest=sha256:c364bf7da054b66b155b779cd66cdb2599d0cc16c6989d42bd2e710d0b4b765e

Observation b21af5fd-a71a-4e57-a297-453a0d1f8cd8 · inbound

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining cites this paper.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:00:40.765255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:329db2a2a76a04ca6ff95a1de4c17414a1736f16ddeda751f80ab8b204b87e93

Observation ef827935-40c0-4108-87f3-cef628ce9116 · inbound

EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction cites this paper.

EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T01:09:36.990898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T01:09:07.983785Z digest=sha256:5e8f3021940563169c44e11410e984aee5890ff713afc57cad1b7ee433c1fcab

Observation f78f4f86-0f53-4092-a17d-fcd360ec2729 · inbound

Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward cites this paper.

Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:20:47.770772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T19:59:19.379119Z digest=sha256:6ed51146a2449defbdedbe31bf5b3cd7242e58630a3d1d751b092a74a68c9579

Observation 842f3199-2da6-41b7-be1e-3e9db7153cfb · inbound

TriAttention: Efficient Long Reasoning with Trigonometric KV Compression cites this paper.

TriAttention: Efficient Long Reasoning with Trigonometric KV Compression Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:20:49.203487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T19:55:45.739280Z digest=sha256:184d1dd05dd919beb0f1ad8efa71a51b248c7c385a73a43d4767f9bf01eaee37

Observation 5da7ee58-3e80-436a-8159-41aac99f4f6c · inbound

How Much Cache Does Reasoning Need? Depth-Cache Tradeoffs in KV-Compressed Transformers cites this paper.

How Much Cache Does Reasoning Need? Depth-Cache Tradeoffs in KV-Compressed Transformers Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:28:39.191396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T05:17:52.344313Z digest=sha256:48f5be722ee2ccf332cf0427f627c48174576176e9a4c46832922fc292aaac37

Observation 8dce6e59-b5f3-4860-ba32-eb1112d25229 · inbound

Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective cites this paper.

Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-09T01:54:34.032488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-07T16:41:23.234607Z digest=sha256:349ce2ac2d4c0324594cb15a202e45981651f9e49baae3e3ba994d35c17c806d

Observation 2b58e1d1-177a-4766-91d7-f114ac55f69b · inbound

On the (In-)Security of the Shuffling Defense in the Transformer Secure Inference cites this paper.

On the (In-)Security of the Shuffling Defense in the Transformer Secure Inference Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:36:06.555949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-08T17:24:04.123827Z digest=sha256:4f401e43c8dde07e8d41fda5cce314f64cfdc724a4119cf50ecd99d3314a374c

Observation 2d31888f-da9b-495c-a3fa-7b03b3942266 · inbound

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost cites this paper.

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 168

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:06:09.876717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-08T10:19:08.451445Z digest=sha256:58b53ea540aea455be11eff49f41ac967a253de5ba49d1c4165db58f2ba936b2

Observation 7804a3b1-2b9e-4312-924b-2cf77aa08c07 · inbound

Agents Should Replace Narrow Predictive AI as the Orchestrator in 6G AI-RAN cites this paper.

Agents Should Replace Narrow Predictive AI as the Orchestrator in 6G AI-RAN Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:17:06.969873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T02:12:21.409833Z digest=sha256:e1b1fe75aa18f962821f901a4d1213c221d40972c92a15b666b68c27ccd8a967

Observation 444c5df5-e6c5-478e-b03c-62bc57f60cfa · inbound

FlowNar: Scalable Streaming Narration for Long-Form Videos cites this paper.

FlowNar: Scalable Streaming Narration for Long-Form Videos Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:12:34.742615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T19:08:36.655886Z digest=sha256:ebeb75aa26500d68e5a683dec45c5143f77cd5c937f7a019a570963e48fd62cd

Observation 95bdb966-0c58-4681-837b-d792a80376a8 · inbound

FastTPS: An Optimized Method for LLM Token Phase for AI accelerators cites this paper.

FastTPS: An Optimized Method for LLM Token Phase for AI accelerators Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-14T06:16:09.070414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:16:09.070414Z digest=sha256:c93495132b42820baf98e87a35c31729a4085cd179543169103758c56b7fd459

Observation 6c6b5074-bdbc-4779-8613-04ebeaf226a6 · inbound

CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents cites this paper.

CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:51.051073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:50:51.051073Z digest=sha256:6e59d67ef27d5ca214c304dc12ed3efeac2f2f457a3dae62a43253ad755eca3a

Observation 7f9d27ab-a6ec-4cb4-b4e3-a59308de6cc1 · inbound

KVDiagnosis: A Diagnostic Benchmark for KV-Cache Compression in Long-Context Language Models cites this paper.

KVDiagnosis: A Diagnostic Benchmark for KV-Cache Compression in Long-Context Language Models Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T17:52:00.794462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:52:00.794462Z digest=sha256:01f7a6b8bcbd167b0ce2f07cf4b73c3d2b75e76e4a07a8e1485fc499f2082589