Pith. sign in

Paper Citation Record · LEDGER

PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2405.12532.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.12532 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:37:07.674267Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T01:36:44.048156Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1d273cea-bb4c-45d8-9942-948c028785f0 · inbound

PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling cites this paper.

PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:58:29.224557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T09:58:29.057357Z digest=sha256:71ca5bb7ea338b89c312f577a8e67811683788f0166e5736294019fc28ff7e08

Observation b3e5f2d2-b669-4ca8-81eb-47de6ea97466 · inbound

LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation cites this paper.

LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-23T18:33:19.489905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T18:31:35.391674Z digest=sha256:f635488598c11df8d940ad2011da466fd17dd3616c8f42870c0c8b356f943771

Observation 7ced9945-c15e-4e04-939c-711b8dbb8902 · inbound

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache cites this paper.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.674267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.674267Z digest=sha256:cc70a29b2224e68c7d13bb9c72fd951d2179a1874db6a9d08517b055f3a8463d

Observation 06d812ad-d052-4c96-b8a5-43a226a96667 · inbound

ZigZagkv: Dynamic KV Cache Compression for Long-context Modeling based on Layer Uncertainty cites this paper.

ZigZagkv: Dynamic KV Cache Compression for Long-context Modeling based on Layer Uncertainty PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T17:24:49.256708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:24:49.256708Z digest=sha256:e449874196b5b7c2c60e4b64a550dcdc9b6b9060805d4c916b79ac54377044c5

Observation dcae64ab-e48c-4d7b-98f6-450c255989ef · inbound

SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator cites this paper.

SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T14:25:06.072675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:25:06.072675Z digest=sha256:766db1005421fd2ce33125b395eee1e0bcf321d872e85a7731c7a285b884d9c6

Observation f2ec4fe1-747a-4a2b-8a12-5f7720a90062 · inbound

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression cites this paper.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.806959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.806959Z digest=sha256:b92160669d1503d95f356dd43588a5bcee80e091148f319d0fb6f44ebef56e3f

Observation f328f19a-bac3-4371-a5a2-5be7ba46d374 · inbound

TreeKV: Smooth Key-Value Cache Compression with Tree Structures cites this paper.

TreeKV: Smooth Key-Value Cache Compression with Tree Structures PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T21:25:41.251024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:25:41.251024Z digest=sha256:556fa3705a888d17dd6660b40723592e71eb9eef1b71647b7baae1ad0451ea4b

Observation 46fdbb1a-83a3-4239-8abf-ed34e63c0914 · inbound

Task-KV: Task-aware KV Cache Optimization via Semantic Differentiation of Attention Heads cites this paper.

Task-KV: Task-aware KV Cache Optimization via Semantic Differentiation of Attention Heads PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T14:41:05.004324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:41:05.004324Z digest=sha256:683710d28458e386941b69744c7f76898a0b7f06dc36ccd7accf6c157c8dc929

Observation 35e38757-05ef-4846-815e-4eba4dbf94ba · inbound

Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression cites this paper.

Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:17:31.141533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T04:15:36.906263Z digest=sha256:ecc257923d6a21258d4c343d158735ed0971428b2f857024cfc0311fff49c9ad

Observation cd6b8efe-e3fa-4474-a515-8ec7e9cc7aae · inbound

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference cites this paper.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T04:32:01.408211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:32:01.408211Z digest=sha256:ad4429a23e29675d6f352f21744a0a07ffb565026bb91a51d5173610b49fe772

Observation 2239970b-c879-4ee0-a22c-4c9f7a1fa120 · inbound

Efficient-vDiT: Efficient Video Diffusion Transformers With Attention Tile cites this paper.

Efficient-vDiT: Efficient Video Diffusion Transformers With Attention Tile PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T16:36:00.827553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:36:00.827553Z digest=sha256:79cd24c4f688f7437b958f3b01d8a8cb3d0aa1c671354f783430c6a74fe89cc9

Observation d88e456d-ea6b-40d7-aaea-51187e3df623 · inbound

Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents cites this paper.

Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-08T14:13:38.413409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:13:38.413409Z digest=sha256:ab9c64766023ef4ec58de2e83f844332341b4bea9d923ee39b493eb1f0c8adef

Observation 827a1ae8-2935-4ac9-96d7-895870c2b1be · inbound

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression cites this paper.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:19.360068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:19.360068Z digest=sha256:410f98b97c406f511a1e26d2130d10ba58c5f16c72c357f1755a4f764272d511

Observation 923033ba-0edb-4294-9ebd-9e6604817a56 · inbound

AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity cites this paper.

AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:55.794210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:55.794210Z digest=sha256:0c0af32f8dce7a8e1f8a0dbdd171d6ec963020d10176fa21156cb3e4e2ab9e5b

Observation 58f35e4a-a465-400d-b3a3-ceb2549dc751 · inbound

MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference cites this paper.

MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T10:20:44.590087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:20:44.590087Z digest=sha256:e28aa27ecd527706b2d720c221a16985d849280a1b3ccb3af5e7c9577642761c

Observation 72268e6b-f74b-4d8c-bf88-5a2f34360f62 · inbound

Breaking the Boundaries of Long-Context LLM Inference: Adaptive KV Management on a Single Commodity GPU cites this paper.

Breaking the Boundaries of Long-Context LLM Inference: Adaptive KV Management on a Single Commodity GPU PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T23:00:11.211770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:00:11.211770Z digest=sha256:8593f56d43f86557827189b40dbb0c262bd1e23a8c4a00546d868aa3ae4b90f4

Observation 5c385049-5d20-4389-bb81-fc98a8695fd0 · inbound

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs cites this paper.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.863938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.863938Z digest=sha256:17f088c6c5c4480cc5d638603954a083d8f07aac620c9d7d101d803cbe0c9eea

Observation 8c455c61-e622-48db-b715-3f0ab16003a7 · inbound

Think Clearly: Improving Reasoning via Redundant Token Pruning cites this paper.

Think Clearly: Improving Reasoning via Redundant Token Pruning PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:18.752863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:18.752863Z digest=sha256:8025443abc6c1245e4d67dc4d4962079f243fa749faebe750cee2f7665398791

Observation 08f65a1c-9488-4181-974d-2413965dcaed · inbound

KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding cites this paper.

KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T17:20:28.474622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:20:28.474622Z digest=sha256:44f40f411c84639171127f347a4b6ccad3e8b97e0524a28db010190fffa0bdb3

Observation 1280fc6c-c0e7-48ed-90fa-0b80692b91d2 · inbound

CaliDrop: KV Cache Compression with Calibration cites this paper.

CaliDrop: KV Cache Compression with Calibration PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.807556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.807556Z digest=sha256:d5e63bd9d8223d618844dab3695e2f87a92189dd733e01e183016bfb68b46aeb

Observation 7a133e17-f04e-46f2-a209-9d66bbdc268e · inbound

GraphKV: Breaking the Static Selection Paradigm with Graph-Based KV Cache Eviction cites this paper.

GraphKV: Breaking the Static Selection Paradigm with Graph-Based KV Cache Eviction PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T13:44:18.879414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:44:18.879414Z digest=sha256:7231a45c1576e44bbc404016583ebf4827cb756668767a802557d442fa9ee611

Observation 9ac1f9ac-caa9-43a1-85fd-33854434a692 · inbound

EvolKV: Evolutionary KV Cache Compression for LLM Inference cites this paper.

EvolKV: Evolutionary KV Cache Compression for LLM Inference PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T20:53:30.523155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T20:53:30.523155Z digest=sha256:33e1b23604a3fda65d5375f589aceb4c39097cba4879c8856bd6736ce6b5d207

Observation ec0782cc-b44f-435c-b9b1-4cb23cfdf3bd · inbound

ClusterFusion++: Expanding Cluster-Level Fusion to Full Transformer-Block Decoding cites this paper.

ClusterFusion++: Expanding Cluster-Level Fusion to Full Transformer-Block Decoding PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:26:17.341241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-08T05:34:08.596351Z digest=sha256:e6c0d9032d8d85cde7e26ed98e8f73c4b042e823bd67098a02364d531f0ec5d8

Observation 7adfc9b3-a699-43a6-8f06-1303ddf5a6ee · inbound

Sparse Prefix Caching for Hybrid and Recurrent LLM Serving cites this paper.

Sparse Prefix Caching for Hybrid and Recurrent LLM Serving PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:53:04.417631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T08:49:24.880528Z digest=sha256:2d1f3d5bb47655d3887e666c04f1ffacc0d98959d70e0976ff2ded8ac836a4b6

Observation 0a7177e5-8175-4beb-8971-d857ab6ec030 · inbound

Reformulating KV Cache Eviction Problem for Long-Context LLM Inference cites this paper.

Reformulating KV Cache Eviction Problem for Long-Context LLM Inference PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:10:53.563323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T02:37:52.545943Z digest=sha256:8b192242c98e6ce4a46f59900a3427eb5158058293dbfb79acf2bd172afd0eb1

Observation e3b5b830-d887-42bb-9dc9-50d00a1e4f32 · inbound

ReST-KV: Robust KV Cache Eviction with Layer-wise Output Reconstruction and Spatial-Temporal Smoothing cites this paper.

ReST-KV: Robust KV Cache Eviction with Layer-wise Output Reconstruction and Spatial-Temporal Smoothing PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:26:30.415977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T02:52:33.076123Z digest=sha256:23f1535546c4b78066f556efd6661c65f619c9bf5a382bd0352f98396259d7fc

Observation 81b036f3-9b64-4a9b-93be-9e841b39eecb · inbound

Protection Is (Nearly) All You Need: Structural Protection Dominates Scoring in Globally Capped KV Eviction cites this paper.

Protection Is (Nearly) All You Need: Structural Protection Dominates Scoring in Globally Capped KV Eviction PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:15.968577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T12:13:14.295659Z digest=sha256:1ae6ad6f6d0a2422a4de9f8f8149c29114fbc63a5519546ecb973bb66967de4b

Observation 102101e7-0036-4788-aa5f-aeec894e557d · inbound

ProactiveLLM: Learning Active Interaction for Streaming Large Language Models cites this paper.

ProactiveLLM: Learning Active Interaction for Streaming Large Language Models PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 112

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:12:34.806502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T19:07:48.425181Z digest=sha256:3e58c843f250619150dfd90971353b6056ab2753189149864bb3c2270ef8a572

Observation f4f64744-e14b-46f3-9e34-bc9bd6626770 · inbound

From Rigid to Dynamic: Entropy-Guided Adaptive Inference for Long-Context LLMs cites this paper.

From Rigid to Dynamic: Entropy-Guided Adaptive Inference for Long-Context LLMs PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:57:30.010828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T16:57:02.709967Z digest=sha256:16f7d26a2a204f39c4d8ca481ed9923e8b68b55983fe2fced5507451183e54f1

Observation ec731856-38ec-41c6-b8ce-74bd7eec91d6 · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 135

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T01:36:44.049434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:dd5bce27576f71e789390b15312457d8fb43a42cd1b55d52b39a22f73f2d49fb