Pith. sign in

Paper Citation Record · LEDGER

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference

As of 18 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2505.21919.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21919 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:22:45.314876Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 375ff02c-a8a4-408b-b728-6828018802f6 · outbound

This paper cites CHIME: A Cache-Efficient and High-Performance Hybrid Index on Disaggregated Memory,.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference CHIME: A Cache-Efficient and High-Performance Hybrid Index on Disaggregated Memory,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:47.290344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:22:44.081676Z digest=sha256:e15b7f127ae89eb909de328eab4f3c4ba2df35a00ca0f8078ba304dcc575b0f6

Observation 0799148e-4176-41c6-be15-11073ca27ef8 · outbound

This paper cites Sherman: A Write-Optimized Distributed B+Tree Index on Disaggregated Memory,.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference Sherman: A Write-Optimized Distributed B+Tree Index on Disaggregated Memory,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:47.181456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:22:44.174156Z digest=sha256:33237c5f5124d94fa94b8d0938102bd7d47e1800cc71907367246387fa82bc87

Observation f05966c6-ad4a-4800-b5df-a7844b415cc3 · outbound

This paper cites Unlocking Longer Generation with Key-Value Cache Quantization.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference Unlocking Longer Generation with Key-Value Cache Quantization

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:47.010693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:22:44.282406Z digest=sha256:19ea914ea04b7a5b34109a77fbb07f2765f3881c8b91b6e1de8503592fefb409

Observation fc9c2d9e-0206-48e6-a911-c3301b27228b · outbound

This paper cites vLLM vs TensorRT- LLM 12, Automatic Prefix Caching.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference vLLM vs TensorRT- LLM 12, Automatic Prefix Caching

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:46.914847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:22:44.386461Z digest=sha256:9bc07af58e7422e8d08174235e122bcb0cc08385ac32d992c9eff7edaf60a575

Observation 9f28759a-2cf3-4dd8-8106-86b6acf5f8e9 · outbound

This paper cites More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:44.472480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:44.472480Z digest=sha256:67d67464ea0321410550413fd1d1c521726e543acdaa290dd5e93229ca219c2f

Observation ba612c96-afe4-4af9-9450-4a4d25d12738 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:44.554910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:44.554910Z digest=sha256:e63876c22d63b60895028be39bfff0e9b5cbc09bff82d85d966d01482bdf2a56

Observation 6f72440f-ebc4-4f90-b435-07f7a5963d91 · outbound

This paper cites Longformer: The long- document transformer,.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference Longformer: The long- document transformer,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:46.729746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:22:44.706234Z digest=sha256:e349ac519b29a39a9b2cac61f8d7b73b5fb746ec171b5e93f0e2ff25000c5cf8

Observation d6e904c3-9c8b-42a7-95ff-0b9f54fb0e8d · outbound

This paper cites Pie: Pooling CPU Memory for LLM Inference.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference Pie: Pooling CPU Memory for LLM Inference

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:44.784746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:44.784746Z digest=sha256:19c4cbedd1f40b94fb6c547552b852338310480e61bc0e99c084049a9d5dd682

Observation a97db814-edd7-46eb-84ce-9e6adf73d081 · outbound

This paper cites Mooncake: Trading More Storage for Less Computation—A KVCache-centric Architecture for Serving LLM Chatbot,.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference Mooncake: Trading More Storage for Less Computation—A KVCache-centric Architecture for Serving LLM Chatbot,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:46.548558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:22:44.875176Z digest=sha256:5e87c9d3a0e6334b61faf83014f6396947eaf8ad2ba642b759e8614dc9fd6dd5

Observation 0fb4072f-0a8b-49d0-9160-848e7435ae17 · outbound

This paper cites CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving,.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:46.334911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:22:44.988250Z digest=sha256:0e9739c9c62ccda60df678566bbb6b0544076e6ca96c0b546b85f44db1a204dd

Observation 5aeb4ef5-c030-4ff3-a977-6aa2a4c68dd7 · outbound

This paper cites DeepSeek 3FS.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference DeepSeek 3FS

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:46.087572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:22:45.106737Z digest=sha256:1a688787331a8be78087485c87026c5dc5f3bb27c4ec6285463b230351129c88

Observation 75256a89-6565-4e20-b516-4c4f2b4578ab · outbound

This paper cites IMPRESS: An Importance-Informed Multi-Tier Prefix KV Storage System for Large Language Model Inference,.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference IMPRESS: An Importance-Informed Multi-Tier Prefix KV Storage System for Large Language Model Inference,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:45.827026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:22:45.218533Z digest=sha256:634e5814382837a353aacec4f1c722f5967b6bb00abdecb891d683373a9cc483

Observation e12c1413-4397-4f13-abee-b623222209de · outbound

This paper cites Exploring cxl-based kv cache storage for llm serving,.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference Exploring cxl-based kv cache storage for llm serving,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:45.549127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:22:45.314876Z digest=sha256:68096ae74e6409d3554e89dd67299a1eff3bac5a6deec44e26f54e94bdd80a32

Pith citing papers

No inbound Pith citation observations are available.