Pith. sign in

Paper Citation Record · LEDGER

KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation

As of 14 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 3 inbound Pith citation observations for arXiv:2411.17089.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17089 v2

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:40:51.752812Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:38:49.204953Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T16:31:10.505851Z

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 61015e1e-527a-4049-971a-e38927b0e88f · outbound

This paper cites FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines.

KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T12:40:51.717925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:40:51.717925Z digest=sha256:96829b3cfcffa6b18a1fb6fcebef0ddd48ce48ac58b4cfcaa1ed7979e9fec0ab

Observation ba28753b-5599-4c3b-a9e7-a479bece9677 · outbound

This paper cites NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference.

KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T12:40:51.722346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:40:51.722346Z digest=sha256:c19321d5a75e5729a7966dabc3834da8620e6b2d93c7de5af1ef8bbcfb4fa0bf

Observation 2197fee8-9bca-4351-b9e9-3a7d3ce257eb · outbound

This paper cites an unresolved cited work.

KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:40:51.834636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T12:40:51.752812Z digest=sha256:edff817ad2cae98fb03931146a77e918b65cb281d1a31eb51f950478113e45ff

Observation a36735cf-9553-4a9f-a276-929e4682d7c3 · outbound

This paper cites GPT-4 Technical Report.

KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation GPT-4 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T12:40:51.729348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:40:51.729348Z digest=sha256:f984e9a176ecae0ad29216a9cd458c5aed54db11e4b9bedb72111a6279cf69ba

Observation a5bc4129-4258-40c9-9011-9a6bd8b5d65f · outbound

This paper cites In Proceedings of the 2024 International Conference on Parallel Architectures and Compila- tion Techniques, pages 233–245.

KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation In Proceedings of the 2024 International Conference on Parallel Architectures and Compila- tion Techniques, pages 233–245

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:40:51.873477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T12:40:51.732486Z digest=sha256:72204f6b9d6608be8bc469bcb1e9baca410712416a860737f0b41ee3e98d966f

Observation 6e0145da-ef53-45aa-bb07-2b890420585c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation LLaMA: Open and Efficient Foundation Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T12:40:51.735584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:40:51.735584Z digest=sha256:629468b26785bd0f38609fdd0905c911095aae3290a553030f9cf1d88a49bed9

Observation 321d3bc7-532e-4b41-a0c7-ebaa4b1b51bf · outbound

This paper cites In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations.

KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:40:51.864328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T12:40:51.739789Z digest=sha256:80e47cfd0210dae6edc7b8b0e4f6fa6979102885fa30f7c8428ab9b9a7e44ba3

Observation 2f7dcad6-803d-4be6-968d-67200567f8ce · outbound

This paper cites In Thirty-seventh Conference on Neural In- formation Processing Systems.

KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation In Thirty-seventh Conference on Neural In- formation Processing Systems

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:40:51.854848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T12:40:51.746157Z digest=sha256:510573d5af13517ebd8624a2c55894d8ef01145e2c8e7c7046c28d2bcbf58f35

Observation d9a9397c-483b-4153-a1f6-3954de36c5a1 · outbound

This paper cites In Findings of the Association for Computational Linguistics: NAACL 2024, pages 2765–2781.

KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation In Findings of the Association for Computational Linguistics: NAACL 2024, pages 2765–2781

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:40:51.845682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T12:40:51.749527Z digest=sha256:ccd0648941b0ea780caa4d1ec3cacdc4b18971581a800767383cb60d3838811c

Observation 1b7bebca-5983-431a-a6c8-dad9adfda00a · outbound

This paper cites In Ad- vances in Neural Information Processing Systems , volume 33, pages 1877–1901.

KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation In Ad- vances in Neural Information Processing Systems , volume 33, pages 1877–1901

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:40:51.883105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T12:40:51.710750Z digest=sha256:122a4f950ec662932e710e462136670d726812993e9a06ef26667f2ba8fdf3d4

Observation dfdf8128-9bb0-437c-869a-c5bb40be82e8 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation OPT: Open Pre-trained Transformer Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-12T12:40:51.742791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:40:51.742791Z digest=sha256:29cd0c0a04df0887ac2527313540fb1a9da5a26c396a2309816eff183abae101

Observation d80a4b07-3dc5-466e-9bb0-dfb5d39a2cee · outbound

This paper cites CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society.

KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T12:40:51.725849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:40:51.725849Z digest=sha256:d4ef8cee551f2f0e73db3e6684b88b5112ab434e03d001a1ba82a37c523a6ce6

Observation d553e17d-7660-49ab-917c-2d8cf9ed06df · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation Gemini: A Family of Highly Capable Multimodal Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T12:40:51.714413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:40:51.714413Z digest=sha256:c58cf2667ca011cd6dccc22b87079809f3ef741d6c32b6d19e0267a6a0f09d99

Pith citing papers

Observation f13a4da3-a693-466f-a635-19f3ee40eb49 · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation

Reference 244

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:49.204953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:49.204953Z digest=sha256:23b20b4ea44f8a51f5daf3f5c400eb0d35d6e689497c284e55adfaa219a3f54f

Observation 8dd36822-23d0-4a69-a797-88e0c5ff4a97 · inbound

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding cites this paper.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:37.745707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:37.745707Z digest=sha256:e6fc9ad107004abb3073f613879958d3dc2b3b2bd455216638f37f3f597c87f9

Observation d9eba12a-550a-4c38-b750-529da45dadfd · inbound

Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving cites this paper.

Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:31:10.512778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-09T16:19:33.613685Z digest=sha256:2a47f7ce658a7c478f7d33ef3bf66b363f31313b080e4bc9d5d74f7821859b01