Pith. sign in

Paper Citation Record · LEDGER

RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2404.12457.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.12457 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T05:40:21.499545Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

5
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3fbfd308-5d65-4a28-af0f-9a707ad817b9 · inbound

Retrieval-Augmented Generation for Natural Language Processing: A Survey cites this paper.

Retrieval-Augmented Generation for Natural Language Processing: A Survey RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-23T23:08:35.645490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T23:06:41.081461Z digest=sha256:77662cd0d653c660ec459831fdce45b7dda9df4dfa2d26655da8645d45d55e3a

Observation 6536fa6a-dff0-45bb-85db-6cfc04398c6b · inbound

BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching cites this paper.

BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-23T16:58:11.921159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T16:57:46.645061Z digest=sha256:240a1b6d7a9d9e2f39b732fff08daf64a50025037b960a11eea98178eb97a31e

Observation fc23ad34-ebe3-4bd3-b794-beb7a386bbcd · inbound

Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation cites this paper.

Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T05:40:21.499545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:40:21.499545Z digest=sha256:d1edd100de96fd24378713ce7192ab698031f4278a268a397cd1f9c993f88681

Observation 1b79fb77-a836-4cdb-bc81-674d147e7c19 · inbound

From Human Memory to AI Memory: A Survey on Memory Mechanisms in the Era of LLMs cites this paper.

From Human Memory to AI Memory: A Survey on Memory Mechanisms in the Era of LLMs RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

Reference 131

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:05:09.865438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T11:05:09.588491Z digest=sha256:8bf6fa763e29b39d397e4677939d8875f0d821d4cabbeefbef9371871ad26ab9

Observation c899251b-df6c-4db9-9098-eb0f070ef054 · inbound

A Survey of LLM $\times$ DATA cites this paper.

A Survey of LLM $\times$ DATA RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

Reference 197

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:13.236325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:13.236325Z digest=sha256:56af9169ea866cf97b8c730cc19563bfb6dec687d807deb4a9b9c8c7247f88bb

Observation 6d08e020-7250-440d-834b-0e527e0b44c3 · inbound

Retrieval-Augmented Generation: A Comprehensive Survey of Architectures, Enhancements, and Robustness Frontiers cites this paper.

Retrieval-Augmented Generation: A Comprehensive Survey of Architectures, Enhancements, and Robustness Frontiers RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:03:10.832053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:03:10.832053Z digest=sha256:242052e4d5478666234511c60b3cbe5e8ad8f871be623cf86248598ffbf54e81

Observation 69bb8b5b-c3e9-48e6-8d79-e03804a72126 · inbound

WebANNS: Fast and Efficient Approximate Nearest Neighbor Search in Web Browsers cites this paper.

WebANNS: Fast and Efficient Approximate Nearest Neighbor Search in Web Browsers RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:42.977310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:42.977310Z digest=sha256:330cecd8c9a012660ceed23d114403391db2fc163078cf958c97a2d8d658eaec

Observation a35af8c7-2b84-4357-891c-e51626203c4d · inbound

A Survey on Proactive Defense Strategies Against Misinformation in Large Language Models cites this paper.

A Survey on Proactive Defense Strategies Against Misinformation in Large Language Models RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:03:46.686010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:03:46.686010Z digest=sha256:7b0b4a6222ca2bbf077b96ee4a6cb791abe34733b971b061ee48da42eef06bae

Observation 74408f15-ecf0-4635-b034-f7e1da15f41e · inbound

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows cites this paper.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.296114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.296114Z digest=sha256:96c6aad5c9f1f716c30357bdcbcbf8d2fcd21f514d0575378da60a17859720f4

Observation 3030bba9-51d7-4025-8545-8f6d73185e4d · inbound

HedraRAG: Coordinating LLM Generation and Database Retrieval in Heterogeneous RAG Serving cites this paper.

HedraRAG: Coordinating LLM Generation and Database Retrieval in Heterogeneous RAG Serving RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:32.840194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:08:32.840194Z digest=sha256:d756c37967f2abd447817ae86bd1871842f584fbab7bdfc73bb4723354591ef1

Observation 2c7c98ed-3e3d-4601-a701-51d152537171 · inbound

Beyond Isolated Capabilities: Bridging Long CoT Reasoning and Long-Context Understanding cites this paper.

Beyond Isolated Capabilities: Bridging Long CoT Reasoning and Long-Context Understanding RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:50:51.161173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:50:51.161173Z digest=sha256:1f7acba529698b70272d101e0f7da99c00edfffcde7b2f90f46b885bb7f2ae7b

Observation e2801588-2fe0-4b20-8828-917c861eb531 · inbound

MultiFluxAI Enhancing Platform Engineering with Advanced Agent-Orchestrated Retrieval Systems cites this paper.

MultiFluxAI Enhancing Platform Engineering with Advanced Agent-Orchestrated Retrieval Systems RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T14:28:37.837782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:28:37.837782Z digest=sha256:24aa176c8cca6a1696e308c331ba60a0ca64596955463c4ad788c525b2cfdac0

Observation 5612ddf2-f1c9-470b-af6a-634537b6e382 · inbound

ConceptBot: Enhancing Robot's Autonomy through Task Decomposition with Large Language Models and Knowledge Graph cites this paper.

ConceptBot: Enhancing Robot's Autonomy through Task Decomposition with Large Language Models and Knowledge Graph RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T13:32:47.824170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:32:47.824170Z digest=sha256:9c185059fde19a25f77b161c08868563786e895f542a87988585ebabf4a00293

Observation e2486d64-1e9b-49aa-915d-a7375c11376e · inbound

CacheClip: Accelerating RAG with Effective KV Cache Reuse cites this paper.

CacheClip: Accelerating RAG with Effective KV Cache Reuse RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:41:33.782933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T12:36:41.630618Z digest=sha256:9557a260d3ab0786b3712140eed1d04d8b4a24bb1c976b99fcf81c9a00e42cfa

Observation 961964b4-2e43-44ad-a432-487a63a32c31 · inbound

AtlasKV: Augmenting LLMs with Billion-Scale Knowledge Graphs in 20GB VRAM cites this paper.

AtlasKV: Augmenting LLMs with Billion-Scale Knowledge Graphs in 20GB VRAM RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-18T06:02:25.443013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T06:00:59.224521Z digest=sha256:8f08bd7ccadd172f40a71f1cb29608ff403176784aaecab17a477c290d38f0be

Observation aaef1185-2944-46aa-a466-43a06277731e · inbound

Efficient Remote KV Cache Reuse with GPU-native Video Codec cites this paper.

Efficient Remote KV Cache Reuse with GPU-native Video Codec RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:22:22.800333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:9c852d6444d5824cb374b38fe580f7c9cb42f39027e59b9358010b1022e35f67

Observation 953aa14f-794d-40c9-b223-598a4216eb2d · inbound

Mobility-Aware Cache Framework for Scalable LLM-Based Human Mobility Simulation cites this paper.

Mobility-Aware Cache Framework for Scalable LLM-Based Human Mobility Simulation RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T22:50:29.864637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:50:29.864637Z digest=sha256:7e12e3cf51adec63ec75d2e6679a26a18337e3da69ca5807ab872a26db8959ff

Observation 10c677c8-8ccb-49b1-9b1a-9c79a39741a7 · inbound

Evaluating Temporal Semantic Caching and Workflow Optimization in Agentic Plan-Execute Pipelines cites this paper.

Evaluating Temporal Semantic Caching and Workflow Optimization in Agentic Plan-Execute Pipelines RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:23:58.622882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T05:20:39.385107Z digest=sha256:d83add98c4f33559ec826dd9bfbe856bd87d9364cd1007cbd8b12bdc275330fb

Observation eae7e26a-3f74-4197-90fb-2797f408ca5c · inbound

Grounded Cache Routing for Retrieval-Augmented Generation: When Is It Safe to Reuse an Answer? cites this paper.

Grounded Cache Routing for Retrieval-Augmented Generation: When Is It Safe to Reuse an Answer? RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:23:45.128328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T17:16:48.588593Z digest=sha256:12a0d5c457a55202cda7ecdbffb79532a6fe29247d585d8720fb3d9fa8cabf45

Observation 60f51884-5529-4661-aea2-4a6d2349146b · inbound

LazyAttention: Efficient Retrieval-Augmented Generation with Deferred Positional Encoding cites this paper.

LazyAttention: Efficient Retrieval-Augmented Generation with Deferred Positional Encoding RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-28T07:11:45.594415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T07:04:28.959310Z digest=sha256:9488d194b5e8b0931e8f44754f767046fded55e03b9197c863f654b2ef6f205d

Observation 26bc2e9c-f250-41f2-862d-b4f2ed2a3a48 · inbound

SIFT: Selective-Index For Fast Compute of RAG Prefill by Exploiting Attention Invariance cites this paper.

SIFT: Selective-Index For Fast Compute of RAG Prefill by Exploiting Attention Invariance RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:37:31.248569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T16:24:31.109508Z digest=sha256:acee18fa849739c1adb1207c4616c8b56deb2d64e2c9b86ac785bfb602292c8e

Observation a49c5a09-3819-4e03-b42a-79d7426aafcb · inbound

Functional Cache Grafting: Robust and Rapid Code-Policy Synthesis for Embodied Agents cites this paper.

Functional Cache Grafting: Robust and Rapid Code-Policy Synthesis for Embodied Agents RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-27T05:30:35.773180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T05:26:48.431739Z digest=sha256:beede186e50add64e48a7144b715ca5b41169492f68e08f15ca5d7b051b092b0

Observation 10a75f28-7150-4e66-9bb2-2e980e055c52 · inbound

Functional Cache Grafting: Robust and Rapid Code-Policy Synthesis for Embodied Agents cites this paper.

Functional Cache Grafting: Robust and Rapid Code-Policy Synthesis for Embodied Agents RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

Reference 1977

Resolution
unresolved
no resolver link, observed 2026-08-02T11:44:03.602575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:44:03.602575Z digest=sha256:222883ca64805b98c43d914a991ea870f93d0e81a9adce684c7574fc900a57bd

Observation 0b498f68-9bdf-4e7d-9612-343f6446a0da · inbound

Models Take Notes at Prefill: KV Cache Can Be Editable and Composable cites this paper.

Models Take Notes at Prefill: KV Cache Can Be Editable and Composable RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:58:46.974184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T03:37:27.513196Z digest=sha256:554bb8d7d7cadc01a64553f41108556a0af9b85b8522c5af76f2b11e52f705b9

Observation e82ca0f9-5e10-4282-8839-84d60ebae099 · inbound

KAP: Bridging the Knowledge Selection-Runtime Consumption Gap in LLM Systems cites this paper.

KAP: Bridging the Knowledge Selection-Runtime Consumption Gap in LLM Systems RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-31T19:52:41.311484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T19:52:41.311484Z digest=sha256:d755e01e2e360b0610dd319cc8cabefdc243b3eda4511d138997dda959fb1a2d

Observation 3350c761-cafa-4c99-89dd-681053936c9b · inbound

FinCacheServe: Dependency-Consistent Answer Reuse for Cost-Efficient RAG Serving over Mutable Enterprise Documents cites this paper.

FinCacheServe: Dependency-Consistent Answer Reuse for Cost-Efficient RAG Serving over Mutable Enterprise Documents RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T06:40:18.439382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:40:18.439382Z digest=sha256:651ceefd74a07dc045d6778c7f098a702f52a4174beb0a5304d269729962feee

Observation 14ff5225-b6c4-4a1c-948b-b7bfa0922871 · inbound

Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory cites this paper.

Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-31T12:56:44.993564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T12:56:44.993564Z digest=sha256:30685f7d1ac1e1de333c3c37d2108e09d658ecae6e1ac793a2e63a4d980a4697