Pith. sign in

Paper Citation Record · LEDGER

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models

As of 20 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 2 inbound Pith citation observations for arXiv:2505.11271.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.11271 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:58:48.391076Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T06:17:20.819486Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cd33428e-6458-444b-8ef3-9d97e679c23e · outbound

This paper cites GPTCache: An open-source semantic cache for llm applications enabling faster answers and cost savings.

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models GPTCache: An open-source semantic cache for llm applications enabling faster answers and cost savings

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:58:49.197291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:58:47.825323Z digest=sha256:750fefbe8e474845e7c92203f2303f7c504ce56433db7c9802a1145e92c35a50

Observation ffb6cf9f-2042-4823-a7cf-bdf5a7a8f4b3 · outbound

This paper cites Language Models are Few-Shot Learners.

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Language Models are Few-Shot Learners

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:58:47.861219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:58:47.861219Z digest=sha256:209ffe787dabf6cce876e20afb9e341fc2e902107ebdbe9fbd1d8816e685db99

Observation c32646be-2548-44ff-8dd5-cd1c3132ae04 · outbound

This paper cites Don't Do RAG: When Cache-Augmented Generation is All You Need for Knowledge Tasks.

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Don't Do RAG: When Cache-Augmented Generation is All You Need for Knowledge Tasks

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:58:47.866645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:58:47.866645Z digest=sha256:ae2ec89a3aeaf4d67c6126245da5c94205bf6aba4d32b3801abe791c78e41bc5

Observation fbf82a27-7d6f-4cd4-94ae-4aca67b539d4 · outbound

This paper cites codefuse-ai/ModelCache, June 2024.

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models codefuse-ai/ModelCache, June 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:58:49.178625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:58:47.873691Z digest=sha256:2263167a470221660cc21618e8c4dda2192c19e0baf3805aecf61b50a9f32442

Observation 78cb19b7-9c97-46f6-bca4-c1048a37fb5a · outbound

This paper cites Langchain: Build applications with llms through composability.

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Langchain: Build applications with llms through composability

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:58:49.163241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:58:47.880205Z digest=sha256:5dca9e0542334ec96895cfbcd76f5d1a953790ba823abab7cd5c158c6ff4471e

Observation 59ff05e7-bbc7-4d02-803d-dda275165778 · outbound

This paper cites Model caches in LangChain.

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Model caches in LangChain

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:58:49.144289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:58:47.884990Z digest=sha256:0a69ca1204e5888cda932ea5c79ae4c855f9fd2091c38c15fea5487892bfc4da

Observation 18a497a0-47bb-44bd-9057-957e9eb4d3e6 · outbound

This paper cites Semantic cache for RAG using FAISS.

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Semantic cache for RAG using FAISS

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:58:49.119966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:58:47.893507Z digest=sha256:c1f337ad6e0da4bdd9da1b5a845a9be41a955a957e45d6356d6522acf827e425

Observation be467206-ff87-4076-aaac-155e0bca71d9 · outbound

This paper cites MeanCache: User-Centric Semantic Caching for LLM Web Services.

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models MeanCache: User-Centric Semantic Caching for LLM Web Services

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:58:47.899143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:58:47.899143Z digest=sha256:bb7abd9c332b1c5e66b7689edc03fb4ef19ebccb6adf2ac27157c0eaef98fe67

Observation 16ed6f4d-bfae-4ca9-baa0-85949ef8450c · outbound

This paper cites Prompt Cache: Modular Attention Reuse for Low-Latency Inference.

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:58:47.904448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:58:47.904448Z digest=sha256:057b4d170807ed2884d3d14b61c50756c50debd8470116439ba5f1fd9e873135

Observation 928ea50b-3914-431e-ae90-2c0002094ad7 · outbound

This paper cites EPIC: Efficient Position-Independent Caching for Serving Large Language Models.

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models EPIC: Efficient Position-Independent Caching for Serving Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:58:47.910040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:58:47.910040Z digest=sha256:828fc08ec0705b3196f00d7ff4d913cc2cd7bb1c0984cebbccf69db3573e96d1

Observation e4058a9f-2dcd-4658-b504-75b78e3cd4c5 · outbound

This paper cites LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models.

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:58:47.983020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:58:47.983020Z digest=sha256:e914b2e80a8533e94f682cd9189fea2ffa0d1013fd313e68d14cdb860619e019

Observation 7fc3b256-d5a7-4834-9354-f3de788b3547 · outbound

This paper cites LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression.

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:58:48.117812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:58:48.117812Z digest=sha256:880c553d083f7db12f944d23feb9815892b3728d8bd993c7dc6c8fb46e585d0d

Observation cfbdc121-e183-4f81-82f4-38ea3c3899be · outbound

This paper cites RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation.

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:58:48.148683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:58:48.148683Z digest=sha256:0cd251eb8079dc0d16c903fc366b680dc4df32abd1380579483e912e1295d705

Observation 966519c4-0449-46dd-b5ef-6209bcdaa207 · outbound

This paper cites Billion-scale similarity search with GPUs.

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Billion-scale similarity search with GPUs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:58:48.154091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:58:48.154091Z digest=sha256:f5df5ea87429e90e0462545d824838a6825fb9f22244e5e0240ab72e59d8eda6

Observation b0917477-83a9-4128-94c7-b98068c814a8 · outbound

This paper cites Weld, and Luke Zettlemoyer.

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Weld, and Luke Zettlemoyer

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:58:49.103682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:58:48.160403Z digest=sha256:a5a4080bcfbf99e4c8f500f27e8c2baf5765886e64cf4a842dde0d371e6fa82e

Observation f3a7dafd-539c-465e-a1ef-cc0917e7dab7 · outbound

This paper cites Dai, Jakob Uszko- reit, Quoc Le, and Slav Petrov.

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Dai, Jakob Uszko- reit, Quoc Le, and Slav Petrov

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:58:49.085213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:58:48.165768Z digest=sha256:8960c9ed56db1312e46a1baf57a11cf27601fbd74baa3f6c7320644a89b48309

Observation 9280a1d8-fcc1-40cd-a948-c95e244db4d4 · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention.

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:58:48.174430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:58:48.174430Z digest=sha256:37ea8511224c16080db76d39ee10b13e8bfbe911b0b755abec5db4502fe44845

Observation fe4127fd-aa75-4f0f-bee7-023313c7d340 · outbound

This paper cites SCALM: Towards Semantic Caching for Automated Chat Services with Large Language Models.

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models SCALM: Towards Semantic Caching for Automated Chat Services with Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:58:48.179534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:58:48.179534Z digest=sha256:76937c6e2b9dc5a383b69516ac9132d0170715ca14d1e3c7d52d7a68d47bd2aa

Observation a0f4fc94-cbe1-44fe-a23a-4464b548cc8d · outbound

This paper cites Context-based Semantic Caching for LLM Applications.

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Context-based Semantic Caching for LLM Applications

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:58:49.067541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:58:48.185101Z digest=sha256:1098bb43428ced0cb9fdab67e84afa6414a17bb99d0442207a8df723fa648f25

Observation df834436-ec53-4f93-a7ff-acac92568f97 · outbound

This paper cites Gpt-3.5 turbo model on OpenAI Platform.

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Gpt-3.5 turbo model on OpenAI Platform

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:58:49.028196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:58:48.191441Z digest=sha256:d3a64eda09cd5b813e4e933fc039424fcae602fd230efc70e3bfa7ae20b95e97

Observation e87241dd-17e3-4a1e-b6e6-87e2bc369140 · outbound

This paper cites LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression.

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:58:48.197690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:58:48.197690Z digest=sha256:255c1823ccdb0944e656dd1efa195b9382d82f3392145039e48a443a4dd792c3

Observation 6e52e12e-8b76-47aa-89de-addda9929c87 · outbound

This paper cites Sentence-bert: Sentence embeddings using siamese bert-networks, 2019.

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Sentence-bert: Sentence embeddings using siamese bert-networks, 2019

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:58:48.272255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:58:48.272255Z digest=sha256:7ac808026805f720f8ba0de57717b990fdc3e9988ff381b5f6ed677190e08de2

Observation 4363324f-61ae-4bfa-9c18-a1785c12e6c2 · outbound

This paper cites Sentence-BERT: Sentence embeddings using siamese bert-networks.

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Sentence-BERT: Sentence embeddings using siamese bert-networks

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:58:48.920842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:58:48.358604Z digest=sha256:873c2f5486174880f2aca247dd236d1e818682fb1b6416d44e51ca8a206d280c

Observation 83a41446-0680-4566-adff-4e1da676dc71 · outbound

This paper cites Hybrid-RACA: Hybrid retrieval-augmented composition assistance for real- time text prediction, 2024.

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Hybrid-RACA: Hybrid retrieval-augmented composition assistance for real- time text prediction, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:58:48.892886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:58:48.365047Z digest=sha256:4da866cbcf7b0ba7da3110d4a2e25e4019c92c9ebe3ff67d381515b3fb430c84

Observation e459a123-2da6-4993-b6c6-07d88dab5e34 · outbound

This paper cites CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion.

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:58:48.371131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:58:48.371131Z digest=sha256:eaad2a764538a2f594d89d97f37b0b664f34f68a67e61a3859face1a507fc32e

Observation 5444d93a-259e-44d9-b226-bbb4f6e5743e · outbound

This paper cites A Systematic Survey of Text Summarization: From Statistical Methods to Large Language Models.

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models A Systematic Survey of Text Summarization: From Statistical Methods to Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:58:48.376954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:58:48.376954Z digest=sha256:55c9d1b00b82ae3085c245fad0b9e7300b80cfe8e862bcf3b46f774fb1421b40

Observation 5d19d9de-cc62-49b4-8c6c-c4a9a08c03fd · outbound

This paper cites Hammerla.

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Hammerla

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:58:48.872405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:58:48.384443Z digest=sha256:b7903a783ef7a228b6a336b391d2a87be3e16b1c1a5c8d5c8188248f8c56f508

Observation 24ac40a4-c5f4-4249-81ae-72b87d789cfd · outbound

This paper cites Instcache: A predictive cache for llm serving, 2024.

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Instcache: A predictive cache for llm serving, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:58:48.820906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:58:48.391076Z digest=sha256:58d0cb89e68aaac4af7dfb0fc8ca507ce8ebc8ca99c84177835a2de48ee69d88

Pith citing papers

Observation cb43a98a-3435-4722-a3fc-6519ed874057 · inbound

From Similarity to Vulnerability: Key Collision Attack on LLM Semantic Caching cites this paper.

From Similarity to Vulnerability: Key Collision Attack on LLM Semantic Caching Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models

Reference 2002

Resolution
unresolved
no resolver link, observed 2026-08-03T06:17:20.819486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:17:20.819486Z digest=sha256:b043adae1b6131cd8b338b7a3f735ab73c68022bf92a3b222c1b69fa7e71d119

Observation 0e1c964f-bce9-49cf-8f53-f2cd2127aa30 · inbound

Mobility-Aware Cache Framework for Scalable LLM-Based Human Mobility Simulation cites this paper.

Mobility-Aware Cache Framework for Scalable LLM-Based Human Mobility Simulation Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T22:50:28.959689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:50:28.959689Z digest=sha256:61deb6bf431c541b7a755637e17cfea90d7c1aca4aff33f1025185e8de586e2b