Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:58:48.391076Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 2 inbound Pith citation observations for arXiv:2505.11271.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:58:48.391076Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T06:17:20.819486Z
A source-named dated measurement, never combined with another source.
Source: cited_works
28 of 28 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cd33428e-6458-444b-8ef3-9d97e679c23e · outbound
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models GPTCache: An open-source semantic cache for llm applications enabling faster answers and cost savings
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ffb6cf9f-2042-4823-a7cf-bdf5a7a8f4b3 · outbound
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Language Models are Few-Shot Learners
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c32646be-2548-44ff-8dd5-cd1c3132ae04 · outbound
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Don't Do RAG: When Cache-Augmented Generation is All You Need for Knowledge Tasks
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbf82a27-7d6f-4cd4-94ae-4aca67b539d4 · outbound
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models codefuse-ai/ModelCache, June 2024
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 78cb19b7-9c97-46f6-bca4-c1048a37fb5a · outbound
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Langchain: Build applications with llms through composability
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 59ff05e7-bbc7-4d02-803d-dda275165778 · outbound
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Model caches in LangChain
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 18a497a0-47bb-44bd-9057-957e9eb4d3e6 · outbound
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Semantic cache for RAG using FAISS
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation be467206-ff87-4076-aaac-155e0bca71d9 · outbound
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models MeanCache: User-Centric Semantic Caching for LLM Web Services
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16ed6f4d-bfae-4ca9-baa0-85949ef8450c · outbound
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Prompt Cache: Modular Attention Reuse for Low-Latency Inference
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 928ea50b-3914-431e-ae90-2c0002094ad7 · outbound
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models EPIC: Efficient Position-Independent Caching for Serving Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4058a9f-2dcd-4658-b504-75b78e3cd4c5 · outbound
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fc3b256-d5a7-4834-9354-f3de788b3547 · outbound
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfbdc121-e183-4f81-82f4-38ea3c3899be · outbound
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 966519c4-0449-46dd-b5ef-6209bcdaa207 · outbound
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Billion-scale similarity search with GPUs
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0917477-83a9-4128-94c7-b98068c814a8 · outbound
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Weld, and Luke Zettlemoyer
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f3a7dafd-539c-465e-a1ef-cc0917e7dab7 · outbound
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Dai, Jakob Uszko- reit, Quoc Le, and Slav Petrov
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9280a1d8-fcc1-40cd-a948-c95e244db4d4 · outbound
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe4127fd-aa75-4f0f-bee7-023313c7d340 · outbound
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models SCALM: Towards Semantic Caching for Automated Chat Services with Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0f4fc94-cbe1-44fe-a23a-4464b548cc8d · outbound
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Context-based Semantic Caching for LLM Applications
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation df834436-ec53-4f93-a7ff-acac92568f97 · outbound
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Gpt-3.5 turbo model on OpenAI Platform
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e87241dd-17e3-4a1e-b6e6-87e2bc369140 · outbound
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e52e12e-8b76-47aa-89de-addda9929c87 · outbound
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Sentence-bert: Sentence embeddings using siamese bert-networks, 2019
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4363324f-61ae-4bfa-9c18-a1785c12e6c2 · outbound
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Sentence-BERT: Sentence embeddings using siamese bert-networks
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 83a41446-0680-4566-adff-4e1da676dc71 · outbound
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Hybrid-RACA: Hybrid retrieval-augmented composition assistance for real- time text prediction, 2024
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e459a123-2da6-4993-b6c6-07d88dab5e34 · outbound
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5444d93a-259e-44d9-b226-bbb4f6e5743e · outbound
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models A Systematic Survey of Text Summarization: From Statistical Methods to Large Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d19d9de-cc62-49b4-8c6c-c4a9a08c03fd · outbound
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Hammerla
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 24ac40a4-c5f4-4249-81ae-72b87d789cfd · outbound
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Instcache: A predictive cache for llm serving, 2024
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation cb43a98a-3435-4722-a3fc-6519ed874057 · inbound
From Similarity to Vulnerability: Key Collision Attack on LLM Semantic Caching Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models
Reference 2002
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e1c964f-bce9-49cf-8f53-f2cd2127aa30 · inbound
Mobility-Aware Cache Framework for Scalable LLM-Based Human Mobility Simulation Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.