Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2311.04934.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:27:45.534694Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
12
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation dd159a32-1bc2-4141-8132-23950eee6907 · inbound
SGLang: Efficient Execution of Structured Language Model Programs Prompt Cache: Modular Attention Reuse for Low-Latency Inference
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation faa06f4f-055a-4740-b990-f77bbe3509f4 · inbound
Accelerating Retrieval-Augmented Generation Prompt Cache: Modular Attention Reuse for Low-Latency Inference
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5c29039-d30e-42c7-8632-b2361c11eb78 · inbound
Offline Learning for Combinatorial Multi-armed Bandits Prompt Cache: Modular Attention Reuse for Low-Latency Inference
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e53a4bae-80f1-46c9-b5d3-d5a35ebd824d · inbound
The Hitchhikers Guide to Production-ready Trustworthy Foundation Model powered Software (FMware) Prompt Cache: Modular Attention Reuse for Low-Latency Inference
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16ed6f4d-bfae-4ca9-baa0-85949ef8450c · inbound
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Prompt Cache: Modular Attention Reuse for Low-Latency Inference
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0963909e-d08c-4d8d-a699-ec5caa12d701 · inbound
Efficient Remote KV Cache Reuse with GPU-native Video Codec Prompt Cache: Modular Attention Reuse for Low-Latency Inference
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 92fa110f-13f9-48be-b398-b59bb357c2cb · inbound
PrefixWall: Mitigating Prefix Caching Side Channels in Shared LLM Systems Prompt Cache: Modular Attention Reuse for Low-Latency Inference
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2a2e0a21-0aea-4f25-ba68-87e6aa32af2d · inbound
HieraSparse: Hierarchical Semi-Structured Sparse KV Attention Prompt Cache: Modular Attention Reuse for Low-Latency Inference
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bd4add9d-645e-4228-a9c8-d058c812b0d9 · inbound
Continuous Semantic Caching for Low-Cost LLM Serving Prompt Cache: Modular Attention Reuse for Low-Latency Inference
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b3ac640f-1592-41ff-baf2-0a54928892e1 · inbound
Rethinking LLMOps for Fraud and AML: Building a Compliance-Grade LLM Serving Stack Prompt Cache: Modular Attention Reuse for Low-Latency Inference
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c60f87ed-cae7-406e-aecf-459bebd1366a · inbound
CacheProbe: Auditing Prompt Cache Isolation in Gateway APIs Prompt Cache: Modular Attention Reuse for Low-Latency Inference
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bfaf9d39-e4d1-4421-be0b-24ae52176da5 · inbound
Move the Query, Not the Cache: Characterizing Cross-Instance Latent Attention Redistribution Across GPU Fabrics Prompt Cache: Modular Attention Reuse for Low-Latency Inference
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 85dadf70-ad99-4a12-bbfd-c82eb014cff8 · inbound
SIFT: Selective-Index For Fast Compute of RAG Prefill by Exploiting Attention Invariance Prompt Cache: Modular Attention Reuse for Low-Latency Inference
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 37c68cb6-0a91-4c8a-87bd-31d350c349f0 · inbound
MiniPIC: Flexible Position-Independent Caching in <100LOC Prompt Cache: Modular Attention Reuse for Low-Latency Inference
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a4ce376b-a86c-4d18-8d68-2843cdf4baaa · inbound
Execution-State Capsules: Graph-Bound Execution-State Checkpoint and Restore for Low-Latency, Small-Batch, On-Device Physical-AI Serving Prompt Cache: Modular Attention Reuse for Low-Latency Inference
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 26caed26-3309-46d9-89ef-878ff58d3e8a · inbound
CRAwLeR -- Cross-Reference Aware Legal Retrieval Prompt Cache: Modular Attention Reuse for Low-Latency Inference
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 624a7fda-bf59-4d3b-ac12-10ace44a4d26 · inbound
A Deterministic Control Plane for LLM Coding Agents Prompt Cache: Modular Attention Reuse for Low-Latency Inference
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 93d08a1b-9ecd-4804-b9fc-8f6a1bac1d30 · inbound
KernelSight-LM: A Kernel-Level LLM Inference Simulator Prompt Cache: Modular Attention Reuse for Low-Latency Inference
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 32144030-bfe6-4ae6-9006-c298a7780836 · inbound
KernelSight-LM: A Kernel-Level LLM Inference Simulator Prompt Cache: Modular Attention Reuse for Low-Latency Inference
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a13ebabe-ba70-4929-aba6-63c486ffc31a · inbound
Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel Prompt Cache: Modular Attention Reuse for Low-Latency Inference
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de1d7fba-23a8-4471-8cb9-b49677e9a0b0 · inbound
Pixels for Programs? A Cross-Provider Case Study of Input-Token Accounting for Source Code as Text and Images Prompt Cache: Modular Attention Reuse for Low-Latency Inference
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0be0343d-266a-43a5-9403-52a7359ec686 · inbound
Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation Prompt Cache: Modular Attention Reuse for Low-Latency Inference
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5490f1a0-917e-4e28-9f39-e4d82775af4e · inbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Prompt Cache: Modular Attention Reuse for Low-Latency Inference
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 406282b1-80b3-4997-bf03-855d288ec60d · inbound
Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory Systems Prompt Cache: Modular Attention Reuse for Low-Latency Inference
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.