Pith. sign in

Paper Citation Record · LEDGER

Prompt Cache: Modular Attention Reuse for Low-Latency Inference

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2311.04934.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.04934 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:27:45.534694Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

12
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation dd159a32-1bc2-4141-8132-23950eee6907 · inbound

SGLang: Efficient Execution of Structured Language Model Programs cites this paper.

SGLang: Efficient Execution of Structured Language Model Programs Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:20:01.060678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:84f3a183e206aff7355ae821eee18b5dbd675c094db91ab6b5d87501445bb996

Observation faa06f4f-055a-4740-b990-f77bbe3509f4 · inbound

Accelerating Retrieval-Augmented Generation cites this paper.

Accelerating Retrieval-Augmented Generation Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T15:45:23.028887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:45:23.028887Z digest=sha256:31b1d00b97b6ced7a4df9e9f777428cabd1bf47dc1084d67f73b832ee739bce6

Observation f5c29039-d30e-42c7-8632-b2361c11eb78 · inbound

Offline Learning for Combinatorial Multi-armed Bandits cites this paper.

Offline Learning for Combinatorial Multi-armed Bandits Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-09T20:45:51.269548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:45:51.269548Z digest=sha256:04b9fa959d84fc0368f07c13e98ba9bc02755b6a19c3e600373ab6e53aaec3d0

Observation e53a4bae-80f1-46c9-b5d3-d5a35ebd824d · inbound

The Hitchhikers Guide to Production-ready Trustworthy Foundation Model powered Software (FMware) cites this paper.

The Hitchhikers Guide to Production-ready Trustworthy Foundation Model powered Software (FMware) Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:19.864940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:19.864940Z digest=sha256:8251c9837b19b84a444fcb31cb030d98c561bceb23eebace1c83e51e75459c16

Observation 16ed6f4d-bfae-4ca9-baa0-85949ef8450c · inbound

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models cites this paper.

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:58:47.904448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:58:47.904448Z digest=sha256:057b4d170807ed2884d3d14b61c50756c50debd8470116439ba5f1fd9e873135

Observation 0963909e-d08c-4d8d-a699-ec5caa12d701 · inbound

Efficient Remote KV Cache Reuse with GPU-native Video Codec cites this paper.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:22:22.751793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:6fe92421b083aa326ffdcbfcc9289175a4e43c4fa1465d3494c7d6e0df0910ed

Observation 92fa110f-13f9-48be-b398-b59bb357c2cb · inbound

PrefixWall: Mitigating Prefix Caching Side Channels in Shared LLM Systems cites this paper.

PrefixWall: Mitigating Prefix Caching Side Channels in Shared LLM Systems Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-21T12:15:06.759644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T12:14:09.509302Z digest=sha256:76934cbfc7b7f89a8b06d8aded25fccb34e2ce65e7e28620b69237dac5070f43

Observation 2a2e0a21-0aea-4f25-ba68-87e6aa32af2d · inbound

HieraSparse: Hierarchical Semi-Structured Sparse KV Attention cites this paper.

HieraSparse: Hierarchical Semi-Structured Sparse KV Attention Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T07:16:54.115010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T07:15:19.184970Z digest=sha256:71bafc72ecfb6abfbe4e0280b81ff208abb8ce4245fb4cd9fee3de4893320661

Observation bd4add9d-645e-4228-a9c8-d058c812b0d9 · inbound

Continuous Semantic Caching for Low-Cost LLM Serving cites this paper.

Continuous Semantic Caching for Low-Cost LLM Serving Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:01:03.329620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T02:32:28.923367Z digest=sha256:82d27cc5b4b596607815f69422fc40d27adb1a1dfe57e18a080233291c19cf75

Observation b3ac640f-1592-41ff-baf2-0a54928892e1 · inbound

Rethinking LLMOps for Fraud and AML: Building a Compliance-Grade LLM Serving Stack cites this paper.

Rethinking LLMOps for Fraud and AML: Building a Compliance-Grade LLM Serving Stack Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:12:07.117712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T02:11:18.259161Z digest=sha256:a117427b721c97fbcf4a6d0beae4d758ba1f44607fe1ba1a01dce7801a53f8f6

Observation c60f87ed-cae7-406e-aecf-459bebd1366a · inbound

CacheProbe: Auditing Prompt Cache Isolation in Gateway APIs cites this paper.

CacheProbe: Auditing Prompt Cache Isolation in Gateway APIs Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T06:23:08.727622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T06:22:28.511930Z digest=sha256:ae6efc23800c82956adbf6f5b633d6a49068db6fb84179a7489add18c87433c6

Observation bfaf9d39-e4d1-4421-be0b-24ae52176da5 · inbound

Move the Query, Not the Cache: Characterizing Cross-Instance Latent Attention Redistribution Across GPU Fabrics cites this paper.

Move the Query, Not the Cache: Characterizing Cross-Instance Latent Attention Redistribution Across GPU Fabrics Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:56:15.724672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T16:02:53.294235Z digest=sha256:e8c96d11e13f8f77273e8cb6907a6a135fa64d1b72ad47cbf940250e3fa71606

Observation 85dadf70-ad99-4a12-bbfd-c82eb014cff8 · inbound

SIFT: Selective-Index For Fast Compute of RAG Prefill by Exploiting Attention Invariance cites this paper.

SIFT: Selective-Index For Fast Compute of RAG Prefill by Exploiting Attention Invariance Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:37:31.262604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T16:24:31.109508Z digest=sha256:52ddb67b96e11f1034a12e191178274a1a7e3b4d9c12e8c8158e5e4c22728e9e

Observation 37c68cb6-0a91-4c8a-87bd-31d350c349f0 · inbound

MiniPIC: Flexible Position-Independent Caching in <100LOC cites this paper.

MiniPIC: Flexible Position-Independent Caching in <100LOC Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-27T07:20:42.194167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T07:12:45.449658Z digest=sha256:ce3321bcd8f745780f91dd720534bbc16ef40ae40362f5e50ed3a44f37890061

Observation a4ce376b-a86c-4d18-8d68-2843cdf4baaa · inbound

Execution-State Capsules: Graph-Bound Execution-State Checkpoint and Restore for Low-Latency, Small-Batch, On-Device Physical-AI Serving cites this paper.

Execution-State Capsules: Graph-Bound Execution-State Checkpoint and Restore for Low-Latency, Small-Batch, On-Device Physical-AI Serving Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:49:29.942545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T17:39:47.477338Z digest=sha256:ac5e8884b0ef9ccbed38a4903c29972b245716b93c02fcd5a629d4b77cba5e0c

Observation 26caed26-3309-46d9-89ef-878ff58d3e8a · inbound

CRAwLeR -- Cross-Reference Aware Legal Retrieval cites this paper.

CRAwLeR -- Cross-Reference Aware Legal Retrieval Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T07:49:39.642868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T12:37:33.877749Z digest=sha256:9f2b7a6edd3a14631baabfa72731763c803ba006eb6f4c4154933539c8a76461

Observation 624a7fda-bf59-4d3b-ac12-10ace44a4d26 · inbound

A Deterministic Control Plane for LLM Coding Agents cites this paper.

A Deterministic Control Plane for LLM Coding Agents Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-26T03:58:57.012444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-26T03:58:28.523545Z digest=sha256:2bbd51e337e26d9a72fba1579746fec28df8cbf59f8be1a5ba2efb6cad0ebcd6

Observation 93d08a1b-9ecd-4804-b9fc-8f6a1bac1d30 · inbound

KernelSight-LM: A Kernel-Level LLM Inference Simulator cites this paper.

KernelSight-LM: A Kernel-Level LLM Inference Simulator Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T00:54:06.243600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T00:48:19.207465Z digest=sha256:83824046344d9d4b2278956212aeb15d984fa1963c00d36774091e53cd9c534b

Observation 32144030-bfe6-4ae6-9006-c298a7780836 · inbound

KernelSight-LM: A Kernel-Level LLM Inference Simulator cites this paper.

KernelSight-LM: A Kernel-Level LLM Inference Simulator Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:19:02.214792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-03T23:09:38.092583Z digest=sha256:23d08539ae5e5fe884950926e76d5092064702aff0e1d08f4231abb870293653

Observation a13ebabe-ba70-4929-aba6-63c486ffc31a · inbound

Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel cites this paper.

Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T02:10:49.354541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:10:49.354541Z digest=sha256:63424dec249873b418fa33671ac3b5193a2bfece0c8226ec98c38e0a1348ba78

Observation de1d7fba-23a8-4471-8cb9-b49677e9a0b0 · inbound

Pixels for Programs? A Cross-Provider Case Study of Input-Token Accounting for Source Code as Text and Images cites this paper.

Pixels for Programs? A Cross-Provider Case Study of Input-Token Accounting for Source Code as Text and Images Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T08:35:41.852997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:35:41.852997Z digest=sha256:71ab50e3c4a5b14de83682eccb4430396dcd008e402d33ef3d3a7f994cf03582

Observation 0be0343d-266a-43a5-9403-52a7359ec686 · inbound

Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation cites this paper.

Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-30T23:47:57.858093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T23:47:57.858093Z digest=sha256:8c97aeb96535e5b762e61ea8809efd7a1f83cd45e8631590872e70601a75248d

Observation 5490f1a0-917e-4e28-9f39-e4d82775af4e · inbound

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving cites this paper.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.662090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.662090Z digest=sha256:363a8ec6d7d1453bea98fc8548af2440b43e30816fd38bbbe1311fd3f1e87ee3

Observation 406282b1-80b3-4997-bf03-855d288ec60d · inbound

Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory Systems cites this paper.

Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory Systems Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:45.534694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:45.534694Z digest=sha256:5ac15e09f8481b2ce5a1688eb3f73267f25b789e4b81e46039db16943c17f2fe