Pith. sign in

Paper Citation Record · LEDGER

MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2406.17565.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.17565 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T04:32:01.202215Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:39:46.803577Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3b653b39-854e-4c09-bc2c-fb4c64f5f2fc · inbound

BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching cites this paper.

BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-23T16:58:11.904020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T16:57:46.645061Z digest=sha256:fc45f4a314342d55f3574f350dde8db96f608855cc422c34b1ecc5e64862d330

Observation 08a01091-abe2-435f-a039-308069137136 · inbound

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference cites this paper.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T04:32:01.202215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:32:01.202215Z digest=sha256:20c090a18ed2637cb6486e9cdf4102b02e20a6a41b14756ae9b75a748d8f3a88

Observation 0b30afa7-a920-4a3a-92ee-b8aab4eef3cd · inbound

From Human Memory to AI Memory: A Survey on Memory Mechanisms in the Era of LLMs cites this paper.

From Human Memory to AI Memory: A Survey on Memory Mechanisms in the Era of LLMs MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 122

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:05:09.848606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T11:05:09.588491Z digest=sha256:a4daa8679e5c464735415f2a28a78a7184820fcb69ae403b1fe32ac5db1f7f4d

Observation 943e6b8a-7059-48d4-b290-02b90832b38b · inbound

CoDec: Prefix-Shared Decoding Kernel for LLMs cites this paper.

CoDec: Prefix-Shared Decoding Kernel for LLMs MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:55.773082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:55.773082Z digest=sha256:3d2ce5717957c243c95491075f87de84dffa823a748debfe4933bf6c205307cf

Observation 3f1b93ad-e1b0-450a-ac38-19ff0a11b44e · inbound

Beyond the Buzz: A Pragmatic Take on Inference Disaggregation cites this paper.

Beyond the Buzz: A Pragmatic Take on Inference Disaggregation MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:26:32.777159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:32.777159Z digest=sha256:d89ccd9789d37a5dd391d45bb2fc1ba189a0e41e9e925071523bb33bc83db4f0

Observation 964a02ee-e3b0-467c-b77d-b3cd16ae2297 · inbound

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration cites this paper.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.319925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.319925Z digest=sha256:0750246203eefe5a9205b939da423cf9ec365aa19553109f994bed4bc2cbdf11

Observation cb7ffe8d-b83e-4d3b-9d8f-7859cb9a33bf · inbound

Rethinking Caching for LLM Serving Systems: Beyond Traditional Heuristics cites this paper.

Rethinking Caching for LLM Serving Systems: Beyond Traditional Heuristics MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T16:21:09.133223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:21:09.133223Z digest=sha256:fee9f12f5f9e8d090109e2549a9503f66595e31e1cd64bf070df755cebd7e4c8

Observation 3b62d94d-a129-4d69-9b26-5a56356e42ca · inbound

Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference cites this paper.

Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:36.304749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:36.304749Z digest=sha256:57790b7a1e1020f11b7b49a015c997f1159aeb22399be4735311605bf4de0175

Observation 26143fc2-cb77-4548-87f0-3d769c800d34 · inbound

Secure Multi-LLM Agentic AI and Agentification for Edge General Intelligence by Zero-Trust: A Survey cites this paper.

Secure Multi-LLM Agentic AI and Agentification for Edge General Intelligence by Zero-Trust: A Survey MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 121

Resolution
unresolved
no resolver link, observed 2026-08-05T15:26:33.726798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:26:33.726798Z digest=sha256:e9a37d3f057804a48daa6f90e34c7f21a2fbc5787b3bd5ef39bfb3c7cbc9766e

Observation 26a19a2d-3907-4e66-92b0-bba1b5cab032 · inbound

Boosting Embodied AI Agents through Perception-Generation Disaggregation and Asynchronous Pipeline Execution cites this paper.

Boosting Embodied AI Agents through Perception-Generation Disaggregation and Asynchronous Pipeline Execution MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:12.680361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:12.680361Z digest=sha256:76b0fd2b1e4f60be6a4a4dec418fc261fa822e94f8af4a0398cf5af8fcc451a2

Observation ae877024-34f8-43a3-9549-b818ea3d9c92 · inbound

DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing cites this paper.

DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T23:39:06.591452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:39:06.591452Z digest=sha256:a4638f7459498b2618163c9496e90acde420bc47282de14ee56c0faeea682be4

Observation 4066d4c7-8537-4b24-91ea-a2cc0bb9a66a · inbound

SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips cites this paper.

SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:30:17.918779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T15:26:01.283448Z digest=sha256:6e72dc13576bb2f023bb9b31bca378342149292d3512e6a3880a83978a71f4b8

Observation 47bd091b-30b8-416a-a99f-bbff14be2f8f · inbound

Efficient Remote KV Cache Reuse with GPU-native Video Codec cites this paper.

Efficient Remote KV Cache Reuse with GPU-native Video Codec MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:22:22.835574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:1cea382c1e483a919a282f8cd493119fe5fe7372dcf383b407b59ba9db52a5f3

Observation da2b1e96-3bd8-47b1-81d1-29da03e930ce · inbound

Towards Efficient Large Vision-Language Models: A Comprehensive Survey on Inference Strategies cites this paper.

Towards Efficient Large Vision-Language Models: A Comprehensive Survey on Inference Strategies MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:43:00.694286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T21:40:55.343169Z digest=sha256:833035714585512d7722b97924d0f9ad1672983fe64260775e1e31993b5dbd11

Observation 3101e9b5-02e8-46ba-8f2e-d1bca161f339 · inbound

InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models cites this paper.

InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:55:59.356849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:23:40.872418Z digest=sha256:6a3151ac60c3050493b7a8ba50d58354df7d6faba3f432739185722bb678ccdf

Observation fd481473-3656-4800-8567-ddac1ac91067 · inbound

Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference cites this paper.

Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:25:46.412499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T18:27:13.781231Z digest=sha256:637a8f7e57b76341a70118e55d91eb4aefdc81de68192aaf0959a07711484659

Observation aef8e12f-c4e3-4c25-9137-6f58e5ec7722 · inbound

Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference cites this paper.

Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:35:10.226400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T00:31:12.965178Z digest=sha256:562930649dec1cae5f8095220696c64f7d9677765a5ea4d2b314b707ec10f742

Observation e463e356-0d2e-4706-94b3-2b2a0c63b35a · inbound

ObjectCache: Layerwise Object-Storage Retrieval for KV Cache Reuse cites this paper.

ObjectCache: Layerwise Object-Storage Retrieval for KV Cache Reuse MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-25T00:25:08.358107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:22:05.867914Z digest=sha256:8c76f11a682bef51a23ff6017643fbe8a23d52ae8f5c1da799f219993b006a20

Observation e820ca28-003b-4c8e-a73c-6b4fa10c1ba3 · inbound

AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System cites this paper.

AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:15:17.299404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T03:12:49.028342Z digest=sha256:a9741d26a6b934be8b90d1c676ae2a4034438880de6388423e87c5494d61982c

Observation 9c02aad4-b481-4ca1-877e-d54c2f24c980 · inbound

Idleness is Relative: Exploiting Tool-Call Idle Windows for Offloading in Agentic Systems with MORI cites this paper.

Idleness is Relative: Exploiting Tool-Call Idle Windows for Offloading in Agentic Systems with MORI MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:06:13.627198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T17:30:56.324289Z digest=sha256:f45ec0ed1a100a9d51fba6393309aa95aeed8d725e2b8f9c404a433bd29ca650

Observation 2fd66820-e86d-4091-940f-115992c883d0 · inbound

Leyline: KV Cache Directives for Agentic Inference cites this paper.

Leyline: KV Cache Directives for Agentic Inference MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T21:36:14.553901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T16:51:08.114092Z digest=sha256:f429d40d564187e501e1f8a343c8b99d75d1df7727bc13aef1ea69ccdc72a5fb

Observation 116ba9ea-3c2a-4e3b-8b0d-22aa9632bda0 · inbound

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill cites this paper.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:39:46.805074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:dbeb769ce8c361f9e41f6e1d5ca9aaa515d3405abaee67d2e855cf5dd0393f34

Observation ef9aefc8-c8cb-4722-b9f5-3ab05fb142b0 · inbound

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding cites this paper.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-30T03:04:14.223576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:a6e7edab97afd26a4bcf20db049ad5e5a997bffd3bee00036aa2acb0dabab728

Observation fd7edb7a-853e-4df5-93bf-7fe941c9f6ca · inbound

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving cites this paper.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:4c7725a895b196eb0d55484dc4fbe7a2db978113b7190354a6d4fad460bc2c46

Observation 9bd133b5-8c0d-498d-af8e-8c50ce925c61 · inbound

[AAFLOW+] Stateful Operator Abstraction with Zero-Copy Distributed KV Cache Orchestration for Multi-Agent Workflows cites this paper.

[AAFLOW+] Stateful Operator Abstraction with Zero-Copy Distributed KV Cache Orchestration for Multi-Agent Workflows MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T07:51:15.549754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:51:15.549754Z digest=sha256:8f1022dc3d18e82ab105e7e474837293be06e6726527f6c939ec3cdde5aed62e

Observation 98f1939f-4852-44c6-ad80-d2c249913367 · inbound

DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch cites this paper.

DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T15:06:54.222623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:06:54.222623Z digest=sha256:da60e1d912da0ca44ab49d01e86878caf83b2075e74043efc553ee3e477842c4

Observation c63fd332-4148-4190-a355-98bd631e8bc3 · inbound

Rethinking AI Cloud Infrastructure for Agentic Serving Systems with the Aries Experimentation Framework cites this paper.

Rethinking AI Cloud Infrastructure for Agentic Serving Systems with the Aries Experimentation Framework MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T14:14:10.290433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:14:10.290433Z digest=sha256:cc469be3d0949a3f5c004764af5b9f55b7b1b74623e531044eefdec34bc67239

Observation 251d395c-dede-4ccd-bd4c-523d8ccf8a65 · inbound

Spatial Prefix Caching for Wireless Edge LLM Inference: A Stochastic-Geometry and Queueing Framework cites this paper.

Spatial Prefix Caching for Wireless Edge LLM Inference: A Stochastic-Geometry and Queueing Framework MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T00:33:27.394771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:33:27.394771Z digest=sha256:37b96e199d7ace553ea5457920429ce5951ad1f295c0cf887eac9cee9ce34993

Observation 2b17282e-2cbb-4913-896d-4eaa09d08b5c · inbound

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling cites this paper.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:22.840559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:22.840559Z digest=sha256:00e379c93ebc584806217b518564f24bc798546dac8dcc70c4c865296b03fd42