Pith. sign in

Paper Citation Record · LEDGER

LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2410.00428.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.00428 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T10:10:23.098928Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T17:24:56.959321Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 702c9932-288c-49ce-8502-41fabbff8863 · inbound

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees cites this paper.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:23.098928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:23.098928Z digest=sha256:7c36e3f4edf46303ef5aaa9a5615c53dc7e5734607451df028fb7d064e1687e4

Observation f770c5da-57e3-4456-a076-a60800a339bf · inbound

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration cites this paper.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.359687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.359687Z digest=sha256:6df4a224d6d748314fd9c5143039d7653e84e918f27840f2daba4d3cbc72c7c7

Observation bca2aa03-58fd-45a2-adcc-31e71317e029 · inbound

Efficient Remote KV Cache Reuse with GPU-native Video Codec cites this paper.

Efficient Remote KV Cache Reuse with GPU-native Video Codec LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:22:22.824323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:ae6931212425c6929feb0b11f4ebaf96cc2961a981060d196ebfcdf0d79090e3

Observation 0b7c41ff-21a0-42bd-b047-83131df9db7a · inbound

ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache cites this paper.

ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:10:54.903875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:16:49.292491Z digest=sha256:0f604d8e8d4370779b6137633dd66eaeea25bd5b92830ac5a1fe220a292ef744

Observation 92e32207-aca9-4950-81da-403be932bc1d · inbound

Adaptive KV Cache Reuse for Fast Long-Context LLM Serving cites this paper.

Adaptive KV Cache Reuse for Fast Long-Context LLM Serving LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:24:56.960758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T17:23:13.154458Z digest=sha256:90f5960efbd86ed44cf4cca6683507adee390357c84106626a0f30f2b63731b0

Observation baea7759-e46b-49ec-9f53-6230a91772e4 · inbound

PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization cites this paper.

PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:10.825037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:10.825037Z digest=sha256:d0de93727bf0b2fdb3729649ade51c1087268a5474bfe61b7c8b8f8b18c5d521

Observation 2a1b2494-8811-4498-b14e-4b0f6b4fe1db · inbound

Beyond Storage: State as a Runtime Control Problem in Parallel and Distributed Systems cites this paper.

Beyond Storage: State as a Runtime Control Problem in Parallel and Distributed Systems LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

Reference 230

Resolution
unresolved
no resolver link, observed 2026-08-01T19:51:22.631756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:51:22.631756Z digest=sha256:71882f0b95fed0a7dfe0c21c82e7f30a394840d768c68afbc0d6dd653b743459

Observation d91cb980-1b46-4f98-b2e4-c768f8990f3f · inbound

Persistent Computational State: A Session-Centric Runtime for Generative World Models cites this paper.

Persistent Computational State: A Session-Centric Runtime for Generative World Models LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T07:46:27.035461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:46:27.035461Z digest=sha256:2be6b10f180f0f02a4cb57bd7f3058fc1875ac152a235d69e3ff1135e03281f5