Pith. sign in

Paper Citation Record · LEDGER

LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2410.00428.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.00428 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T10:10:23.098928Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T17:24:56.959321Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 702c9932-288c-49ce-8502-41fabbff8863 · inbound

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees cites this paper.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:23.098928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:23.098928Z digest=sha256:71af0882eb02073c9241d580b112ae746b441863f5543734279575bcc17ad7c2

Observation f770c5da-57e3-4456-a076-a60800a339bf · inbound

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration cites this paper.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.359687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.359687Z digest=sha256:519faf2178e2c36090482cda30b36eb09312cf3119f4a81fd8a211172150f1dd

Observation bca2aa03-58fd-45a2-adcc-31e71317e029 · inbound

Efficient Remote KV Cache Reuse with GPU-native Video Codec cites this paper.

Efficient Remote KV Cache Reuse with GPU-native Video Codec LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:22:22.824323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:3de4d95217f7429e3385ea91baeb1c85b22986bae6425a9d812dce914bfe547c

Observation 0b7c41ff-21a0-42bd-b047-83131df9db7a · inbound

ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache cites this paper.

ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:10:54.903875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:16:49.292491Z digest=sha256:adc07f339d74d8b41bf65970b17cc532568fd8219906699ac41413c670d2da63

Observation 92e32207-aca9-4950-81da-403be932bc1d · inbound

Adaptive KV Cache Reuse for Fast Long-Context LLM Serving cites this paper.

Adaptive KV Cache Reuse for Fast Long-Context LLM Serving LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:24:56.960758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T17:23:13.154458Z digest=sha256:3f8982ba1c1fac09bf2c583ed4b7343dbdf745092d23e3d38150e04a6732c6ce

Observation baea7759-e46b-49ec-9f53-6230a91772e4 · inbound

PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization cites this paper.

PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:10.825037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:10.825037Z digest=sha256:744d18f87748e0b0bf1d6adfb068c81df0e4555d3616564c7dca7424c9d80b55

Observation 2a1b2494-8811-4498-b14e-4b0f6b4fe1db · inbound

Beyond Storage: State as a Runtime Control Problem in Parallel and Distributed Systems cites this paper.

Beyond Storage: State as a Runtime Control Problem in Parallel and Distributed Systems LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

Reference 230

Resolution
unresolved
no resolver link, observed 2026-08-01T19:51:22.631756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:51:22.631756Z digest=sha256:d19bb14aef82c29143144db78f1f4d672f12da7bc619da8b55ad6c9dd23bdc89

Observation d91cb980-1b46-4f98-b2e4-c768f8990f3f · inbound

Persistent Computational State: A Session-Centric Runtime for Generative World Models cites this paper.

Persistent Computational State: A Session-Centric Runtime for Generative World Models LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T07:46:27.035461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:46:27.035461Z digest=sha256:15056cfd15d5a84954c00dc321d666d50d6e2f5db2cdff83b1ee80e5f76c9bb5