Pith. sign in

Paper Citation Record · LEDGER

vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2405.04437.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.04437 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T18:19:21.033324Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T01:36:44.326797Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7ed2f0e2-9307-4a4f-a90d-0755c556733b · inbound

Position: AI Scaling: From Up to Down and Out cites this paper.

Position: AI Scaling: From Up to Down and Out vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T18:19:21.033324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:19:21.033324Z digest=sha256:20665138089534b6abd03b9b9a3e661a8f7cffe4a0e2c45f8cdc447551efa1c9

Observation e187508e-9ad4-4c8d-8242-60c8b694c8f3 · inbound

AutoChemSchematic AI: Agentic Physics-Aware Automation for Chemical Manufacturing Scale-Up cites this paper.

AutoChemSchematic AI: Agentic Physics-Aware Automation for Chemical Manufacturing Scale-Up vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:29.351522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:29.351522Z digest=sha256:b4e761e5d58d624abdd91cfe35ebda88007475c34d96b82caf66969e3c3977ae

Observation 4cbac68f-45a7-42af-88ab-5209a864596e · inbound

eLLM: Elastic Memory Management Framework for Efficient LLM Serving cites this paper.

eLLM: Elastic Memory Management Framework for Efficient LLM Serving vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:42:13.910745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T09:41:45.544399Z digest=sha256:42690f313b75e5fc37f76a828fd64776b6f3114da413c6d401c975bc4d50fd58

Observation 5cdbeddd-2d21-4162-aa5f-5982cab441b1 · inbound

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference cites this paper.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:09.183968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:09.183968Z digest=sha256:77f8cabb34033d83922d51d555897ebc4b273a6c92e76a0102d093e4f212cc77

Observation 0f85f268-14ee-492e-93a6-fd93f1c5731b · inbound

Parallelizing Tool Execution and LLM Generation for Low-Latency Agent Serving cites this paper.

Parallelizing Tool Execution and LLM Generation for Low-Latency Agent Serving vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-13T22:20:51.838309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:20:51.838309Z digest=sha256:179f3fddfda2c8e2e9fb87b9ed40a0f7fcc8eb3394fc46b66f835903c7c34ff2

Observation deea181f-38ac-4696-886d-b92bb444cd29 · inbound

cuRAMSES: Scalable AMR Optimizations for Large-Scale Cosmological Simulations cites this paper.

cuRAMSES: Scalable AMR Optimizations for Large-Scale Cosmological Simulations vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-13T09:09:32.932051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:09:32.932051Z digest=sha256:250b46abb522056801db67c7f1987328f89e0a3681b325531560a9c2e51b42fe

Observation 9d0d39ea-3a51-4936-870d-f1591f511c78 · inbound

CodecSight: Leveraging Video Codec Signals for Efficient Streaming VLM Inference cites this paper.

CodecSight: Leveraging Video Codec Signals for Efficient Streaming VLM Inference vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:00:52.993784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:45:26.555395Z digest=sha256:58ece175e48e5514e87baad0730773b43143768c48180dd29c823e4e3a44eeb0

Observation 1421fa3b-f095-4da7-88e5-d953020acd85 · inbound

KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving cites this paper.

KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:51:27.063199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:57:30.104609Z digest=sha256:d444986c04d3154db65b847221a3cb1082a5c2edf6e666b16a79e74726f3ac6e

Observation 85c56d4b-216e-499d-8bbe-7f370039632a · inbound

KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving cites this paper.

KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.450518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T08:05:44.256565Z digest=sha256:1d5c4f78a8e0d95669c53396fa2cfd564317eb24b53f7c1cc1c92c5e3302a8e9

Observation 861b87d2-2475-4f59-9a8c-87be60491b3a · inbound

Token-Operations-Oriented Inference Optimization Techniques for Large Models cites this paper.

Token-Operations-Oriented Inference Optimization Techniques for Large Models vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

Reference 200

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:09:36.541010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T16:15:22.543601Z digest=sha256:35105f7298974aa27523b485cc75d03adbfb43e08bc71f8c344bd9c21856217d

Observation 420d3c7e-fe8b-4d58-91c7-782b6db6bb72 · inbound

Token-Operations-Oriented Inference Optimization Techniques for Large Models cites this paper.

Token-Operations-Oriented Inference Optimization Techniques for Large Models vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

Reference 187

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:22.213295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:22.213295Z digest=sha256:02746f52f830c86dac7d1b83147f318dcf1f237ab3f1b53064ac82b2b99ff947

Observation 2ead7c60-3ad8-49b8-a333-4a79cadf0d32 · inbound

Execution-State Capsules: Graph-Bound Execution-State Checkpoint and Restore for Low-Latency, Small-Batch, On-Device Physical-AI Serving cites this paper.

Execution-State Capsules: Graph-Bound Execution-State Checkpoint and Restore for Low-Latency, Small-Batch, On-Device Physical-AI Serving vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:49:29.923241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T17:39:47.477338Z digest=sha256:cfd1d38208b25852c0994f023ad4b2169432f0e3f4b74b15750bc2e14df76c47

Observation c6687955-f488-49d7-b25e-5c4ee1f5523a · inbound

Keyless Attention: Value-Space Routing and Value-Only Caching for Efficient Transformers cites this paper.

Keyless Attention: Value-Space Routing and Value-Only Caching for Efficient Transformers vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T10:40:08.018921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:40:08.018921Z digest=sha256:c3cb9db976e47c12755e1eec40c1506e51a37ee355fcffb1e8811f3a794cf4de

Observation 0778cff1-9044-4a61-a263-8fbf6d1c9011 · inbound

Keyless Attention: Value-Space Routing and Value-Only Caching for Efficient Transformers cites this paper.

Keyless Attention: Value-Space Routing and Value-Only Caching for Efficient Transformers vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T02:48:35.073858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:48:35.073858Z digest=sha256:1878fd53827d565a31baf709139716e0e919def7ce0ac912dfbbfd3bae04000b

Observation d400d8ad-2f2c-46c2-9112-85c7cd931a85 · inbound

PersistentKV: Page-Aware Decode Scheduling for Long-Context LLM Serving on Commodity GPUs cites this paper.

PersistentKV: Page-Aware Decode Scheduling for Long-Context LLM Serving on Commodity GPUs vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:19:50.308008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T05:23:08.673299Z digest=sha256:0892cf5c57d6223e91bf9a2efa01eba4c990cb58920e4e57ea3aeee1b096d4df

Observation 1b961289-45f5-427e-93d7-a5c03cd83641 · inbound

PersistentKV: Page-Aware Decode Scheduling for Long-Context LLM Serving on Commodity GPUs cites this paper.

PersistentKV: Page-Aware Decode Scheduling for Long-Context LLM Serving on Commodity GPUs vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:07:23.419736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-02T21:04:43.874528Z digest=sha256:9c974b2d0c16072c5481b3a6d2c9d3930dad1dd775d8596c9cec47a812d1fc93

Observation a56edc14-4e35-4800-815f-b329b3a2a82f · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention

Reference 95

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:44.327959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:fc567809a438092fee8c6878f03e9e3f7cf8714e59c8f907e5c7ca3ed7f41c9c