Pith. sign in

Paper Citation Record · LEDGER

InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2409.04992.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.04992 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:03:52.186188Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T15:08:33.737665Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3cc82a1b-c19a-45d0-89b6-ae7e0ea360c1 · inbound

Recursive Offloading for LLM Serving in Multi-tier Networks cites this paper.

Recursive Offloading for LLM Serving in Multi-tier Networks InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:52.186188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:03:52.186188Z digest=sha256:66055da898f6f9c34676d106d3231dae868d7ae2d6d96102893565e6936404e7

Observation ce740891-810a-48c2-b92e-2e3cb7a9f7f1 · inbound

Breaking the Boundaries of Long-Context LLM Inference: Adaptive KV Management on a Single Commodity GPU cites this paper.

Breaking the Boundaries of Long-Context LLM Inference: Adaptive KV Management on a Single Commodity GPU InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T23:00:11.154210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:00:11.154210Z digest=sha256:93f70260d45a21286c7e9c69063d47f32258a5d2a5ed1924b7f2210143820da1

Observation a6a83586-5834-405a-a6b6-f5e81933883a · inbound

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference cites this paper.

HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:08.996865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:08.996865Z digest=sha256:94adec27989343820dea3fd8772ca70bcfe20c97123181b921becf6094f37558

Observation da8dd679-23b6-47f7-b5a0-7e1daeef5325 · inbound

Efficient Remote KV Cache Reuse with GPU-native Video Codec cites this paper.

Efficient Remote KV Cache Reuse with GPU-native Video Codec InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:22:22.775393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:b59445b02588a999d0e50c7aa3799350f7da1f255701770a0a3aa73f1315aa24

Observation 7c106dcc-4808-49d2-8e4d-bcfd094ef7bf · inbound

SinkRouter: Sink-Aware Routing for Efficient Long-Context Decoding in Large Language and Multimodal Models cites this paper.

SinkRouter: Sink-Aware Routing for Efficient Long-Context Decoding in Large Language and Multimodal Models InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:57:15.446640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T07:56:05.583390Z digest=sha256:7be9eb698d16c195e784352344c2a61133c403ddd59f35d4efa5366166749f81

Observation 5cbaa4b8-62c4-4477-b3ba-0b73e4df97db · inbound

MCAP: Deployment-Time Layer Profiling for Memory-Constrained LLM Inference cites this paper.

MCAP: Deployment-Time Layer Profiling for Memory-Constrained LLM Inference InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:54:48.398544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T00:54:33.897112Z digest=sha256:14743982ab1af435435c5fe4c3cbee137a3f5979c3e90a54cf058c85ee220049

Observation d7d8662d-d962-417a-9f9e-fac135a34a56 · inbound

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference cites this paper.

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:06:26.525019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T12:51:42.916410Z digest=sha256:ab61ebe255cef7de9bf93afe91e250263063d4925bd585865bda3830d61969d3

Observation e031f882-aef6-4d03-8ef2-204c5f93705c · inbound

Can I Buy Your KV Cache? cites this paper.

Can I Buy Your KV Cache? InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:08:33.739576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T06:37:39.592571Z digest=sha256:fc9975692ab0646f238a6ffdf6fefcfd2ca20f70370411a8775637906c16d80a

Observation 36d558ef-bfc5-44bc-88d3-b2e0487f09cd · inbound

NELSSA: A GPU-PNM Heterogeneous System for Mixed-Length LLM Serving via Length-based Request Placement cites this paper.

NELSSA: A GPU-PNM Heterogeneous System for Mixed-Length LLM Serving via Length-based Request Placement InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T12:00:24.365208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:00:24.365208Z digest=sha256:0169f9a3c66a4d21b54a5f14d5130330a71b4bd537be3a9b189ad2ad757cabbe