Pith. sign in

Paper Citation Record · LEDGER

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees

As of 8 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 1 inbound Pith citation observation for arXiv:2502.08182.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.08182 v1

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T10:10:23.140836Z

measured 88 of 88 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-21T15:26:01.283448Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T15:30:17.957796Z

Reference resolution

87 of 87 outbound references displayed

  • verified exact18
  • verified fuzzy14
  • unresolved55
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dfd3e3b5-9f1e-48c8-bd5b-a897ed731d4c · outbound

This paper cites URL https://huggingface.co/ docs/accelerate/index.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees URL https://huggingface.co/ docs/accelerate/index

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.710232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.710232Z digest=sha256:11d5ad2bc5b59432a53f83153511537c1dbd19079eb39ac721cc8ce33f347ed4

Observation ea0366ba-6bb8-4a30-9625-dce2bdc4c27f · outbound

This paper cites URL https://sharegpt.com/.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees URL https://sharegpt.com/

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.716168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.716168Z digest=sha256:16c43008be538ad601aabfb18b6d9ce70d0b8bcaef4b7aade9dced0b49ac5eb9

Observation c809b2a0-2bd0-4602-9c5b-dbf0129b93cd · outbound

This paper cites Acharya, B.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Acharya, B

Reference 3

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T10:10:26.250091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:22.721928Z digest=sha256:06fd0a66387778b0eb833396636d75fe9927a5603319acc35dfc4f0e556db39e

Observation aea5f611-4591-4f01-8d04-77904f1a748f · outbound

This paper cites SYMPHONY: Improving Memory Management for LLM Inference Workloads.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees SYMPHONY: Improving Memory Management for LLM Inference Workloads

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-08T10:10:26.078778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:22.727386Z digest=sha256:e5f978d6469a18810f665b4524208a04378592729d480ba08733f312c531f774

Observation b16632af-cffe-4664-a125-b340c71fcf6f · outbound

This paper cites SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.732763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.732763Z digest=sha256:07bcba37ce8bc4b5b9385eb6d118d2ced8d16855453dd792315eb168e9c6fa9b

Observation e73bab42-7498-48f5-bf9d-7c33882b4062 · outbound

This paper cites LLM in a flash: Efficient Large Language Model Inference with Limited Memory.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.738166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.738166Z digest=sha256:486e15a5cf7b7e98013ed3d104269ff5f0d1f9ef66de3eb65d9f51086f3327e2

Observation cf37bd70-da3b-4841-97b7-fd4028c59005 · outbound

This paper cites Alomari, N.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Alomari, N

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.743842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.743842Z digest=sha256:99785a08610e6322cca5b9a35dbecb45ddc73b88d30b90053ad8aef2bb6489f9

Observation 64d1de85-13b0-46c8-b1a0-474069fc9f6b · outbound

This paper cites Neutrino Production via $e^-e^+$ Collision at $Z$-boson Peak.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Neutrino Production via $e^-e^+$ Collision at $Z$-boson Peak

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.753134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.753134Z digest=sha256:2c6dbd80e4bc697df49144b89a229a3533bf2c2b15a13c34106ba38e6e581403

Observation bd779b54-fef4-4499-b02c-d0d4df4e3041 · outbound

This paper cites Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.757378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.757378Z digest=sha256:af899523f40379340c0280fecc5227e30f78f49e1b01b4b8197df65314a04714

Observation 99c060a7-b00f-4a3d-be13-bf570ed796c5 · outbound

This paper cites AWESOME: GPU Memory-constrained Long Document Summarization using Memory Mechanism and Global Salient Content.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees AWESOME: GPU Memory-constrained Long Document Summarization using Memory Mechanism and Global Salient Content

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-08T10:10:25.989475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:22.761709Z digest=sha256:e5c9265238d0cc22ece253746d0aab86eb66de8545c15d0f37cfeaccdf57771d

Observation 2e3514d0-fad9-46f4-b93b-1a6e6adea279 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Evaluating Large Language Models Trained on Code

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.765947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.765947Z digest=sha256:8ac3e944c476a5b27bf255d2bde24413e676045befc8fee89a1e2bc9dcfda6cf

Observation 31fa23cf-a362-426c-bd5e-0524c2d95eaa · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.770073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.770073Z digest=sha256:ce645cf520229c0ab56183afb7b03afc26ed8f8db04971336efb84b25dcb1df2

Observation 3ec6d3e6-3e3c-4a2c-8d75-392c61dbe549 · outbound

This paper cites Cheng, Y.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Cheng, Y

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.779090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.779090Z digest=sha256:35d54ade662b56eea97ef98aac37e8290786a7cd016001c18120f16a10d92355

Observation 361882a4-e08b-49c1-b7f5-b9e10f42196d · outbound

This paper cites Choquette, W.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Choquette, W

Reference 14

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T10:10:25.739015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:22.783485Z digest=sha256:3c59fdfd3d48d561b696779674b7f1cea485106b7a4de31caa7dff172bd1c228

Observation a217b87c-3eb6-4156-8223-f9e31b9318d4 · outbound

This paper cites Crankshaw, X.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Crankshaw, X

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.788626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.788626Z digest=sha256:20c1a25ce98cae6a9cf128d00a91c14c0231409801bde9ad69e2a4a38c3e7b6d

Observation 0eb26b63-d62b-433a-b9dd-49bd525b9a5e · outbound

This paper cites Crankshaw, G.-E.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Crankshaw, G.-E

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.794421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.794421Z digest=sha256:a52e66b08ed92afc3aaca57b8ba4d8aacd8bd8dee21411e6f45790e71a921abd

Observation 8864959f-64f8-4c43-8edb-a01c5363656d · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.804664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.804664Z digest=sha256:7d409ecafe041d4fd8812942bf96480fdc1aaa70189b592c90155a72a64fc82d

Observation 2d267163-9e1d-4904-b99b-ae00d8b36c70 · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-08T10:10:26.626313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:22.809216Z digest=sha256:7ee71a51bf518d55847a7a7498c089660b5c60b10bba0f72e624eca651969820

Observation c055ce8a-9ab6-43ec-9bda-0a119794ef09 · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-08T10:10:26.612178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:22.815367Z digest=sha256:342342c17250ddcfb33a2dffc9aa3bd6e3c88906ee08497ae43b5836f2ebca81

Observation 868c2bb5-b7aa-4d78-b8f9-516d52af42b3 · outbound

This paper cites Improving LLM Abilities in Idiomatic Translation.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Improving LLM Abilities in Idiomatic Translation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.822360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.822360Z digest=sha256:1edfdbdbe0fd77d60f3eadca0c1188fdb3096c1ef424fab1dff297da709479dd

Observation b7d1cb3d-82fc-453c-be89-0574bc0e903d · outbound

This paper cites Efficient Training of Large Language Models on Distributed Infrastructures: A Survey.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Efficient Training of Large Language Models on Distributed Infrastructures: A Survey

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.827490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.827490Z digest=sha256:4d8cb97893dca44f084a681bf42165f8f7d39f11b270d68eadd409d56d412ef1

Observation ad2081ed-afb5-4bab-9329-e09f717e306c · outbound

This paper cites Elliott, M.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Elliott, M

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:10:26.598091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:22.832686Z digest=sha256:3a8f8133de4353f8f8b33cc4bb005b677388cd47e61336e413b491733ffe1608

Observation f513ccb5-2799-4452-84e7-1bf801a016da · outbound

This paper cites Gambhir and V.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Gambhir and V

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:10:26.584171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:22.837711Z digest=sha256:5176ff7d9b37903d07bc0a3597ddde2db189c02c611bf31fb04437af0ebeab4e

Observation aab287a1-5331-43ed-9816-6b4d6f824e8a · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-08T10:10:26.569988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:22.842527Z digest=sha256:db425e894e5f6829d01e232659de2ce43ec510bd87c0a6017ec7047c681254c3

Observation 05f2565b-fce5-4450-aa48-bc54c5e0504e · outbound

This paper cites Fast State Restoration in LLM Serving with HCache.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Fast State Restoration in LLM Serving with HCache

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.848211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.848211Z digest=sha256:54c542078fb50463012e7b298b19d1140da053483b8b6ed54e8a71c721e8bb9b

Observation 0875813d-2c66-4ca6-bea5-74fc4b222a31 · outbound

This paper cites M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.853443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.853443Z digest=sha256:3965657ad418612df1d1e6899ad28ff3df33c0cd4edab27a3eca8ac31595c4c9

Observation 079c3fe8-7d92-4e73-9b8f-29ff9a898f21 · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 27

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T10:10:25.293444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:22.858724Z digest=sha256:5758fdae89e3216f0b2204603297f9e06372b1fbeb5c8d70d4dd7de527e9c01a

Observation 3e9542ac-3d38-458a-9453-dbf985ba63f4 · outbound

This paper cites Gujarati, R.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Gujarati, R

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:10:26.554673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:22.863209Z digest=sha256:25e4aecf40d8b23ce4fda20b78cd2b509defae9bc8aca65411ca1a2e9284ea12

Observation 6c2f1dc1-ed41-4cce-8198-907b1da0d872 · outbound

This paper cites Gujarati, R.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Gujarati, R

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:10:26.540571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:22.867754Z digest=sha256:a78a9ad265720e93246339588713435d21e967fb0a572bf95ae759a45b08c98b

Observation 6bedb766-c3d8-4009-a6f4-8cabf5d05c7c · outbound

This paper cites DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.872509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.872509Z digest=sha256:1fae67ba280544561033f8dfb36d96d97305d74a5c088dc71ba588c72b2d2267

Observation 4102a7aa-35ba-4dfa-8a29-3e9d8efcfe92 · outbound

This paper cites Data Interpreter: An LLM Agent For Data Science.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Data Interpreter: An LLM Agent For Data Science

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.878437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.878437Z digest=sha256:4ca3f008cf438e0549e2478907272f91b5aa2d22d2071254782951cd9445cafb

Observation 7c6ccbed-f64c-4d5b-8fd7-4062c9ac36e2 · outbound

This paper cites Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.887709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.887709Z digest=sha256:9e4c4c0625f6b219b3571b69196b905507c12e8dd5a6711809dfb131f36b8243

Observation c16fe9b7-a4e6-4618-80d9-5f3d3eb0c748 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees LoRA: Low-Rank Adaptation of Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.891857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.891857Z digest=sha256:c383e060011cf07b41dde02369bbad70d2d8aff3b20160e3dbc0e1c0ee88a96c

Observation 080b3edf-e083-4039-befa-b6418483dec4 · outbound

This paper cites Jayaram Subramanya, D.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Jayaram Subramanya, D

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:10:26.525397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:22.896278Z digest=sha256:daa018fcf34becc9f68b68707c796c55356d4d20a33a6457d2f71ac3eb0f7fea

Observation d1d0193d-511a-475f-be6f-ed6b8ca82988 · outbound

This paper cites NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.904998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.904998Z digest=sha256:79258fefb794a55fdc3ec66c97703d3c71dca59192bf9fdd053433a61e1bbd15

Observation 8d8bef64-1fbb-4748-884f-0e9ed2d915f8 · outbound

This paper cites A System for Microserving of LLMs.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees A System for Microserving of LLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.910074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.910074Z digest=sha256:b665f1eea75f8f7f93495601d6215119ffdf0907d0d776f2d752bab1dc684e98

Observation 98edf3ec-67d3-4150-bca6-d89a37b78ddb · outbound

This paper cites P/D-Serve: Serving Disaggregated Large Language Model at Scale.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees P/D-Serve: Serving Disaggregated Large Language Model at Scale

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.915120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.915120Z digest=sha256:a086bfb9d42830a26b3028620ecf9fe9acc581f0ffd8e802b66d872c1bb680ea

Observation 71ab360a-47e1-4829-a2cc-787f9493e12a · outbound

This paper cites Kasner and O.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Kasner and O

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.920018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.920018Z digest=sha256:72f4278a9fe542b782e685eda0d9d5fff1d52ef3173fae58179b95a988d66880

Observation 551ac547-c551-4136-985d-96db252a50a6 · outbound

This paper cites TransLLaMa: LLM-based Simultaneous Translation System.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees TransLLaMa: LLM-based Simultaneous Translation System

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.924765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.924765Z digest=sha256:9aa6ea628bd1e93e011e0b6b3bf8db5398d1161212efcb5f4cdf8320b6991a2a

Observation b6999be8-0db9-4b49-9d4b-110ca6713b84 · outbound

This paper cites Koziolek, S.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Koziolek, S

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:10:26.511870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:22.929391Z digest=sha256:0a98d47151a89b3ac42468887f0be01b7414e851258e505bbcd48c78f59d23a8

Observation 8a8b374e-b10d-4e9a-9594-7daa28bf3f7d · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 42

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T10:10:25.018864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:22.939282Z digest=sha256:31dd029aec14fa54bc6b8caec211291cc9a6afd77566849403ead33929a151f7

Observation 843c7021-a90c-4e97-831c-8d9dcc4d6a57 · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-08T10:10:26.497975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:22.943986Z digest=sha256:80e27c0dbd5f261886b4ea55af1534b84515ae89a5dff3834a0e295cd7c993a9

Observation 3bd7c19c-e72d-450c-8ce4-93341d3e4ac1 · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 44

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T10:10:24.618588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:22.953648Z digest=sha256:44fb13cbcdd9c94cf80995011cb673860c2ed9fdb3f9dc787635fa9ffe55ae2d

Observation 1c8af5f9-0bca-478e-a65f-338ecd210b1d · outbound

This paper cites Li and Y.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Li and Y

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.958296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.958296Z digest=sha256:6b8d755065f79df9d3a54885d2c283a093b48b85bbb9976f611750d5b475b138

Observation 00cfd601-3e0b-45c4-bde4-0a203f62de22 · outbound

This paper cites ISBN 9798400705793.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees ISBN 9798400705793

Reference 46

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T10:10:24.774126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:22.934079Z digest=sha256:806c81343343fb94f4d647c8fcffd7d3290fde57b6af35fa91887d6ece738e15

Observation 67037a46-e217-4161-9602-77ce5d054a37 · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-08T10:10:26.451086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:22.969690Z digest=sha256:0b7937657945a7bbafbc452ebb7080b11cfb770efb14e864c094680cf8414292

Observation 78c06f67-3fea-4181-9542-0d3614066805 · outbound

This paper cites JarviX: A LLM No code Platform for Tabular Data Analysis and Optimization.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees JarviX: A LLM No code Platform for Tabular Data Analysis and Optimization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.978430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.978430Z digest=sha256:7d6a69a13975bd0b2691c13a50718cd851ebc90796190614364d66a0fa91d606

Observation 933940f3-c2ba-45f7-842e-3de9c9c88baa · outbound

This paper cites ISBN 978-1-939133-40-3.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees ISBN 978-1-939133-40-3

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:10:26.482902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:22.948862Z digest=sha256:1cff5ae70bbf276f3280ee3e473ca950fc7787c4f2b7d844ec2533aa67118beb

Observation 3ee6f993-8a04-4e54-b4ac-517074a9be2b · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.989448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.989448Z digest=sha256:027d68c135168e3641fe97f7bb38b00d77c5cffe03f6e845972181d0dc845870

Observation 1869ca4a-04c4-436d-80ab-e9a07dd08f58 · outbound

This paper cites Patel, E.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Patel, E

Reference 51

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T10:10:24.356569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:22.994847Z digest=sha256:3516b12e1cd6d7941cbe9cdf3c68f306518c9e368b5abe8217b7775a3be54ee6

Observation 23f0e975-cf2d-4376-aefd-7619d2c91c38 · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-08T10:10:26.467178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:22.964842Z digest=sha256:ff37406d123a7a143e801c085bcfa97dec15c96799e3e33df9803c6ec2119dce

Observation d3080e40-7188-4226-b8df-5d76f9baf677 · outbound

This paper cites Embedding-based Retrieval with LLM for Effective Agriculture Information Extracting from Unstructured Data.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Embedding-based Retrieval with LLM for Effective Agriculture Information Extracting from Unstructured Data

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:23.008155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:23.008155Z digest=sha256:a5ce1f9e2c8f56c4aae5c4edd1f6268e3429f52750008c0a5f85ce13560a10e2

Observation 2edf95f3-8250-47f2-b3ef-aa5241ad9e9d · outbound

This paper cites Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.973785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.973785Z digest=sha256:fb573041b4f739197d7d973beb1976a19d0af30334ddb2d5d82a060daad21f46

Observation 859140e1-5240-453b-bb61-92f8c698617f · outbound

This paper cites Radford, J.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Radford, J

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:10:26.403802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:23.018461Z digest=sha256:8afe84961e5238b3e2b186d6adfd19f907b54a775c180173e4bfc125c6aa297d

Observation 89689631-2df4-43fd-acbb-f8bd492b7d58 · outbound

This paper cites LLaMAX: Scaling Linguistic Horizons of LLM by Enhancing Translation Capabilities Beyond 100 Languages.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees LLaMAX: Scaling Linguistic Horizons of LLM by Enhancing Translation Capabilities Beyond 100 Languages

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.983923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.983923Z digest=sha256:eab36d0ef41ee567b4c4c384e56e377b755f3d1202ba3b3842087cae94ded886

Observation b5ab214a-be66-4b90-be9b-270b12727c8d · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 57

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T10:10:23.940852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:23.027812Z digest=sha256:20356a9340ca103d816410dbb2b9e1a05b732ea54d36f1e176b6fff5523a8ed5

Observation ee04131e-c1a4-4b82-a082-b917a0c239fc · outbound

This paper cites Sheng, L.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Sheng, L

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:10:26.372442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:23.031820Z digest=sha256:3fc027695b48992524b127b2d1aaf691c3a11a85b7fe2595e6abffa3ae3f5925

Observation b1865967-9974-40fd-9d08-a0f2b497d8ef · outbound

This paper cites Patke, D.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Patke, D

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:10:26.434280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:22.999634Z digest=sha256:1e96fa8675a052de930189a3a37d9df9724f0c7c45d3109d58573aee259e6b24

Observation 9dd51793-6362-43ee-bb19-d1a6b6bfa105 · outbound

This paper cites ISBN 9798400712869.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees ISBN 9798400712869

Reference 60

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T10:10:24.193122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:23.003701Z digest=sha256:bdeaa68a5639508a16d922b821e8e69c6c8856bfb58865f3338b2736b17af8cb

Observation 662abd6e-5878-4a70-b9fb-c2d4c8dd6cf4 · outbound

This paper cites D\'ej\`aVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees D\'ej\`aVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:23.045273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:23.045273Z digest=sha256:08b67466072c8da06a5b839ab17d62443900f55405cb3b22dcac7f32069885ad

Observation a99035e2-d60e-4602-a506-c98c8c941309 · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-08T10:10:26.419027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:23.013101Z digest=sha256:86f140c71b74ed9d16028bb2d9e872d1ed7d6f615f37b7ba594d5daeef3022f4

Observation 39c68a13-1fb8-4926-a96e-7bd0d068d316 · outbound

This paper cites Taori, I.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Taori, I

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:10:26.329163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:23.055114Z digest=sha256:448083ac46c4ceb0f0b1057ffba9ed452f0ccd6252932345ef2bf4636fb633b0

Observation 94410835-311f-4f22-8d66-a29178cc0ccd · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-08T10:10:26.388049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:23.023103Z digest=sha256:efd03b22d558b24ebae2a3c5559c0cb1b165d75dc0b9c632cfe7f19560b01fe2

Observation 57493d46-6d0c-4de4-b325-1d93336b9913 · outbound

This paper cites SynCode: LLM Generation with Grammar Augmentation.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees SynCode: LLM Generation with Grammar Augmentation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:23.065020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:23.065020Z digest=sha256:80357da1feb69c4e223f8b58ace9a7c32ad0e417b055be06a0785b1decf88464

Observation 7c171c9a-a435-4661-abb7-c565c78cbb53 · outbound

This paper cites Attention Is All You Need.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Attention Is All You Need

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:23.069757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:23.069757Z digest=sha256:b823b526fdc10a26c39a178fa6f4519285b2d01ce88891ee7b4fb42c378ffb04

Observation e7d609aa-25e5-40d3-9b0b-0a22d781f7de · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 67

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T10:10:23.745609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:23.036534Z digest=sha256:bd68441215d61d5e48c0cd5098f7f8081c44e586422bb92a3870b7984b563478

Observation 9c88a061-ac65-4cc4-9197-81ade5109dc1 · outbound

This paper cites Sivakumar.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Sivakumar

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:10:26.358754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:23.040912Z digest=sha256:5a94326eab253a5e8dfc6aea04d4f76270756bf11cd1e85649b3dd6b26399110

Observation 04c34b7b-ae75-4b9c-b79b-59d040fa0e2e · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Efficient Streaming Language Models with Attention Sinks

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:23.089056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:23.089056Z digest=sha256:5b7e60cc19fc0a71f259199363a0c148011bb7ad09f66ed3c6a1009259ae2741

Observation 74f6e1bd-4e2e-472b-ad3c-45f13b31fea5 · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-08T10:10:26.345078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:23.050027Z digest=sha256:03effdc551b7e77338921dfb6cf73021939a0ffde616697ecd551ed47e61be5e

Observation 702c9932-288c-49ce-8502-41fabbff8863 · outbound

This paper cites LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:23.098928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:23.098928Z digest=sha256:7c36e3f4edf46303ef5aaa9a5615c53dc7e5734607451df028fb7d064e1687e4

Observation 7ef46cc2-c744-4699-ab3a-3650b3babf5b · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:23.060164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:23.060164Z digest=sha256:be33b43d7614786bfdcb8445d40c88f112aa77f15920045af8f7d65dae939012

Observation 73c879ee-4105-496f-9fea-418b7edf8bb9 · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-08T10:10:26.297051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:23.108846Z digest=sha256:7aaf75f2a59581d4809b9807a918a629d43710b07244c764ffaf8b0fd88bbff9

Observation 6ed6c039-3b37-495c-8cbe-b021ae7d85a0 · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:23.113427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:23.113427Z digest=sha256:0c506c7656d81d9c35116259a1de3327409e678f7e19bb66e6ff6d81edee8c4f

Observation 85a06af6-94c4-4c7d-8068-6b465a061033 · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:23.074632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:23.074632Z digest=sha256:269c1c7204450014c48a99f7277c98af37d904d036fe0d7afc2a6c24ae78307d

Observation 31e23aa1-de2f-45ca-a08b-7f8daffda8d9 · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-08T10:10:26.312598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:23.079314Z digest=sha256:4769032556717205cd83474bc88b5f4633ffb1f73c66a3044323521c22844116

Observation d9e39e33-03ee-45d4-9459-ef5051e604bf · outbound

This paper cites Fast Distributed Inference Serving for Large Language Models.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Fast Distributed Inference Serving for Large Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:23.083720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:23.083720Z digest=sha256:07ae9f95380d6b5752c83798c324c4675685a67e5964fa3f794e026d47a8b48d

Observation 1fb437f5-8eb6-4b00-9e21-9ba31727476f · outbound

This paper cites Zhong, S.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Zhong, S

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:10:26.266298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:23.130873Z digest=sha256:d0049bbe4fe19ab098173f0551994a4d01256195c2ce35ee0aab1834aad428f6

Observation 7d24709d-5531-4835-a6a8-ae200dff1cbb · outbound

This paper cites Enhancing LLM with Evolutionary Fine Tuning for News Summary Generation.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Enhancing LLM with Evolutionary Fine Tuning for News Summary Generation

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:23.094126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:23.094126Z digest=sha256:94e552915daf2ea80af24e57bf67872d6a55cd032d8a8e669d0a905b3d5da4f8

Observation 49859a21-b5f5-4422-afb2-cbf81ec3b713 · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 80

Resolution
verified exact
doi, observed 2026-08-08T10:10:23.181192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:23.140836Z digest=sha256:a19147f5b5854c7f8bd142725f88d18a63cd940be6c17b0cf40a358ebfd3fe7a

Observation e0698959-7015-43d3-ab8a-17a0f7d293ca · outbound

This paper cites A Comparative Study of Offline Models and Online LLMs in Fake News Detection.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees A Comparative Study of Offline Models and Online LLMs in Fake News Detection

Reference 81

Resolution
verified exact
local_arxiv, observed 2026-08-08T10:10:23.316805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:23.103915Z digest=sha256:f1e827c071ae6e5209f3e5a06f787c5ef3d385209bdda90013ec3545bbd9d9a7

Observation 684d01be-6bba-4e3a-b5c0-df6027b1f96b · outbound

This paper cites Zhang, Y.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Zhang, Y

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:10:26.282165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:23.118218Z digest=sha256:dd378ca881cf4326bf811856ce38b166725c2050941cc104fae9a04383ae3f80

Observation d4a23382-f0b9-483a-aa37-fb196362d0f2 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees OPT: Open Pre-trained Transformer Language Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:23.122052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:23.122052Z digest=sha256:e385f567139fdeaf158c3c9acbe132ec5c416fe73c2eb30411655782521ecf25

Observation c1d5ee91-2e11-40cb-943f-04393324bc8c · outbound

This paper cites MEMO: Fine-grained Tensor Management For Ultra-long Context LLM Training.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees MEMO: Fine-grained Tensor Management For Ultra-long Context LLM Training

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-08-08T10:10:23.264959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:23.126229Z digest=sha256:215901f04571553fb466f6a974c80a88e13a60bfab64c4f85dc16cfb67a05d35

Observation 78ff80cd-7750-4a06-bf4b-db00c4e2132a · outbound

This paper cites LLM-Enhanced Data Management.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees LLM-Enhanced Data Management

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:23.135618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:23.135618Z digest=sha256:efff273bcb6e03362fdcf58b4c3a0a6266215c56ea74d5c4c4b298c2d268da00

Observation 9c7f6cc5-0cdd-4cf1-83c4-8069fea15ff8 · outbound

This paper cites ISBN 9781450381376.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees ISBN 9781450381376

Reference 2020

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T10:10:25.541451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:22.799477Z digest=sha256:b9d3054f13b2dfa7c5f5bfc564216f900c573ca2667547fd98f42a18d472cdf9

Observation 863f88bd-2095-4e28-8bc3-5619ec192cf1 · outbound

This paper cites doi: https://doi.org/10.1016/j.csl.2021.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees doi: https://doi.org/10.1016/j.csl.2021

Reference 2022

Resolution
verified exact
doi, observed 2026-08-08T10:10:23.228323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:22.748456Z digest=sha256:5a84a2c10324557dbf82a4822ad8517ee1303de228d83937446c69ec970240ef

Observation c827061c-605c-455c-b76a-ab5a0c378816 · outbound

This paper cites Practical offloading for fine-tuning LLM on commodity GPU via learned sparse projectors.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Practical offloading for fine-tuning LLM on commodity GPU via learned sparse projectors

Reference 2024

Resolution
verified exact
local_arxiv, observed 2026-08-08T10:10:25.951139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T10:10:22.774052Z digest=sha256:27ec5518c118ee4d0beb87133cbbf2eb4c1319edb2ba8bafbf0e058b40a4cc44

Pith citing papers

Observation 0f7bd89d-1367-495f-859a-f4031d98415d · inbound

SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips cites this paper.

SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips Memory Offloading for Large Language Model Inference with Latency SLO Guarantees

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:30:17.959799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T15:26:01.283448Z digest=sha256:096055eb1f615307fddd535d721bb66930801c38d4313605b24028020353b426