Pith. sign in

Paper Citation Record · LEDGER

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees

As of 9 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 1 inbound Pith citation observation for arXiv:2502.08182.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.08182 v1

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T10:10:23.140836Z

measured 88 of 88 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-21T15:26:01.283448Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T15:30:17.957796Z

Reference resolution

87 of 87 outbound references displayed

  • verified exact18
  • verified fuzzy14
  • unresolved55
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dfd3e3b5-9f1e-48c8-bd5b-a897ed731d4c · outbound

This paper cites URL https://huggingface.co/ docs/accelerate/index.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees URL https://huggingface.co/ docs/accelerate/index

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.710232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.710232Z digest=sha256:11d5ad2bc5b59432a53f83153511537c1dbd19079eb39ac721cc8ce33f347ed4

Observation ea0366ba-6bb8-4a30-9625-dce2bdc4c27f · outbound

This paper cites URL https://sharegpt.com/.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees URL https://sharegpt.com/

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.716168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.716168Z digest=sha256:16c43008be538ad601aabfb18b6d9ce70d0b8bcaef4b7aade9dced0b49ac5eb9

Observation c809b2a0-2bd0-4602-9c5b-dbf0129b93cd · outbound

This paper cites Acharya, B.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Acharya, B

Reference 3

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T10:10:26.250091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:22.721928Z digest=sha256:bbacf372ff062e2a4bd2368a9e1297a6533c00ae25915fe415d4b32747230b24

Observation aea5f611-4591-4f01-8d04-77904f1a748f · outbound

This paper cites SYMPHONY: Improving Memory Management for LLM Inference Workloads.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees SYMPHONY: Improving Memory Management for LLM Inference Workloads

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-08T10:10:26.078778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:22.727386Z digest=sha256:59898d860ea1c5d477717f361ee81e615c848d17c6378874b6c34dd6e111d466

Observation b16632af-cffe-4664-a125-b340c71fcf6f · outbound

This paper cites SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.732763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.732763Z digest=sha256:07bcba37ce8bc4b5b9385eb6d118d2ced8d16855453dd792315eb168e9c6fa9b

Observation e73bab42-7498-48f5-bf9d-7c33882b4062 · outbound

This paper cites LLM in a flash: Efficient Large Language Model Inference with Limited Memory.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.738166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.738166Z digest=sha256:486e15a5cf7b7e98013ed3d104269ff5f0d1f9ef66de3eb65d9f51086f3327e2

Observation cf37bd70-da3b-4841-97b7-fd4028c59005 · outbound

This paper cites Alomari, N.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Alomari, N

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.743842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.743842Z digest=sha256:99785a08610e6322cca5b9a35dbecb45ddc73b88d30b90053ad8aef2bb6489f9

Observation 64d1de85-13b0-46c8-b1a0-474069fc9f6b · outbound

This paper cites Neutrino Production via $e^-e^+$ Collision at $Z$-boson Peak.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Neutrino Production via $e^-e^+$ Collision at $Z$-boson Peak

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.753134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.753134Z digest=sha256:2c6dbd80e4bc697df49144b89a229a3533bf2c2b15a13c34106ba38e6e581403

Observation bd779b54-fef4-4499-b02c-d0d4df4e3041 · outbound

This paper cites Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.757378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.757378Z digest=sha256:af899523f40379340c0280fecc5227e30f78f49e1b01b4b8197df65314a04714

Observation 99c060a7-b00f-4a3d-be13-bf570ed796c5 · outbound

This paper cites AWESOME: GPU Memory-constrained Long Document Summarization using Memory Mechanism and Global Salient Content.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees AWESOME: GPU Memory-constrained Long Document Summarization using Memory Mechanism and Global Salient Content

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-08T10:10:25.989475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:22.761709Z digest=sha256:76fd6e667d4aba8ae55237610cca85fc2ff740aa34599f5c5a6e90bd69afc9a5

Observation 2e3514d0-fad9-46f4-b93b-1a6e6adea279 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Evaluating Large Language Models Trained on Code

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.765947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.765947Z digest=sha256:8ac3e944c476a5b27bf255d2bde24413e676045befc8fee89a1e2bc9dcfda6cf

Observation 31fa23cf-a362-426c-bd5e-0524c2d95eaa · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.770073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.770073Z digest=sha256:ce645cf520229c0ab56183afb7b03afc26ed8f8db04971336efb84b25dcb1df2

Observation 3ec6d3e6-3e3c-4a2c-8d75-392c61dbe549 · outbound

This paper cites Cheng, Y.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Cheng, Y

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.779090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.779090Z digest=sha256:35d54ade662b56eea97ef98aac37e8290786a7cd016001c18120f16a10d92355

Observation 361882a4-e08b-49c1-b7f5-b9e10f42196d · outbound

This paper cites Choquette, W.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Choquette, W

Reference 14

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T10:10:25.739015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:22.783485Z digest=sha256:e393b8a56136c7ba2b86c6dd8e9a957084684a1356a87a8764a4e37f70df34bf

Observation a217b87c-3eb6-4156-8223-f9e31b9318d4 · outbound

This paper cites Crankshaw, X.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Crankshaw, X

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.788626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.788626Z digest=sha256:20c1a25ce98cae6a9cf128d00a91c14c0231409801bde9ad69e2a4a38c3e7b6d

Observation 0eb26b63-d62b-433a-b9dd-49bd525b9a5e · outbound

This paper cites Crankshaw, G.-E.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Crankshaw, G.-E

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.794421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.794421Z digest=sha256:a52e66b08ed92afc3aaca57b8ba4d8aacd8bd8dee21411e6f45790e71a921abd

Observation 8864959f-64f8-4c43-8edb-a01c5363656d · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.804664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.804664Z digest=sha256:7d409ecafe041d4fd8812942bf96480fdc1aaa70189b592c90155a72a64fc82d

Observation 2d267163-9e1d-4904-b99b-ae00d8b36c70 · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-08T10:10:26.626313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:22.809216Z digest=sha256:6dd3ee1ebfd9f49166ab3b40633363538f2c5b978b77e388dbfe97c10c32f15d

Observation c055ce8a-9ab6-43ec-9bda-0a119794ef09 · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-08T10:10:26.612178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:22.815367Z digest=sha256:5ada808c07cf097c12328c45a69cbe71a913a2e1ef31768ec70c92b867b41c62

Observation 868c2bb5-b7aa-4d78-b8f9-516d52af42b3 · outbound

This paper cites Improving LLM Abilities in Idiomatic Translation.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Improving LLM Abilities in Idiomatic Translation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.822360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.822360Z digest=sha256:1edfdbdbe0fd77d60f3eadca0c1188fdb3096c1ef424fab1dff297da709479dd

Observation b7d1cb3d-82fc-453c-be89-0574bc0e903d · outbound

This paper cites Efficient Training of Large Language Models on Distributed Infrastructures: A Survey.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Efficient Training of Large Language Models on Distributed Infrastructures: A Survey

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.827490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.827490Z digest=sha256:4d8cb97893dca44f084a681bf42165f8f7d39f11b270d68eadd409d56d412ef1

Observation ad2081ed-afb5-4bab-9329-e09f717e306c · outbound

This paper cites Elliott, M.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Elliott, M

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:10:26.598091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:22.832686Z digest=sha256:99f4c81fd73df26a97fc541ab72eb0f9b685b3321a008fda1b6c68f79d4de8b7

Observation f513ccb5-2799-4452-84e7-1bf801a016da · outbound

This paper cites Gambhir and V.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Gambhir and V

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:10:26.584171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:22.837711Z digest=sha256:dedafb72ac0cf7968b7b14d1f8861562a6a401d9bc610d4e12446ea972618a3c

Observation aab287a1-5331-43ed-9816-6b4d6f824e8a · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-08T10:10:26.569988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:22.842527Z digest=sha256:ca8c33a68099dc6417e1dbe910b7617ac04ca3be7a79d8b3185825a3d4bc704a

Observation 05f2565b-fce5-4450-aa48-bc54c5e0504e · outbound

This paper cites Fast State Restoration in LLM Serving with HCache.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Fast State Restoration in LLM Serving with HCache

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.848211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.848211Z digest=sha256:54c542078fb50463012e7b298b19d1140da053483b8b6ed54e8a71c721e8bb9b

Observation 0875813d-2c66-4ca6-bea5-74fc4b222a31 · outbound

This paper cites M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.853443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.853443Z digest=sha256:3965657ad418612df1d1e6899ad28ff3df33c0cd4edab27a3eca8ac31595c4c9

Observation 079c3fe8-7d92-4e73-9b8f-29ff9a898f21 · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 27

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T10:10:25.293444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:22.858724Z digest=sha256:416ed797988e2be5ba95224f73ceeaf360fccaabbac36dee899037261ab398bf

Observation 3e9542ac-3d38-458a-9453-dbf985ba63f4 · outbound

This paper cites Gujarati, R.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Gujarati, R

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:10:26.554673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:22.863209Z digest=sha256:cb1a4938694b41b20ae363c3654fd71a1f025eb27c3e3d5133fa0b6eff9df12b

Observation 6c2f1dc1-ed41-4cce-8198-907b1da0d872 · outbound

This paper cites Gujarati, R.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Gujarati, R

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:10:26.540571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:22.867754Z digest=sha256:c6d7d9caebeff8c092d051e7fdffd89b3492f51bc4e21d24737eddd2c0e68fa1

Observation 6bedb766-c3d8-4009-a6f4-8cabf5d05c7c · outbound

This paper cites DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.872509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.872509Z digest=sha256:1fae67ba280544561033f8dfb36d96d97305d74a5c088dc71ba588c72b2d2267

Observation 4102a7aa-35ba-4dfa-8a29-3e9d8efcfe92 · outbound

This paper cites Data Interpreter: An LLM Agent For Data Science.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Data Interpreter: An LLM Agent For Data Science

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.878437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.878437Z digest=sha256:09514f06e748b70158b318e081d66258bc741a00f6ee818a0aab23b0265fa96d

Observation 7c6ccbed-f64c-4d5b-8fd7-4062c9ac36e2 · outbound

This paper cites Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.887709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.887709Z digest=sha256:9e4c4c0625f6b219b3571b69196b905507c12e8dd5a6711809dfb131f36b8243

Observation c16fe9b7-a4e6-4618-80d9-5f3d3eb0c748 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees LoRA: Low-Rank Adaptation of Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.891857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.891857Z digest=sha256:c383e060011cf07b41dde02369bbad70d2d8aff3b20160e3dbc0e1c0ee88a96c

Observation 080b3edf-e083-4039-befa-b6418483dec4 · outbound

This paper cites Jayaram Subramanya, D.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Jayaram Subramanya, D

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:10:26.525397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:22.896278Z digest=sha256:a20d57a764521bf76a61c0c063a58650781b895d5d5038688480ccd735735e5a

Observation d1d0193d-511a-475f-be6f-ed6b8ca82988 · outbound

This paper cites NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.904998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.904998Z digest=sha256:79258fefb794a55fdc3ec66c97703d3c71dca59192bf9fdd053433a61e1bbd15

Observation 8d8bef64-1fbb-4748-884f-0e9ed2d915f8 · outbound

This paper cites A System for Microserving of LLMs.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees A System for Microserving of LLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.910074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.910074Z digest=sha256:b665f1eea75f8f7f93495601d6215119ffdf0907d0d776f2d752bab1dc684e98

Observation 98edf3ec-67d3-4150-bca6-d89a37b78ddb · outbound

This paper cites P/D-Serve: Serving Disaggregated Large Language Model at Scale.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees P/D-Serve: Serving Disaggregated Large Language Model at Scale

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.915120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.915120Z digest=sha256:a086bfb9d42830a26b3028620ecf9fe9acc581f0ffd8e802b66d872c1bb680ea

Observation 71ab360a-47e1-4829-a2cc-787f9493e12a · outbound

This paper cites Kasner and O.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Kasner and O

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.920018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.920018Z digest=sha256:72f4278a9fe542b782e685eda0d9d5fff1d52ef3173fae58179b95a988d66880

Observation 551ac547-c551-4136-985d-96db252a50a6 · outbound

This paper cites TransLLaMa: LLM-based Simultaneous Translation System.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees TransLLaMa: LLM-based Simultaneous Translation System

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.924765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.924765Z digest=sha256:9aa6ea628bd1e93e011e0b6b3bf8db5398d1161212efcb5f4cdf8320b6991a2a

Observation b6999be8-0db9-4b49-9d4b-110ca6713b84 · outbound

This paper cites Koziolek, S.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Koziolek, S

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:10:26.511870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:22.929391Z digest=sha256:a5eaaed930bb89f6f3c1dcb9d0edbd44523acbe26bef0bdd7e3651ce67977ef2

Observation 8a8b374e-b10d-4e9a-9594-7daa28bf3f7d · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 42

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T10:10:25.018864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:22.939282Z digest=sha256:93040063ee6060c58ab08ad3498600f7d9c1d6b83e2d372684a2dcac9bdefb56

Observation 843c7021-a90c-4e97-831c-8d9dcc4d6a57 · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-08T10:10:26.497975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:22.943986Z digest=sha256:0169a7b372db809e1ebd55f2dd0cd1202e838a1f3eff4c4ec63cbff42ffe5fec

Observation 3bd7c19c-e72d-450c-8ce4-93341d3e4ac1 · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 44

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T10:10:24.618588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:22.953648Z digest=sha256:d19b9d6412ff6cbb6195188fc64e327024323b2cca16f1571147290c8835ad31

Observation 1c8af5f9-0bca-478e-a65f-338ecd210b1d · outbound

This paper cites Li and Y.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Li and Y

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.958296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.958296Z digest=sha256:6b8d755065f79df9d3a54885d2c283a093b48b85bbb9976f611750d5b475b138

Observation 00cfd601-3e0b-45c4-bde4-0a203f62de22 · outbound

This paper cites ISBN 9798400705793.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees ISBN 9798400705793

Reference 46

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T10:10:24.774126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:22.934079Z digest=sha256:9edc07d065b7f6bb6ea9d50f727d271b8671432c83506437ebb5022ab3032392

Observation 67037a46-e217-4161-9602-77ce5d054a37 · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-08T10:10:26.451086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:22.969690Z digest=sha256:306634dfa8a4e990597919f3db3ffa222a306e9e13e4be53617fd7c7c1a52517

Observation 78c06f67-3fea-4181-9542-0d3614066805 · outbound

This paper cites JarviX: A LLM No code Platform for Tabular Data Analysis and Optimization.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees JarviX: A LLM No code Platform for Tabular Data Analysis and Optimization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.978430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.978430Z digest=sha256:7d6a69a13975bd0b2691c13a50718cd851ebc90796190614364d66a0fa91d606

Observation 933940f3-c2ba-45f7-842e-3de9c9c88baa · outbound

This paper cites ISBN 978-1-939133-40-3.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees ISBN 978-1-939133-40-3

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:10:26.482902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:22.948862Z digest=sha256:f7f563a4d25c62fa961b0788e6c86672142403f9e811558a75595c794ac41fed

Observation 3ee6f993-8a04-4e54-b4ac-517074a9be2b · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.989448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.989448Z digest=sha256:027d68c135168e3641fe97f7bb38b00d77c5cffe03f6e845972181d0dc845870

Observation 1869ca4a-04c4-436d-80ab-e9a07dd08f58 · outbound

This paper cites Patel, E.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Patel, E

Reference 51

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T10:10:24.356569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:22.994847Z digest=sha256:c591db9d538f68995ffd36922691a10bd8f6e5d6b5a1c1176336967785f34ce4

Observation 23f0e975-cf2d-4376-aefd-7619d2c91c38 · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-08T10:10:26.467178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:22.964842Z digest=sha256:cb3d48c52e6309ef26dcd85f17903cd969f4dcf49d4ffca84268cb29f573c2ff

Observation d3080e40-7188-4226-b8df-5d76f9baf677 · outbound

This paper cites Embedding-based Retrieval with LLM for Effective Agriculture Information Extracting from Unstructured Data.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Embedding-based Retrieval with LLM for Effective Agriculture Information Extracting from Unstructured Data

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:23.008155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:23.008155Z digest=sha256:a5ce1f9e2c8f56c4aae5c4edd1f6268e3429f52750008c0a5f85ce13560a10e2

Observation 2edf95f3-8250-47f2-b3ef-aa5241ad9e9d · outbound

This paper cites Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.973785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.973785Z digest=sha256:fb573041b4f739197d7d973beb1976a19d0af30334ddb2d5d82a060daad21f46

Observation 859140e1-5240-453b-bb61-92f8c698617f · outbound

This paper cites Radford, J.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Radford, J

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:10:26.403802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:23.018461Z digest=sha256:dda7eba922c4ae67294944b6c17e992aeaf0d5e1811594c8a58f5f66625314c3

Observation 89689631-2df4-43fd-acbb-f8bd492b7d58 · outbound

This paper cites LLaMAX: Scaling Linguistic Horizons of LLM by Enhancing Translation Capabilities Beyond 100 Languages.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees LLaMAX: Scaling Linguistic Horizons of LLM by Enhancing Translation Capabilities Beyond 100 Languages

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.983923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.983923Z digest=sha256:def05ce5cb6363f3f27d0631b09cf2b7216bd6bc892adc0e9c72bf46652dea2a

Observation b5ab214a-be66-4b90-be9b-270b12727c8d · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 57

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T10:10:23.940852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:23.027812Z digest=sha256:0c96a3c77e16f751290896557455710e02aeceba886bf7cdaa18fb60ce7b3710

Observation ee04131e-c1a4-4b82-a082-b917a0c239fc · outbound

This paper cites Sheng, L.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Sheng, L

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:10:26.372442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:23.031820Z digest=sha256:bff6ca0703f6375ac01f8b61b4b47722d5a4b25389883939d94001561ac406de

Observation b1865967-9974-40fd-9d08-a0f2b497d8ef · outbound

This paper cites Patke, D.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Patke, D

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:10:26.434280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:22.999634Z digest=sha256:9d9fcd6f770e8e8a71a1b6a08c5c6d64c118f5986e75a9aa3d2bbc62ecac3f88

Observation 9dd51793-6362-43ee-bb19-d1a6b6bfa105 · outbound

This paper cites ISBN 9798400712869.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees ISBN 9798400712869

Reference 60

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T10:10:24.193122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:23.003701Z digest=sha256:12b542d749cf9ac1b98185da3b35851f9a369cf5beda8bb04809fbaf1035f2bd

Observation 662abd6e-5878-4a70-b9fb-c2d4c8dd6cf4 · outbound

This paper cites D\'ej\`aVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees D\'ej\`aVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:23.045273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:23.045273Z digest=sha256:08b67466072c8da06a5b839ab17d62443900f55405cb3b22dcac7f32069885ad

Observation a99035e2-d60e-4602-a506-c98c8c941309 · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-08T10:10:26.419027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:23.013101Z digest=sha256:d2eeef38b43c1ee740e6b4c66d8d1c491de7523003a8822dd148c1fa13ec4659

Observation 39c68a13-1fb8-4926-a96e-7bd0d068d316 · outbound

This paper cites Taori, I.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Taori, I

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:10:26.329163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:23.055114Z digest=sha256:c9a6e619eb866bb0ea4efb5eb15cba231040d4ff64ec87bc73e7100b12045ae0

Observation 94410835-311f-4f22-8d66-a29178cc0ccd · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-08T10:10:26.388049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:23.023103Z digest=sha256:e4340cd01c1853d38d174f0381aa70b43cd8d7f042263e4abdcc067a75b5409c

Observation 57493d46-6d0c-4de4-b325-1d93336b9913 · outbound

This paper cites SynCode: LLM Generation with Grammar Augmentation.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees SynCode: LLM Generation with Grammar Augmentation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:23.065020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:23.065020Z digest=sha256:80357da1feb69c4e223f8b58ace9a7c32ad0e417b055be06a0785b1decf88464

Observation 7c171c9a-a435-4661-abb7-c565c78cbb53 · outbound

This paper cites Attention Is All You Need.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Attention Is All You Need

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:23.069757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:23.069757Z digest=sha256:b823b526fdc10a26c39a178fa6f4519285b2d01ce88891ee7b4fb42c378ffb04

Observation e7d609aa-25e5-40d3-9b0b-0a22d781f7de · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 67

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T10:10:23.745609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:23.036534Z digest=sha256:13dc028972c4a93e65e299814a5a4ee067167485355334edb538f76c4514444b

Observation 9c88a061-ac65-4cc4-9197-81ade5109dc1 · outbound

This paper cites Sivakumar.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Sivakumar

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:10:26.358754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:23.040912Z digest=sha256:3d869541fdc3fde915aa7a1a462460a8f7628ec9092160fca549f2de75b5a86c

Observation 04c34b7b-ae75-4b9c-b79b-59d040fa0e2e · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Efficient Streaming Language Models with Attention Sinks

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:23.089056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:23.089056Z digest=sha256:5b7e60cc19fc0a71f259199363a0c148011bb7ad09f66ed3c6a1009259ae2741

Observation 74f6e1bd-4e2e-472b-ad3c-45f13b31fea5 · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-08T10:10:26.345078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:23.050027Z digest=sha256:efcf7033ac55fae5df22f1207a864e2d0e58dea97a37b3533c162a3d4c5f5421

Observation 702c9932-288c-49ce-8502-41fabbff8863 · outbound

This paper cites LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:23.098928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:23.098928Z digest=sha256:71af0882eb02073c9241d580b112ae746b441863f5543734279575bcc17ad7c2

Observation 7ef46cc2-c744-4699-ab3a-3650b3babf5b · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:23.060164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:23.060164Z digest=sha256:be33b43d7614786bfdcb8445d40c88f112aa77f15920045af8f7d65dae939012

Observation 73c879ee-4105-496f-9fea-418b7edf8bb9 · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-08T10:10:26.297051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:23.108846Z digest=sha256:79266bf6605e566b5d02a902bac92c2fbe4bef30d8046d55f2ad39b72b1d6068

Observation 6ed6c039-3b37-495c-8cbe-b021ae7d85a0 · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:23.113427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:23.113427Z digest=sha256:0c506c7656d81d9c35116259a1de3327409e678f7e19bb66e6ff6d81edee8c4f

Observation 85a06af6-94c4-4c7d-8068-6b465a061033 · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:23.074632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:23.074632Z digest=sha256:269c1c7204450014c48a99f7277c98af37d904d036fe0d7afc2a6c24ae78307d

Observation 31e23aa1-de2f-45ca-a08b-7f8daffda8d9 · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-08T10:10:26.312598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:23.079314Z digest=sha256:755140aa78bf68798330fcc09f0dfbc78a3436312d195ea2b15947d30bda3365

Observation d9e39e33-03ee-45d4-9459-ef5051e604bf · outbound

This paper cites Fast Distributed Inference Serving for Large Language Models.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Fast Distributed Inference Serving for Large Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:23.083720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:23.083720Z digest=sha256:3deb3ebd7f6f12f2cccb755a2d040d11865443999ab2f2d497fc3a3bbe25c401

Observation 1fb437f5-8eb6-4b00-9e21-9ba31727476f · outbound

This paper cites Zhong, S.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Zhong, S

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:10:26.266298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:23.130873Z digest=sha256:c541c67f7655ed5579cde2f09e219c75b858c02ef281bbab1f474990fda95126

Observation 7d24709d-5531-4835-a6a8-ae200dff1cbb · outbound

This paper cites Enhancing LLM with Evolutionary Fine Tuning for News Summary Generation.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Enhancing LLM with Evolutionary Fine Tuning for News Summary Generation

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:23.094126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:23.094126Z digest=sha256:94e552915daf2ea80af24e57bf67872d6a55cd032d8a8e669d0a905b3d5da4f8

Observation 49859a21-b5f5-4422-afb2-cbf81ec3b713 · outbound

This paper cites an unresolved cited work.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work

Reference 80

Resolution
verified exact
doi, observed 2026-08-08T10:10:23.181192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:23.140836Z digest=sha256:8aa7a7d66c9a032f3490799e4db81029c1df22234b15275b322eb867988203c3

Observation e0698959-7015-43d3-ab8a-17a0f7d293ca · outbound

This paper cites A Comparative Study of Offline Models and Online LLMs in Fake News Detection.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees A Comparative Study of Offline Models and Online LLMs in Fake News Detection

Reference 81

Resolution
verified exact
local_arxiv, observed 2026-08-08T10:10:23.316805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:23.103915Z digest=sha256:0d28b4bb53b2300369dfafb3b9b549275aef008b1563cd486ea71e2f5102802f

Observation 684d01be-6bba-4e3a-b5c0-df6027b1f96b · outbound

This paper cites Zhang, Y.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Zhang, Y

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:10:26.282165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:23.118218Z digest=sha256:73c1a57add0ff4e6a9ab61d31032ed66201d03d550ae934ababc6041046d94b5

Observation d4a23382-f0b9-483a-aa37-fb196362d0f2 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees OPT: Open Pre-trained Transformer Language Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:23.122052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:23.122052Z digest=sha256:e385f567139fdeaf158c3c9acbe132ec5c416fe73c2eb30411655782521ecf25

Observation c1d5ee91-2e11-40cb-943f-04393324bc8c · outbound

This paper cites MEMO: Fine-grained Tensor Management For Ultra-long Context LLM Training.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees MEMO: Fine-grained Tensor Management For Ultra-long Context LLM Training

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-08-08T10:10:23.264959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:23.126229Z digest=sha256:95474e6796ee4c230a3a9c1c539aeba41e27d562b99cad461c945dec551ce0d8

Observation 78ff80cd-7750-4a06-bf4b-db00c4e2132a · outbound

This paper cites LLM-Enhanced Data Management.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees LLM-Enhanced Data Management

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:23.135618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:23.135618Z digest=sha256:c555c4be5053c54b7607ef4f61a7b83ce3495716cfd4d4dc8b641d7c81d8c146

Observation 9c7f6cc5-0cdd-4cf1-83c4-8069fea15ff8 · outbound

This paper cites ISBN 9781450381376.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees ISBN 9781450381376

Reference 2020

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T10:10:25.541451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:22.799477Z digest=sha256:4d5072086287cf67aa807afece2d91eba3e5183154e909c625aa21024103ccc0

Observation 863f88bd-2095-4e28-8bc3-5619ec192cf1 · outbound

This paper cites doi: https://doi.org/10.1016/j.csl.2021.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees doi: https://doi.org/10.1016/j.csl.2021

Reference 2022

Resolution
verified exact
doi, observed 2026-08-08T10:10:23.228323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:22.748456Z digest=sha256:563d2289dd66f98e376dbb1d73f78db978a0875f3c94cdf8173b15c1548380ab

Observation c827061c-605c-455c-b76a-ab5a0c378816 · outbound

This paper cites Practical offloading for fine-tuning LLM on commodity GPU via learned sparse projectors.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Practical offloading for fine-tuning LLM on commodity GPU via learned sparse projectors

Reference 2024

Resolution
verified exact
local_arxiv, observed 2026-08-08T10:10:25.951139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:10:22.774052Z digest=sha256:20fa77426ba23fcd634b1abbf2ce267aa5ccaeae83186d5631101c608ca7de17

Pith citing papers

Observation 0f7bd89d-1367-495f-859a-f4031d98415d · inbound

SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips cites this paper.

SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips Memory Offloading for Large Language Model Inference with Latency SLO Guarantees

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:30:17.959799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T15:26:01.283448Z digest=sha256:617521872028aca3dd78fa5fe90b550062e6c99c9e731df6d3f1f327585f63f2