Pith. sign in

Paper Citation Record · LEDGER

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees

As of 17 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2506.19677.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.19677 v2

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:34:39.461915Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact1
  • verified fuzzy28
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a7b85b78-d2bc-4062-9493-ee49a6281da0 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Code Llama: Open Foundation Models for Code

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.135939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.135939Z digest=sha256:ca2860fc84c3ea6c72b8e2f49a98e711f10b95adaf930fae7ae19fb02dffbf87

Observation 5ae8e1bb-65c9-45b7-9679-9d03881705af · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.142344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.142344Z digest=sha256:b9af47b88cea3cf953e4fe3a7373f6d7633ffbfff139ccf7a6d10c480822adf2

Observation fa6e4015-90ca-4285-8d57-536cdf6b889d · outbound

This paper cites Qwen2.5-Coder Technical Report.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Qwen2.5-Coder Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.148481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.148481Z digest=sha256:41898160085682c3f09ae9e41d31f1f0609ef1618dc7da10268bd8cace995d6b

Observation 2487f045-311f-4609-b73a-42cec3188f6a · outbound

This paper cites StarCoder 2 and The Stack v2: The Next Generation.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees StarCoder 2 and The Stack v2: The Next Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.154863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.154863Z digest=sha256:6ff252a716a9ac8983d0f0fdd003c29dc792c5b5ee66530a643768abf4232a62

Observation 0419d7b9-9f88-44a2-bc25-02f8c6df7a6b · outbound

This paper cites Fine tuning large language model for secure code generation,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Fine tuning large language model for secure code generation,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.836208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.160386Z digest=sha256:1b23e3362673b6ef5f0f7e69224bc52c195419bdc9b1f10bf167e74befd0a461

Observation 975034c9-d087-43a5-9f57-3da2366bb72c · outbound

This paper cites RepoHyper: Search-Expand-Refine on Semantic Graphs for Repository-Level Code Completion.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees RepoHyper: Search-Expand-Refine on Semantic Graphs for Repository-Level Code Completion

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.164978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.164978Z digest=sha256:5f32895df19d3cab85bff9b4e2474414e3b303fc72eb8f3fe5ac29fc512886ee

Observation d2e1e620-1d0c-4cde-94ba-99ea4809c32d · outbound

This paper cites A Survey on Large Language Models for Code Generation.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees A Survey on Large Language Models for Code Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.171573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.171573Z digest=sha256:f142a7bb09404803dc55c4900b1fc63e30d16b3bd779aafc3fbbca92b9d143f0

Observation 81fda2bd-3062-49c6-b4bd-e9f3b4f75e8f · outbound

This paper cites Language models for code completion: A practical evaluation,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Language models for code completion: A practical evaluation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.808922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.177142Z digest=sha256:e053cc3d7be99d7be75eaaa71b61996f998558644fc63f5fe9c021da1efd6684

Observation 984a41e3-9824-4d89-9e67-147a8cc2a164 · outbound

This paper cites AI-assisted Code Authoring at Scale: Fine-tuning, deploying, and mixed methods evaluation.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees AI-assisted Code Authoring at Scale: Fine-tuning, deploying, and mixed methods evaluation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.182174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.182174Z digest=sha256:ac188a574f38f2729d0787a6149520cbb1c6c63941726b4c5913354099b49030

Observation de5cda9c-beef-486e-8e3b-f51dd0b660cf · outbound

This paper cites A survey on large language model (llm) security and privacy: The good, the bad, and the ugly,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees A survey on large language model (llm) security and privacy: The good, the bad, and the ugly,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.188655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.188655Z digest=sha256:af5c4ec52704e8bddf686cda7c89853a6c067ca571eb0aab077cd737161fe69d

Observation fb1b5d8a-fc0d-4f20-baaa-3956624fb948 · outbound

This paper cites Security and privacy challenges of large language models: A survey,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Security and privacy challenges of large language models: A survey,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.194272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.194272Z digest=sha256:3d69f5cedc1e21ac58ba33fd1ae2578f81aff487c31016b3ed37c05f61e6114c

Observation e4245da9-307c-43c7-9d71-8c95c8245d53 · outbound

This paper cites Hierarchical Repository-Level Code Summarization for Business Applications Using Local LLMs.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Hierarchical Repository-Level Code Summarization for Business Applications Using Local LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.201458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.201458Z digest=sha256:cff7e7451e65fd45b1556971deec293fe94ae69de59f3a154479f30721a5fb7b

Observation a0de50b7-d19b-447a-82fa-61a647a202e0 · outbound

This paper cites Using ollama,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Using ollama,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.206736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.206736Z digest=sha256:b1390bea86937a67e6b67dea6443ab83b2cc856a158da60b0a880ae676a7260f

Observation a86f1378-8681-47e0-a439-8386e407f3f4 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Efficient memory management for large language model serving with pagedattention,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.212463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.212463Z digest=sha256:7010b8c4b08c33fd7c923e464dd2db9d560814d29813cc64797e536b96945a6c

Observation 6a3e9bb0-2cca-41c1-acce-1740743c9e2c · outbound

This paper cites Efficiently programming large language models using sglang.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Efficiently programming large language models using sglang

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.218200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.218200Z digest=sha256:6dca1b7a6737da6e9b379120b9b3f13e62d634fc2dbb0aad2b579ed1866d3e2a

Observation 88e3acd8-986c-4367-ae0a-968afae4c9ba · outbound

This paper cites an unresolved cited work.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:34:40.707301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.223034Z digest=sha256:de94ab3fdcb85cde75a397b6f3ae043e503b4fb7948bf676c3443a9cf125e0ba

Observation f9dfc8c3-0af7-4947-9eb9-76716a57586a · outbound

This paper cites Orca: A distributed serving system for{Transformer-Based}generative models,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Orca: A distributed serving system for{Transformer-Based}generative models,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.682085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.228640Z digest=sha256:1ad463c2924e85af424b0635b41750fb66ce88de81b5575673036ba8a78fc83f

Observation 6b06a26a-ab7e-4c13-8a75-f654254649d1 · outbound

This paper cites Llumnix: Dynamic scheduling for large language model serving,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Llumnix: Dynamic scheduling for large language model serving,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.234092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.234092Z digest=sha256:e06b0b5841f5bbb2b64122d8e12e1cc8ea37b4d0ffa1ad519afa5e748bb5c510

Observation 14004c0c-c6cd-4aca-b686-b31b093e8020 · outbound

This paper cites Batch: Machine learn- ing inference serving on serverless platforms with adaptive batching,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Batch: Machine learn- ing inference serving on serverless platforms with adaptive batching,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.650655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.240589Z digest=sha256:70ab08c6d53a0587c248d5305b538743b0a2efbfa70e27a9d91bc69d54d4c7f6

Observation 4b467c65-bc88-42d3-942d-1f11c637c742 · outbound

This paper cites an unresolved cited work.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:34:40.621627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.246143Z digest=sha256:046274cad22fbbf13c9036ff77367bf4456044a48f9101b4de7f98658ebcb879

Observation 92ebd5ed-4418-4c73-aae6-d5b413ab1033 · outbound

This paper cites Efficient deep neural network serving: Fast and furious,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Efficient deep neural network serving: Fast and furious,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.593282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.251838Z digest=sha256:2316d22352fc997b8d8a51354c96fa161ebb41530c5b0681729cbdeefee9b80a

Observation 7b077c26-7f02-4c97-bc65-36badbd8e441 · outbound

This paper cites A comprehensive review of yolo architectures in computer vision: From yolov1 to yolov8 and yolo-nas,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees A comprehensive review of yolo architectures in computer vision: From yolov1 to yolov8 and yolo-nas,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.257810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.257810Z digest=sha256:5b7360edd2a94cf44f49010604ef7ad27f0b0f9b6f51dd0c9ab103b18aa754ee

Observation 89693262-08b4-4aca-9dad-bd170f4f25b9 · outbound

This paper cites Survey of uncertainty estimation in large language models-sources, methods, applications, and challenge,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Survey of uncertainty estimation in large language models-sources, methods, applications, and challenge,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.544746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.263884Z digest=sha256:b77d5c9f48f305ad0c4532025ad9268969f29c7688496c5a24c2fb76a328cc83

Observation c44c90cf-5c58-4362-ae34-376fab391135 · outbound

This paper cites Enabling efficient batch serving for lmaas via generation length prediction,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Enabling efficient batch serving for lmaas via generation length prediction,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.511175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.270108Z digest=sha256:75c6e3b5eed3a983c44ed9b28f070787f1b76db2e5cbfbef9af36bdbd5fbf1fb

Observation 6a2720c6-3fe5-4df3-8cab-5ac2e9704ee3 · outbound

This paper cites [performance]: [v1] increasing the request batch size causes a significant drop in performance,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees [performance]: [v1] increasing the request batch size causes a significant drop in performance,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.490366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.276247Z digest=sha256:a4f7499f7715944b0dfea8c0f96177f29fc4237a9178a76f6253e5cc89244de6

Observation 7c6a5c16-1a97-460e-8382-232caa403ddf · outbound

This paper cites [performance]: Added request take too much time, and the model will not run untill all the request are added into the cache,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees [performance]: Added request take too much time, and the model will not run untill all the request are added into the cache,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.467521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.282691Z digest=sha256:195ae79042da56635fec540bc272978f2b26fb0d7d83b2d29e5a3abf5cad3140

Observation d333ecd6-60c1-43c3-a91c-7f704869250c · outbound

This paper cites [performance]: Why does the tpot increase with the request rate increase?.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees [performance]: Why does the tpot increase with the request rate increase?

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.445506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.289069Z digest=sha256:d4ff030e5fef8c616d330f83036e249d10b97d054ec127ef258048e058228caf

Observation 1e11c2fc-df77-4507-a78f-0a568eedddaa · outbound

This paper cites [performance]: Ttft spikes when qps increases during deepseek- r1 testing with tp8 and pp2,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees [performance]: Ttft spikes when qps increases during deepseek- r1 testing with tp8 and pp2,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.420072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.295612Z digest=sha256:169b2e5964cb8e5f8fc150f7109b953b8c85bbed813f28daa4a8a0725a175372

Observation 63039161-2ca4-4e2c-a2b2-1f7096b417f9 · outbound

This paper cites [performance]: poor performance in pipeline parallesm when batch-size is large,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees [performance]: poor performance in pipeline parallesm when batch-size is large,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.389296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.302539Z digest=sha256:b54f9ae479e31bc14e357522714a8ebdcfc024321c9c5a1da5a99a563d94c408

Observation 6a219196-1625-497f-9766-d0269f70d488 · outbound

This paper cites [performance]: How to improve performance under concurrency,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees [performance]: How to improve performance under concurrency,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.354487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.313230Z digest=sha256:772198a936435511a5e97abb07a78fcd573ad7ad42632b026fcc1c403f48d9b2

Observation dac3f033-f474-435b-9908-ae705c93400c · outbound

This paper cites {SHEPHERD}: Serv- ing{DNNs}in the wild,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees {SHEPHERD}: Serv- ing{DNNs}in the wild,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.331059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.318307Z digest=sha256:7b5b1aa0a971894439b93688bd91216582f4afac26fd95811e036527eaa1e4ae

Observation fc85e42c-445e-4c5f-9ca3-c736a7f1623f · outbound

This paper cites Serving{DNNs}like clockwork: Performance predictability from the bottom up,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Serving{DNNs}like clockwork: Performance predictability from the bottom up,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.299244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.324673Z digest=sha256:5c359b58012b55301ac8a98f756168a6fa15dc7de45bdb580ab9a7c8fde171b7

Observation 7f1229fe-691c-4f25-841c-a679c64f1c80 · outbound

This paper cites {INFaaS}: Automated model-less inference serving,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees {INFaaS}: Automated model-less inference serving,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.270319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.332023Z digest=sha256:ddf68e6fb0f3dcd31f27b6303b96b92cd2ccb6ad1ac223df4dc928904e4ff4b4

Observation f50f9745-b263-4575-b556-0e150cb23a47 · outbound

This paper cites Llama: A heterogeneous & serverless framework for auto-tuning video analytics pipelines,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Llama: A heterogeneous & serverless framework for auto-tuning video analytics pipelines,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.239989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.337702Z digest=sha256:eaa3dcde92d9ffc9ba3ec180344a3a31d1e8f6b8724818eb6bee0686685cfd58

Observation 0abfbf75-a5bd-4d1e-ae6f-e57c0c0b8f48 · outbound

This paper cites Llumnix: Dynamic Scheduling for Large Language Model Serving.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Llumnix: Dynamic Scheduling for Large Language Model Serving

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.343001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.343001Z digest=sha256:99284db9888244d7f431cb9d2d0e9637daab12a562dc79e494b84bf22c3eed94

Observation 6d0980ed-da06-4666-a1f9-6df49d0f4ce8 · outbound

This paper cites {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.348915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.348915Z digest=sha256:68e9ba19b39be4fdd493c8a57798d0b13df7f51d2053d54fc37a945d662c61cf

Observation 21f521eb-5287-4847-bfeb-ea744a2059ac · outbound

This paper cites SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.354289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.354289Z digest=sha256:4b12b44c4fbc88c54c7bd9dfac1f26fbfb1854fa148c047cfa7bd93a48351407

Observation 6d72c8f2-8d06-49c2-bfb3-2007f5d3bef8 · outbound

This paper cites Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.360289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.360289Z digest=sha256:0b5913db6fbfbafb85980deca6106be32500e38f915533d260e7e0fbd53ccabc

Observation 6b687260-c7e0-4390-a200-4ad5d118d0ff · outbound

This paper cites Splitwise: Efficient generative llm inference using phase splitting,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Splitwise: Efficient generative llm inference using phase splitting,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.366807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.366807Z digest=sha256:799d172e1975c2a6107799842f74b15ec73de20494a5f00880690ebd2752615b

Observation f4dfdf25-cd56-47a8-abc6-fd64155fa57c · outbound

This paper cites ExeGPT: Constraint-Aware Resource Scheduling for LLM Inference.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees ExeGPT: Constraint-Aware Resource Scheduling for LLM Inference

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-15T18:34:39.600636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.373993Z digest=sha256:98c5d9fdc8a820eb215c24a6e7f321ad4a9546cd607d6a8602e3104c93d9fa77

Observation 94df0066-4de7-4408-a830-7d0e450f6759 · outbound

This paper cites Niyama : Breaking the Silos of LLM Inference Serving.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Niyama : Breaking the Silos of LLM Inference Serving

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.381247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.381247Z digest=sha256:4f19becbeca4ddae21bd3591ffaf0ffdb3b6d6a558f00f368d6bc4a149f1076a

Observation 558ef808-996f-4c05-8fbd-566875a77a00 · outbound

This paper cites Hl-codellama-chat-response dataset,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Hl-codellama-chat-response dataset,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.174778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.386759Z digest=sha256:81a54b090bf49ca551a4aa6028493558b64f311370b412cabe857e9b83d89e54

Observation 472ce62f-cb3a-439c-bf13-eacd8a120661 · outbound

This paper cites Synthetic code generations dataset,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Synthetic code generations dataset,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.143508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.392953Z digest=sha256:2c26f71ef8f601e18048e479821ced64607790c1ed08aa43e1b82c44439e0610

Observation 4d48ba10-b845-4ee8-b8c9-133f8d9f1a82 · outbound

This paper cites Code summary java dataset,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Code summary java dataset,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.111151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.398191Z digest=sha256:82ce5b4a91ee584804404780f985e758222a18535fb6ffc6417fc22439617af8

Observation 3e23a315-b542-4f1f-bce9-46d62a2dfad2 · outbound

This paper cites Code translation dataset,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Code translation dataset,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.075428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.403669Z digest=sha256:10923b69bc9c55536bcec12535cd34070ffe01e6b156b9a0b02e93fe2c6c50fd

Observation 651f9bd1-fc9f-4958-902d-92d799ad9696 · outbound

This paper cites Learned Best-Effort LLM Serving.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Learned Best-Effort LLM Serving

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.408803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.408803Z digest=sha256:52aeb341831f92f0bb501bcfb595a1f5aefd4408be757b7451da72a8751b397c

Observation be3f75a8-9c4f-4c8b-9263-3e8fb2d439ec · outbound

This paper cites Preble: Efficient Distributed Prompt Scheduling for LLM Serving.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Preble: Efficient Distributed Prompt Scheduling for LLM Serving

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.414400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.414400Z digest=sha256:6b2386fa4171d01efb5dbf9ea34ee327047743adb951c6f965973e453eeedb81

Observation 20cc696d-a5b2-448a-9bff-b0a632792fa4 · outbound

This paper cites Past-future scheduler for llm serving under sla guarantees,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Past-future scheduler for llm serving under sla guarantees,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.049337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.419712Z digest=sha256:3cec02160db02a640ad4a6b3f90963c270afa3dd570a16e6b2d8300e394860b8

Observation 44a109d9-1384-44a7-9802-babe4e010d8b · outbound

This paper cites ShareGPT Dataset: A Collection of ChatGPT Conversations,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees ShareGPT Dataset: A Collection of ChatGPT Conversations,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.019737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.425745Z digest=sha256:091411b9910d6a30fbaa6c0993a706616d5e39b386312ee80e5b7a3beee0a0d4

Observation 269a077a-b8b5-40d8-83d7-fd363b68b917 · outbound

This paper cites Integrating concurrency control in n-tier application scaling management in the cloud,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Integrating concurrency control in n-tier application scaling management in the cloud,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:39.992542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.431313Z digest=sha256:df9c54eb6ba7080dfffabcf2fd7143947a95138ef7ee22160fde20a79a4716b1

Observation 09d6d1a8-8d6c-4637-b690-604c3accdab1 · outbound

This paper cites An r-square coefficient based on final prediction error,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees An r-square coefficient based on final prediction error,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:39.966406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.436591Z digest=sha256:991806d0a873b848b4bcd4b2831145769d717d4b1e2b89fac72ef326123bca4d

Observation b2a8c6f6-3f6f-4d29-8a61-959cdc3073f1 · outbound

This paper cites Scipy 1.0: fundamental algorithms for scientific computing in python,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Scipy 1.0: fundamental algorithms for scientific computing in python,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.442683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.442683Z digest=sha256:cf258e22908ddbabc8d16651114cd893c82930f2c6c2b2b6e0920ee859e4bd71

Observation 328358dd-936f-4e48-9342-cc8dc09198ec · outbound

This paper cites Multi-dimensional sla-based resource allocation for multi-tier cloud computing systems,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Multi-dimensional sla-based resource allocation for multi-tier cloud computing systems,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:39.919225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.448659Z digest=sha256:9aaef9bc2ba2729612d0011cfe6d2a69172effd80d8cade7fb4137c3c687ac3f

Observation 780e5dfe-4108-4577-84f5-14c7ff256e7c · outbound

This paper cites Decision model for cloud comput- ing under sla constraints,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Decision model for cloud comput- ing under sla constraints,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:39.898075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.455244Z digest=sha256:45185c567c3fac92ff4dd316a565a1cd0de21561dfad0dcc31a9a6b917cf4cf0

Observation 1018c585-1bd8-4c21-88f1-880752fcb863 · outbound

This paper cites When average is not average: large response time fluctuations in n-tier systems,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees When average is not average: large response time fluctuations in n-tier systems,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:39.878885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.461915Z digest=sha256:86558dcaf617f76eaab8f026a383cc2f76511e04067d3bacf91ed1ec520d6bfa

Pith citing papers

No inbound Pith citation observations are available.