Pith. sign in

Paper Citation Record · LEDGER

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees

As of 16 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2506.19677.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.19677 v2

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:34:39.461915Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact1
  • verified fuzzy28
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a7b85b78-d2bc-4062-9493-ee49a6281da0 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Code Llama: Open Foundation Models for Code

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.135939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.135939Z digest=sha256:ca2860fc84c3ea6c72b8e2f49a98e711f10b95adaf930fae7ae19fb02dffbf87

Observation 5ae8e1bb-65c9-45b7-9679-9d03881705af · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.142344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.142344Z digest=sha256:b9af47b88cea3cf953e4fe3a7373f6d7633ffbfff139ccf7a6d10c480822adf2

Observation fa6e4015-90ca-4285-8d57-536cdf6b889d · outbound

This paper cites Qwen2.5-Coder Technical Report.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Qwen2.5-Coder Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.148481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.148481Z digest=sha256:41898160085682c3f09ae9e41d31f1f0609ef1618dc7da10268bd8cace995d6b

Observation 2487f045-311f-4609-b73a-42cec3188f6a · outbound

This paper cites StarCoder 2 and The Stack v2: The Next Generation.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees StarCoder 2 and The Stack v2: The Next Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.154863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.154863Z digest=sha256:6ff252a716a9ac8983d0f0fdd003c29dc792c5b5ee66530a643768abf4232a62

Observation 0419d7b9-9f88-44a2-bc25-02f8c6df7a6b · outbound

This paper cites Fine tuning large language model for secure code generation,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Fine tuning large language model for secure code generation,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.836208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:34:39.160386Z digest=sha256:0360f190d72bb6f11ab244b15c4b65343b9546b1d135e36746d3bddc0ea1c2ba

Observation 975034c9-d087-43a5-9f57-3da2366bb72c · outbound

This paper cites RepoHyper: Search-Expand-Refine on Semantic Graphs for Repository-Level Code Completion.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees RepoHyper: Search-Expand-Refine on Semantic Graphs for Repository-Level Code Completion

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.164978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.164978Z digest=sha256:5f32895df19d3cab85bff9b4e2474414e3b303fc72eb8f3fe5ac29fc512886ee

Observation d2e1e620-1d0c-4cde-94ba-99ea4809c32d · outbound

This paper cites A Survey on Large Language Models for Code Generation.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees A Survey on Large Language Models for Code Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.171573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.171573Z digest=sha256:f142a7bb09404803dc55c4900b1fc63e30d16b3bd779aafc3fbbca92b9d143f0

Observation 81fda2bd-3062-49c6-b4bd-e9f3b4f75e8f · outbound

This paper cites Language models for code completion: A practical evaluation,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Language models for code completion: A practical evaluation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.808922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:34:39.177142Z digest=sha256:9f7374f288cc8f20e84bf4cd3443a7d48ea2390f0334133d8192bb22027fe6b7

Observation 984a41e3-9824-4d89-9e67-147a8cc2a164 · outbound

This paper cites AI-assisted Code Authoring at Scale: Fine-tuning, deploying, and mixed methods evaluation.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees AI-assisted Code Authoring at Scale: Fine-tuning, deploying, and mixed methods evaluation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.182174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.182174Z digest=sha256:ac188a574f38f2729d0787a6149520cbb1c6c63941726b4c5913354099b49030

Observation de5cda9c-beef-486e-8e3b-f51dd0b660cf · outbound

This paper cites A survey on large language model (llm) security and privacy: The good, the bad, and the ugly,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees A survey on large language model (llm) security and privacy: The good, the bad, and the ugly,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.188655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.188655Z digest=sha256:af5c4ec52704e8bddf686cda7c89853a6c067ca571eb0aab077cd737161fe69d

Observation fb1b5d8a-fc0d-4f20-baaa-3956624fb948 · outbound

This paper cites Security and privacy challenges of large language models: A survey,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Security and privacy challenges of large language models: A survey,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.194272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.194272Z digest=sha256:3d69f5cedc1e21ac58ba33fd1ae2578f81aff487c31016b3ed37c05f61e6114c

Observation e4245da9-307c-43c7-9d71-8c95c8245d53 · outbound

This paper cites Hierarchical Repository-Level Code Summarization for Business Applications Using Local LLMs.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Hierarchical Repository-Level Code Summarization for Business Applications Using Local LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.201458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.201458Z digest=sha256:cff7e7451e65fd45b1556971deec293fe94ae69de59f3a154479f30721a5fb7b

Observation a0de50b7-d19b-447a-82fa-61a647a202e0 · outbound

This paper cites Using ollama,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Using ollama,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.206736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.206736Z digest=sha256:b1390bea86937a67e6b67dea6443ab83b2cc856a158da60b0a880ae676a7260f

Observation a86f1378-8681-47e0-a439-8386e407f3f4 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Efficient memory management for large language model serving with pagedattention,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.212463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.212463Z digest=sha256:7010b8c4b08c33fd7c923e464dd2db9d560814d29813cc64797e536b96945a6c

Observation 6a3e9bb0-2cca-41c1-acce-1740743c9e2c · outbound

This paper cites Efficiently programming large language models using sglang.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Efficiently programming large language models using sglang

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.218200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.218200Z digest=sha256:6dca1b7a6737da6e9b379120b9b3f13e62d634fc2dbb0aad2b579ed1866d3e2a

Observation 88e3acd8-986c-4367-ae0a-968afae4c9ba · outbound

This paper cites an unresolved cited work.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:34:40.707301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:34:39.223034Z digest=sha256:ee2d66fb3c0398261a79ca3cc2c745de8074a383df2c2503853987974077b822

Observation f9dfc8c3-0af7-4947-9eb9-76716a57586a · outbound

This paper cites Orca: A distributed serving system for{Transformer-Based}generative models,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Orca: A distributed serving system for{Transformer-Based}generative models,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.682085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:34:39.228640Z digest=sha256:1b167324fdaabff4e8c268329f4946441ce33a067f2819f6bfcb6fb3266547e2

Observation 6b06a26a-ab7e-4c13-8a75-f654254649d1 · outbound

This paper cites Llumnix: Dynamic scheduling for large language model serving,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Llumnix: Dynamic scheduling for large language model serving,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.234092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.234092Z digest=sha256:e06b0b5841f5bbb2b64122d8e12e1cc8ea37b4d0ffa1ad519afa5e748bb5c510

Observation 14004c0c-c6cd-4aca-b686-b31b093e8020 · outbound

This paper cites Batch: Machine learn- ing inference serving on serverless platforms with adaptive batching,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Batch: Machine learn- ing inference serving on serverless platforms with adaptive batching,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.650655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:34:39.240589Z digest=sha256:eb960b6d72bdc2e90567ee06506f36d9d26847ac7c97e28c7974e5efc187304d

Observation 4b467c65-bc88-42d3-942d-1f11c637c742 · outbound

This paper cites an unresolved cited work.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:34:40.621627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:34:39.246143Z digest=sha256:31f68d47a01ab89debcb9310d89430629f03f323b5792d2aa34d139cef459a05

Observation 92ebd5ed-4418-4c73-aae6-d5b413ab1033 · outbound

This paper cites Efficient deep neural network serving: Fast and furious,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Efficient deep neural network serving: Fast and furious,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.593282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:34:39.251838Z digest=sha256:494f62a5ba90e962361eaa31cad741f78479db179d4ae7bf0b4d160103b05acd

Observation 7b077c26-7f02-4c97-bc65-36badbd8e441 · outbound

This paper cites A comprehensive review of yolo architectures in computer vision: From yolov1 to yolov8 and yolo-nas,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees A comprehensive review of yolo architectures in computer vision: From yolov1 to yolov8 and yolo-nas,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.257810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.257810Z digest=sha256:5b7360edd2a94cf44f49010604ef7ad27f0b0f9b6f51dd0c9ab103b18aa754ee

Observation 89693262-08b4-4aca-9dad-bd170f4f25b9 · outbound

This paper cites Survey of uncertainty estimation in large language models-sources, methods, applications, and challenge,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Survey of uncertainty estimation in large language models-sources, methods, applications, and challenge,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.544746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:34:39.263884Z digest=sha256:6d52214e7007e5f0ec69c763ba417df0c4dfc6ab33e601786e9c8d6b81ad6792

Observation c44c90cf-5c58-4362-ae34-376fab391135 · outbound

This paper cites Enabling efficient batch serving for lmaas via generation length prediction,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Enabling efficient batch serving for lmaas via generation length prediction,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.511175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:34:39.270108Z digest=sha256:bec38aad9b2fbbcd868d436197ee01ada163a11058ca13a239c2be3680578efc

Observation 6a2720c6-3fe5-4df3-8cab-5ac2e9704ee3 · outbound

This paper cites [performance]: [v1] increasing the request batch size causes a significant drop in performance,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees [performance]: [v1] increasing the request batch size causes a significant drop in performance,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.490366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:34:39.276247Z digest=sha256:40acc242eeaff6f29c4697c9b896fa6913a4050d35511cf2a21bc19032e6b465

Observation 7c6a5c16-1a97-460e-8382-232caa403ddf · outbound

This paper cites [performance]: Added request take too much time, and the model will not run untill all the request are added into the cache,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees [performance]: Added request take too much time, and the model will not run untill all the request are added into the cache,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.467521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:34:39.282691Z digest=sha256:9fa1d3d55e1f4a76e9fdd584b416fc8f1f30316351e20d6a06fefed4775f75f0

Observation d333ecd6-60c1-43c3-a91c-7f704869250c · outbound

This paper cites [performance]: Why does the tpot increase with the request rate increase?.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees [performance]: Why does the tpot increase with the request rate increase?

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.445506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:34:39.289069Z digest=sha256:f9b432de1470c32680037955f9433511b43cb18ccf30c2e01a5a8a8b017ab877

Observation 1e11c2fc-df77-4507-a78f-0a568eedddaa · outbound

This paper cites [performance]: Ttft spikes when qps increases during deepseek- r1 testing with tp8 and pp2,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees [performance]: Ttft spikes when qps increases during deepseek- r1 testing with tp8 and pp2,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.420072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:34:39.295612Z digest=sha256:b708fe9186dde228ea6f45890ffde1d10e6ec9945127d62ecfacc4cf540ef22e

Observation 63039161-2ca4-4e2c-a2b2-1f7096b417f9 · outbound

This paper cites [performance]: poor performance in pipeline parallesm when batch-size is large,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees [performance]: poor performance in pipeline parallesm when batch-size is large,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.389296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:34:39.302539Z digest=sha256:7a9ff9f00d3f7f6b1e57193214f955babf5f1d33915271b0c5f56a80d08df8b5

Observation 6a219196-1625-497f-9766-d0269f70d488 · outbound

This paper cites [performance]: How to improve performance under concurrency,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees [performance]: How to improve performance under concurrency,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.354487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:34:39.313230Z digest=sha256:59be30c43e7d6049f6bd1fab003333e8c7982322f0591363146161c7627d6368

Observation dac3f033-f474-435b-9908-ae705c93400c · outbound

This paper cites {SHEPHERD}: Serv- ing{DNNs}in the wild,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees {SHEPHERD}: Serv- ing{DNNs}in the wild,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.331059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:34:39.318307Z digest=sha256:1bb2eaef26db25599b4663bf4b355b136d6238177fd4df19a0fc2f74f1913cf8

Observation fc85e42c-445e-4c5f-9ca3-c736a7f1623f · outbound

This paper cites Serving{DNNs}like clockwork: Performance predictability from the bottom up,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Serving{DNNs}like clockwork: Performance predictability from the bottom up,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.299244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:34:39.324673Z digest=sha256:45a337129947b67366994210dc4eac5d628f06237fe198d24ce9a7d3b07599ba

Observation 7f1229fe-691c-4f25-841c-a679c64f1c80 · outbound

This paper cites {INFaaS}: Automated model-less inference serving,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees {INFaaS}: Automated model-less inference serving,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.270319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:34:39.332023Z digest=sha256:e5f9306c0344ad61b927f2a752527387133077ff9abe364a7cf5679ebc26cf51

Observation f50f9745-b263-4575-b556-0e150cb23a47 · outbound

This paper cites Llama: A heterogeneous & serverless framework for auto-tuning video analytics pipelines,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Llama: A heterogeneous & serverless framework for auto-tuning video analytics pipelines,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.239989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:34:39.337702Z digest=sha256:e0a45d2cfc27ca2f2454e2b6e01b805e46ed02219f7c0571629d5b0ba75a11b2

Observation 0abfbf75-a5bd-4d1e-ae6f-e57c0c0b8f48 · outbound

This paper cites Llumnix: Dynamic Scheduling for Large Language Model Serving.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Llumnix: Dynamic Scheduling for Large Language Model Serving

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.343001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.343001Z digest=sha256:99284db9888244d7f431cb9d2d0e9637daab12a562dc79e494b84bf22c3eed94

Observation 6d0980ed-da06-4666-a1f9-6df49d0f4ce8 · outbound

This paper cites {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.348915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.348915Z digest=sha256:68e9ba19b39be4fdd493c8a57798d0b13df7f51d2053d54fc37a945d662c61cf

Observation 21f521eb-5287-4847-bfeb-ea744a2059ac · outbound

This paper cites SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.354289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.354289Z digest=sha256:4b12b44c4fbc88c54c7bd9dfac1f26fbfb1854fa148c047cfa7bd93a48351407

Observation 6d72c8f2-8d06-49c2-bfb3-2007f5d3bef8 · outbound

This paper cites Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.360289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.360289Z digest=sha256:0b5913db6fbfbafb85980deca6106be32500e38f915533d260e7e0fbd53ccabc

Observation 6b687260-c7e0-4390-a200-4ad5d118d0ff · outbound

This paper cites Splitwise: Efficient generative llm inference using phase splitting,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Splitwise: Efficient generative llm inference using phase splitting,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.366807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.366807Z digest=sha256:799d172e1975c2a6107799842f74b15ec73de20494a5f00880690ebd2752615b

Observation f4dfdf25-cd56-47a8-abc6-fd64155fa57c · outbound

This paper cites ExeGPT: Constraint-Aware Resource Scheduling for LLM Inference.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees ExeGPT: Constraint-Aware Resource Scheduling for LLM Inference

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-15T18:34:39.600636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:34:39.373993Z digest=sha256:f0bca03828428f056519af65b3613b487f11cc293e4c9e5682c67dbd782d7677

Observation 94df0066-4de7-4408-a830-7d0e450f6759 · outbound

This paper cites Niyama : Breaking the Silos of LLM Inference Serving.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Niyama : Breaking the Silos of LLM Inference Serving

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.381247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.381247Z digest=sha256:4f19becbeca4ddae21bd3591ffaf0ffdb3b6d6a558f00f368d6bc4a149f1076a

Observation 558ef808-996f-4c05-8fbd-566875a77a00 · outbound

This paper cites Hl-codellama-chat-response dataset,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Hl-codellama-chat-response dataset,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.174778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:34:39.386759Z digest=sha256:8853aa3f383701f5883efcd0da7d1c6d5930fbf7aec57fea5c519b41fa23cc18

Observation 472ce62f-cb3a-439c-bf13-eacd8a120661 · outbound

This paper cites Synthetic code generations dataset,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Synthetic code generations dataset,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.143508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:34:39.392953Z digest=sha256:4cdfb14ff4ee0ea0b5d198b5e70acbfdb7f1fa65f13e1f84f3dc9c9a59a27b3d

Observation 4d48ba10-b845-4ee8-b8c9-133f8d9f1a82 · outbound

This paper cites Code summary java dataset,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Code summary java dataset,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.111151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:34:39.398191Z digest=sha256:765c6bbcd8756124cecaa3971b350cd7c7de91db45c405774a2581ec10214153

Observation 3e23a315-b542-4f1f-bce9-46d62a2dfad2 · outbound

This paper cites Code translation dataset,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Code translation dataset,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.075428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:34:39.403669Z digest=sha256:e58d15cad241e0fb08a90c8b359f553866b6f3008de70feaa8f03793314cd314

Observation 651f9bd1-fc9f-4958-902d-92d799ad9696 · outbound

This paper cites Learned Best-Effort LLM Serving.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Learned Best-Effort LLM Serving

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.408803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.408803Z digest=sha256:52aeb341831f92f0bb501bcfb595a1f5aefd4408be757b7451da72a8751b397c

Observation be3f75a8-9c4f-4c8b-9263-3e8fb2d439ec · outbound

This paper cites Preble: Efficient Distributed Prompt Scheduling for LLM Serving.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Preble: Efficient Distributed Prompt Scheduling for LLM Serving

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.414400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.414400Z digest=sha256:6b2386fa4171d01efb5dbf9ea34ee327047743adb951c6f965973e453eeedb81

Observation 20cc696d-a5b2-448a-9bff-b0a632792fa4 · outbound

This paper cites Past-future scheduler for llm serving under sla guarantees,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Past-future scheduler for llm serving under sla guarantees,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.049337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:34:39.419712Z digest=sha256:4a7f58ad1744a6467e2cf1914d9ccf077466ca7c960ce199b92a205f2f621a2e

Observation 44a109d9-1384-44a7-9802-babe4e010d8b · outbound

This paper cites ShareGPT Dataset: A Collection of ChatGPT Conversations,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees ShareGPT Dataset: A Collection of ChatGPT Conversations,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.019737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:34:39.425745Z digest=sha256:01dc55c037461f9315477d1dc4f6f499b847b0d6e88fee37fc6cff6300f12f98

Observation 269a077a-b8b5-40d8-83d7-fd363b68b917 · outbound

This paper cites Integrating concurrency control in n-tier application scaling management in the cloud,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Integrating concurrency control in n-tier application scaling management in the cloud,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:39.992542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:34:39.431313Z digest=sha256:dc46f85f33a1a24c4783e578156c51a2d6287f36fb05ebaea7ff1403c794549c

Observation 09d6d1a8-8d6c-4637-b690-604c3accdab1 · outbound

This paper cites An r-square coefficient based on final prediction error,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees An r-square coefficient based on final prediction error,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:39.966406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:34:39.436591Z digest=sha256:996f6b0ba1364bd1716c449d81b824e005f65f8df19d47d7f5af95556c9b2cad

Observation b2a8c6f6-3f6f-4d29-8a61-959cdc3073f1 · outbound

This paper cites Scipy 1.0: fundamental algorithms for scientific computing in python,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Scipy 1.0: fundamental algorithms for scientific computing in python,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.442683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.442683Z digest=sha256:cf258e22908ddbabc8d16651114cd893c82930f2c6c2b2b6e0920ee859e4bd71

Observation 328358dd-936f-4e48-9342-cc8dc09198ec · outbound

This paper cites Multi-dimensional sla-based resource allocation for multi-tier cloud computing systems,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Multi-dimensional sla-based resource allocation for multi-tier cloud computing systems,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:39.919225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:34:39.448659Z digest=sha256:e05912d09b2364c32dbea7d7c5f50f399b5994f18640d940a86a02c78c364699

Observation 780e5dfe-4108-4577-84f5-14c7ff256e7c · outbound

This paper cites Decision model for cloud comput- ing under sla constraints,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Decision model for cloud comput- ing under sla constraints,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:39.898075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:34:39.455244Z digest=sha256:36bc1d4afbf2e6315be5458e0c1139248265913072d384ccbac5f0b93c56e23c

Observation 1018c585-1bd8-4c21-88f1-880752fcb863 · outbound

This paper cites When average is not average: large response time fluctuations in n-tier systems,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees When average is not average: large response time fluctuations in n-tier systems,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:39.878885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T18:34:39.461915Z digest=sha256:f0870e7bf5e9bc349ea63e9fdad79b24965df7d75cf2c1a94a7a03c85c9ff214

Pith citing papers

No inbound Pith citation observations are available.