Pith. sign in

Paper Citation Record · LEDGER

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving

As of 18 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2608.06557.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06557 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:37:20.782815Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact3
  • verified fuzzy20
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f7c086e7-da91-4cf2-9e73-416473ca3692 · outbound

This paper cites Qwen2 technical report,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Qwen2 technical report,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.614826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.614826Z digest=sha256:5a3341174a1340325215571e6a05902a558ff89dcdc7644b8208098c14eb3374

Observation 6df7e7c6-661c-4448-829f-664448a2ff39 · outbound

This paper cites Vidur: A large-scale simulation framework for llm inference,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Vidur: A large-scale simulation framework for llm inference,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.996527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T14:37:20.619844Z digest=sha256:fd3996e0d26d62a02e8bad4a06cd26c81b0c46b51f4977ac638af4a16704b525

Observation e40999ce-2470-470b-a2f0-1178eb087af6 · outbound

This paper cites Taming throughput-latency tradeoff in llm inference with sarathi-serve,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Taming throughput-latency tradeoff in llm inference with sarathi-serve,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.623715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.623715Z digest=sha256:b22631557f18092863086a006fc2dea59b6aff1ca0e48a49ecd41ba1783a0918

Observation 31e0c590-f737-4ce3-8d93-0fcd5148538c · outbound

This paper cites No request left behind: Tackling heterogeneity in long-context llm inference with medha,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving No request left behind: Tackling heterogeneity in long-context llm inference with medha,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.981307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T14:37:20.627580Z digest=sha256:7aceec6ad82aa169dbe38dd2343a4a5242d9a8c3331c345dd5356b4775b41d4f

Observation 2c8554d3-2816-43c6-90f4-9db367884daa · outbound

This paper cites Llama 3 model card,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Llama 3 model card,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.969707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T14:37:20.634509Z digest=sha256:35681a85cc4e1d70f405b78310c7be05a23e699f929a411bc3899259f0e96594

Observation d35df27b-d3f1-4666-867e-a1851c6c7424 · outbound

This paper cites GQA: Training generalized multi-query transformer models from multi-head checkpoints,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving GQA: Training generalized multi-query transformer models from multi-head checkpoints,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.638243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.638243Z digest=sha256:256768bc2137bec2b5535713acaa3bab038301e18cb153ce9df42de1f513d554

Observation 09985aab-a5fe-4d0a-b00c-c0fca85c91eb · outbound

This paper cites Qwen-bailian anonymous dataset,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Qwen-bailian anonymous dataset,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.954386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T14:37:20.641506Z digest=sha256:190fe1cba46dc181bef6ae874a7f88b13440d4e981c627d19b52c60e816fe362

Observation 0b689cdd-67dd-48e1-8f04-25bec86b029b · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Evaluating Large Language Models Trained on Code

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.644989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.644989Z digest=sha256:1d605b506f0b0523c6c81de6b5be263974ef87b0fd3e0dfc02f911f259dfca4f

Observation 10c15315-24f1-4f0c-8d2d-322151282dd5 · outbound

This paper cites SLOs-Serve: Optimized Serving of Multi-SLO LLMs.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving SLOs-Serve: Optimized Serving of Multi-SLO LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.648215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.648215Z digest=sha256:195e61ed94834b00f84428deca324c11aade122e580919b8fe07db1684831b45

Observation b5a122f2-266f-42e2-9cab-25cc238628ab · outbound

This paper cites ATP: Adaptive Tensor Parallelism for Foundation Models.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving ATP: Adaptive Tensor Parallelism for Foundation Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.651825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.651825Z digest=sha256:87daee70b6f8249444665057b84001f24b8b40f06161b32c9272fb4cbb93ef08

Observation 293345e1-d02f-42ba-b9db-730dc9fa4821 · outbound

This paper cites Jockey: guaranteed job latency in data parallel clusters,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Jockey: guaranteed job latency in data parallel clusters,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.655496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.655496Z digest=sha256:8c5fb33e8962849d55d3d586951ae7f15728eb80c996ff6ed5f2cb11b8d1337c

Observation c648caa0-8e40-47cf-afd1-8f33e914deb2 · outbound

This paper cites Prompt cache: Modular attention reuse for low-latency inference,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Prompt cache: Modular attention reuse for low-latency inference,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.658651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.658651Z digest=sha256:9f82aecec548902dd96593cb71400a3381ee3470d6f1f611dd65d6c6e8c1865e

Observation a4c25151-548c-47f8-8a7e-89bac272047a · outbound

This paper cites Qoserve: Breaking the silos of llm inference serving,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Qoserve: Breaking the silos of llm inference serving,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.665753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.665753Z digest=sha256:dae36de8c4ded848234ef59ec058b5c2af8c96c4c34e4711a14bea64fc269a91

Observation 84208f07-9b5c-4c1c-9016-7f977104d077 · outbound

This paper cites [Online].

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving [Online]

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.939258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T14:37:20.669126Z digest=sha256:e3d4c35182cb2c9c79a2160321907d220295979ba2f745146250d4c7ca317b49

Observation c7baa3a3-8134-4d35-9bb9-30ab31b8392e · outbound

This paper cites Kvquant: towards 10 million context length llm inference with kv cache quantization,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Kvquant: towards 10 million context length llm inference with kv cache quantization,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.930056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T14:37:20.672231Z digest=sha256:8beb478198c733b33c0be05ea4b999a6136b9412eb05e90deda0ede494c704a3

Observation 6c523305-35dd-42e4-a6ee-e89a4984e8a5 · outbound

This paper cites Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.675426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.675426Z digest=sha256:23c5cb6cdedb4083750a228a1e90025533cb4c0b27dc60e31a2012fdd23c431c

Observation 7c541652-eada-4457-8dda-348ec2c2091c · outbound

This paper cites A quantitative measure of fairness and discrimination for resource allocation in shared computer systems,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving A quantitative measure of fairness and discrimination for resource allocation in shared computer systems,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.921044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T14:37:20.678834Z digest=sha256:10fcec7219b2be6759c2d4d0a47fd4d52dc760fcbcda08dda5d9835836bbcdfc

Observation 666e0fd2-14f4-44d4-9676-d687734bee90 · outbound

This paper cites Sageserve: Optimizing llm serving on cloud data centers with forecast aware auto-scaling,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Sageserve: Optimizing llm serving on cloud data centers with forecast aware auto-scaling,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.686402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.686402Z digest=sha256:777bd6fd755ae92a559fba0122f12543a132752cc5521197347a89c10be97f06

Observation 638bcc91-5ef2-475b-a47f-6f2599c2e29d · outbound

This paper cites Learned Best-Effort LLM Serving.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Learned Best-Effort LLM Serving

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:37:21.557527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T14:37:20.689429Z digest=sha256:051238206aa347c1cbb01cec06f1e9bd03f7e70fd96bea4e0d60763612cc3776

Observation 8870396d-c2fa-4414-909d-9ff6483b0ca2 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Efficient memory management for large language model serving with pagedattention,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.692788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.692788Z digest=sha256:a4952df0ca1e380ac8bde424ffe0d30fcbeab31b7457eff0aee3156547918f47

Observation c75f11ff-066f-45bc-97cd-339a9e211cd0 · outbound

This paper cites Tokenscale: Timely and accurate autoscaling for disaggregated llm serving with token velocity,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Tokenscale: Timely and accurate autoscaling for disaggregated llm serving with token velocity,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.696067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.696067Z digest=sha256:a9783b0bdbc3c75adf2777ef3f0c0b67d357f4c3cebf33b1468527eb24a944c3

Observation 0075d4b1-e848-48d3-8f0e-b3bc98dd3a63 · outbound

This paper cites Revisiting disaggregated large language model serving for performance and energy implications,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Revisiting disaggregated large language model serving for performance and energy implications,

Reference 22

Resolution
metadata mismatch
raw_fallback, observed 2026-08-15T14:37:21.434037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T14:37:20.699193Z digest=sha256:32a5290b372922d143bc9c9fb82b5078333528a2abf44c042ab542389a9da6e2

Observation 35e1ccb2-db4c-417a-b66f-869f1f0a0f41 · outbound

This paper cites AlpaServe: Statistical multiplexing with model parallelism for deep learning serving,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving AlpaServe: Statistical multiplexing with model parallelism for deep learning serving,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.911478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T14:37:20.702290Z digest=sha256:598b834b4541f25567e489311f62be5f2762718a27be2390cd5efb2cac20f0a1

Observation 09a2ad14-8c11-4154-931d-55f0e279e7c7 · outbound

This paper cites Scheduling algorithms for multiprogramming in a hard-real-time environment,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Scheduling algorithms for multiprogramming in a hard-real-time environment,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.705254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.705254Z digest=sha256:52fad03b043326a5d896e5a1ad368331d9ba4cb7eb1e681ffbb3b7d3a352bf63

Observation 2c7146dc-68e0-4829-9115-ddf59a1a2310 · outbound

This paper cites Lmcache: An efficient kv cache layer for enterprise-scale llm inference,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Lmcache: An efficient kv cache layer for enterprise-scale llm inference,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.708627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.708627Z digest=sha256:c95339e30c972afb7bf0253f7df2f05f862e21a4989b96860e373e499d424cb0

Observation 9bac5c43-334a-4fb1-9e22-81af19b95c07 · outbound

This paper cites Ai-dynamo,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Ai-dynamo,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.902193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T14:37:20.711621Z digest=sha256:ab712aedbe12a3d789ffd8fdcfa834b40b95da79ced707a4f5932f6bb355d17d

Observation 8427720a-7675-417a-85e9-870ec2923103 · outbound

This paper cites Nvidia gb200 nvl partition,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Nvidia gb200 nvl partition,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.892252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T14:37:20.714811Z digest=sha256:543994e035f076054e7a4c76389c989e4024b930623b979323cfeb74ea32ad9c

Observation 61c4a06c-d981-41cb-8d49-e3032ab53b05 · outbound

This paper cites Nvidia gb200 nvl72 delivers trillion-parameter llm training and real-time inference,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Nvidia gb200 nvl72 delivers trillion-parameter llm training and real-time inference,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.882768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T14:37:20.718007Z digest=sha256:4e78bbe5fe9813fef8e0bb01a3f102217d6edc47a0bad13a9d572b11a80f1931

Observation 503f3e8e-a236-4824-8677-3a68d0798b69 · outbound

This paper cites Splitwise: Efficient generative llm inference using phase splitting,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Splitwise: Efficient generative llm inference using phase splitting,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.721320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.721320Z digest=sha256:ac7aa715b6185a67a4a9488ce24139328cf505a70890e5b18e2a85684b35ec92

Observation 536b8c62-c3b9-4e5d-81ac-a0e129befb8b · outbound

This paper cites Conserve: Fine-grained gpu harvesting for llm online and offline co-serving,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Conserve: Fine-grained gpu harvesting for llm online and offline co-serving,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.872872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T14:37:20.724311Z digest=sha256:6a57a51f91090be6f8cb8594ed158694fbb3de2bc18edaf5343eb5af39e48d09

Observation 4541a13c-fa37-471c-9027-06615b4e1c41 · outbound

This paper cites Mooncake: trading more storage for less computation — a kvcache-centric architecture for serving llm chatbot,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Mooncake: trading more storage for less computation — a kvcache-centric architecture for serving llm chatbot,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.862931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T14:37:20.730929Z digest=sha256:b180efe4e6f871a407ebba045d3210ecc211b569d9b4c9ac53b0c89d2f14d1e2

Observation 39b8a3ff-44f9-43c0-b4e8-e3347737ddb0 · outbound

This paper cites Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:37:21.206726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T14:37:20.733967Z digest=sha256:8cc2511e8201aff7d2a3c1c339ebf08837d9884451dff0f2a78ec39f8895b860

Observation 651a8d8a-658b-4efc-a648-8015e9df3fe2 · outbound

This paper cites Timecard: controlling user-perceived delays in server-based mobile applications,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Timecard: controlling user-perceived delays in server-based mobile applications,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.737451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.737451Z digest=sha256:4866b7765b4942fd93a27677140a869f40dec6cd8457b7ed6b8cf8a636c77e2b

Observation a41a1345-3f21-4cee-b901-5143ae3a48ea · outbound

This paper cites ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.727436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.727436Z digest=sha256:24dd3c632c70842e9e09fe95d0ea5fd8e1f3095833d3dc4bd22eca08c9ec02e1

Observation df6804cb-9c70-46ad-9ab9-a7373a2c6000 · outbound

This paper cites Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:37:21.116238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T14:37:20.743928Z digest=sha256:784108e3eb55db1892bdd42a161f2467a18fbf4eecf148e071f636304ea5bdc9

Observation b3e2c9db-a907-49e0-80fc-9c80050748cc · outbound

This paper cites Kvcache cache in the wild: Characterizing and optimizing kvcache cache at a large cloud provider,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Kvcache cache in the wild: Characterizing and optimizing kvcache cache at a large cloud provider,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.852825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T14:37:20.747222Z digest=sha256:1fbcf1b8873920fd65f3daae760e053732396f8afc97ff11d498f63630925561

Observation 3b7e6a22-325d-42ac-8e81-ba10fa374696 · outbound

This paper cites Better never than late: meeting deadlines in datacenter networks,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Better never than late: meeting deadlines in datacenter networks,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.750430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.750430Z digest=sha256:28f50cc06d06fdfddd0cef558a02abff80d78490a93d9eb753b447c731a4ca1e

Observation 9c690888-fb43-4994-b527-8f98516dfc4c · outbound

This paper cites Preble: Efficient Distributed Prompt Scheduling for LLM Serving.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Preble: Efficient Distributed Prompt Scheduling for LLM Serving

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.740627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.740627Z digest=sha256:ac988de85631f37067dbbc4ea766e3a928882491e34ae23c9864225ff022880f

Observation 6fd0007d-fff8-4e4c-b1d6-3162977ac553 · outbound

This paper cites Aegaeon: Effective gpu pooling for concurrent llm serving on the market,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Aegaeon: Effective gpu pooling for concurrent llm serving on the market,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.757321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.757321Z digest=sha256:d318357f84d980004741352b9089b08a736778db3a7efd3ecbb430bdd6a60296

Observation 613b04c9-77a8-40b0-98c2-98966bceb4b5 · outbound

This paper cites Orca: A distributed serving system for transformer-based generative models,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Orca: A distributed serving system for transformer-based generative models,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.831594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T14:37:20.760766Z digest=sha256:4bd551d9e894c8e0277919b5d80ff4af34b1076666795f35b98da255e38998c9

Observation 8c723fcd-c9af-4d71-9627-8c8a942f1dd0 · outbound

This paper cites Superinfer: Slo-aware rotary scheduling and memory management for llm inference on superchips,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Superinfer: Slo-aware rotary scheduling and memory management for llm inference on superchips,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.821546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T14:37:20.763697Z digest=sha256:a242995ea5878eb08ae7d7414b6753357246e1824d7f86fb338005ef2f102602

Observation 9c7f6120-152e-4850-ad47-5545febf8105 · outbound

This paper cites FastServe: Iteration-Level preemptive scheduling for large language model inference,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving FastServe: Iteration-Level preemptive scheduling for large language model inference,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.842741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T14:37:20.754259Z digest=sha256:2cb55aaa6237eadd65750ee6853e53d423ba59d644d6640c29aded3996dcea00

Observation 3485f82c-ef46-496b-8e9d-8a9e3430313f · outbound

This paper cites Sglang: efficient execution of structured language model programs,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Sglang: efficient execution of structured language model programs,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.801320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T14:37:20.776514Z digest=sha256:8b100d8ebbd06aa636708d6f001d1c05b22644b1cc3ba14b3d0aad19165a1fd8

Observation d1cdf254-95b2-4983-8adf-992d781c83ce · outbound

This paper cites Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.791256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T14:37:20.779425Z digest=sha256:69dccd0ca50d5827bd770701348f5d1b5c936af61bed3ae1940c845ce2864304

Observation 1d212a09-726c-41d7-bfa2-d0fa4512fd9f · outbound

This paper cites PolyServe: Efficient Multi-SLO Serving at Scale.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving PolyServe: Efficient Multi-SLO Serving at Scale

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.782815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.782815Z digest=sha256:803b664161dc42606be6980c63a04aec3f68525e598e6e2f84c6e044aec7a45a

Observation db9f465e-5014-4f7e-b25d-6a60a126d96f · outbound

This paper cites Jitserve: Slo-aware llm serving with imprecise request information,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Jitserve: Slo-aware llm serving with imprecise request information,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.811217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T14:37:20.770173Z digest=sha256:77cb2580e44ea33ba3c2430893f6be3e364d3d0e70efe279ff8d1fcca55b1ae2

Observation a4fb85d8-adb6-42b8-aaff-e03ec7b03973 · outbound

This paper cites Available: https://arxiv.org/abs/2504.20068.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Available: https://arxiv.org/abs/2504.20068

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.773331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.773331Z digest=sha256:c0c498d9684ce821219b201109e6cddf614660c5932c8d2fd0cb6b480b894997

Observation 950d1db5-909a-46ba-b81d-2e1bfa2e58ac · outbound

This paper cites A Quantitative Measure Of Fairness And Discrimination For Resource Allocation In Shared Computer Systems.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving A Quantitative Measure Of Fairness And Discrimination For Resource Allocation In Shared Computer Systems

Reference 1998

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.682201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.682201Z digest=sha256:6c52b8c69f001a4fbdc4f60c70895d44207a4c6e72a68d0ecd46a19b9313461d

Observation 5490f1a0-917e-4e28-9f39-e4d82775af4e · outbound

This paper cites Prompt Cache: Modular Attention Reuse for Low-Latency Inference.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.662090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.662090Z digest=sha256:c99c031f96e08a850818783d90d19d237a807e4e08f092c0974c1e0c98611053

Observation 75f0e0c6-d426-4116-b53f-a4db4bc05a15 · outbound

This paper cites Available: https://arxiv.org/abs/2409.17264.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Available: https://arxiv.org/abs/2409.17264

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.631206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.631206Z digest=sha256:3621aab7ac171454cde9de2f8d1329d780ee57ab3d734d60f4064bc5ec87f1a5

Observation 84c5990a-d01b-4f37-a842-c87d29bd2cc4 · outbound

This paper cites SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.766968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.766968Z digest=sha256:1dae02c0b966207d1171b361abda2915c9491505ca5e7a335e5c1a2d6e2649bb

Pith citing papers

No inbound Pith citation observations are available.