Pith. sign in

Paper Citation Record · LEDGER

LLM Serving Optimization with Variable Prefill and Decode Lengths

As of 15 August 2026, this Paper Citation Record lists 2 of 2 outbound references and 11 inbound Pith citation observations for arXiv:2508.06133.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.06133 v4

Coverage vector

measured 2 of 2 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T22:59:19.872813Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:48:23.929970Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T08:39:41.944225Z

Reference resolution

2 of 2 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 638c69ce-eb7c-43c9-be03-16cdedd8f489 · outbound

This paper cites SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models.

LLM Serving Optimization with Variable Prefill and Decode Lengths SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models

Reference 2

Resolution
malformed identifier
no resolver link, observed 2026-08-05T22:59:19.872813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:59:19.872813Z digest=sha256:f38d9912606adc1bec81f7f4605b1bad939892eedfb86f05c551be37710a4ae5

Observation 5321b12d-c7cb-4b99-b21f-f664bdd575c0 · outbound

This paper cites Stochastic Monkeys at Play: Random Augmentations Cheaply Break LLM Safety Alignment.

LLM Serving Optimization with Variable Prefill and Decode Lengths Stochastic Monkeys at Play: Random Augmentations Cheaply Break LLM Safety Alignment

Reference 2024

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T22:59:20.132431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T22:59:19.726224Z digest=sha256:b5d9af8f63703f765971c34265374cc06086f28a44dfae79a09194b5584c4b46

Pith citing papers

Observation 8fbed5c4-d278-4e34-9a12-ac32dbf7e135 · inbound

Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints cites this paper.

Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints LLM Serving Optimization with Variable Prefill and Decode Lengths

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-30T02:16:09.189250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-22T19:44:35.756018Z digest=sha256:42a5ed92363758881c74e15d390cb6fe8f644b20ce04c88885c8f76aecfc2fa8

Observation 3f4fb693-23e0-43ed-b0ab-89bd60590180 · inbound

SAM Encoder Breach by Adversarial Simplicial Complex Triggers Downstream Model Failures cites this paper.

SAM Encoder Breach by Adversarial Simplicial Complex Triggers Downstream Model Failures LLM Serving Optimization with Variable Prefill and Decode Lengths

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T22:55:55.448797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:55:55.448797Z digest=sha256:90ee049de4856637b36695e6a739676c07825097f4079a9488cd71d1bab2ce9b

Observation de56bfa0-519f-4f09-86af-2d8fa95a00e3 · inbound

Flow-Controlled Scheduling for LLM Inference with Provable Stability Guarantees cites this paper.

Flow-Controlled Scheduling for LLM Inference with Provable Stability Guarantees LLM Serving Optimization with Variable Prefill and Decode Lengths

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T02:16:09.189250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:51:11.343200Z digest=sha256:0f227fb81ecfc377ef0cf925d0d5f587d04cea88c748491f0189b7964e60a08e

Observation 48e8820b-e5c1-4798-b0ed-61f1b997a5ab · inbound

Position: LLM Serving Needs Mathematical Optimization and Algorithmic Foundations, Not Just Heuristics cites this paper.

Position: LLM Serving Needs Mathematical Optimization and Algorithmic Foundations, Not Just Heuristics LLM Serving Optimization with Variable Prefill and Decode Lengths

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T02:16:09.189250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-09T18:53:29.089882Z digest=sha256:86c118b241f58097a4973e858d9764f044c3067ca8baa30acd742c2fc6e98d82

Observation 580b7487-c84c-4279-956f-9af48ae0fddb · inbound

A Queueing-Theoretic Framework for Stability Analysis of LLM Inference with KV Cache Memory Constraints cites this paper.

A Queueing-Theoretic Framework for Stability Analysis of LLM Inference with KV Cache Memory Constraints LLM Serving Optimization with Variable Prefill and Decode Lengths

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T02:16:09.189250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-08T17:28:05.492863Z digest=sha256:071dacb0ff0ae8c165c41f4f4a374a4a6fc2cfa09bd072084c5cafc5ae2b0fe9

Observation 44518a3f-da22-4fcf-a503-ffab3c8ead8a · inbound

Tackling the Data-Parallel Load Balancing Bottleneck in LLM Serving: Practical Online Routing at Scale cites this paper.

Tackling the Data-Parallel Load Balancing Bottleneck in LLM Serving: Practical Online Routing at Scale LLM Serving Optimization with Variable Prefill and Decode Lengths

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T02:16:09.189250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-08T05:22:14.122156Z digest=sha256:6a6d2a8f819f0cfa3abfb5a09c48fafbda23195af209d31fae158fc651b675a2

Observation 1bfc7601-42bb-439c-91e2-0b68510eb68f · inbound

Tackling the Data-Parallel Load Balancing Bottleneck in LLM Serving: Practical Online Routing at Scale cites this paper.

Tackling the Data-Parallel Load Balancing Bottleneck in LLM Serving: Practical Online Routing at Scale LLM Serving Optimization with Variable Prefill and Decode Lengths

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T02:16:09.189250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-11T00:59:57.639525Z digest=sha256:89a1d9399c31335a79db4d9532e30cd98d3c10aedf465bbbb674901e929357bb

Observation 44cf0a7c-fd4b-4bed-bedf-18c5ecf0e7f8 · inbound

Geometry-Aware Online Scheduling for LLM Serving: From Theoretical Bound to System Practice cites this paper.

Geometry-Aware Online Scheduling for LLM Serving: From Theoretical Bound to System Practice LLM Serving Optimization with Variable Prefill and Decode Lengths

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T08:39:41.945441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T11:17:35.860286Z digest=sha256:87ea487cc04767c73d8ece03c8671480e634c15229765a08bd1e47a90b0b5a3a

Observation 8e3a61ed-affe-44a7-bf55-0e7f4295ef5c · inbound

General Non-Clairvoyant KV-Cache Scheduling via Regime-Aware Routing cites this paper.

General Non-Clairvoyant KV-Cache Scheduling via Regime-Aware Routing LLM Serving Optimization with Variable Prefill and Decode Lengths

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-13T04:25:48.447055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T04:25:48.447055Z digest=sha256:9d760bfa45350d2f36c8fc0cd83d3a5df1939e5858f19ff91e6adfbe48b65c60

Observation 7fb6e405-dcab-4fb9-be5a-386235643ee9 · inbound

Mixture-of-Experts Serving cites this paper.

Mixture-of-Experts Serving LLM Serving Optimization with Variable Prefill and Decode Lengths

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T15:48:23.929970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:48:23.929970Z digest=sha256:4dcf4f4bccee06967ef81f6b01f808c63aac6d0cc4b09b3287ef0f9f1ef428df

Observation a0baf015-af7d-44f8-a8b5-2be04de56ed9 · inbound

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling cites this paper.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling LLM Serving Optimization with Variable Prefill and Decode Lengths

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:03.079660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:03.079660Z digest=sha256:094f0cc4697d4433ef0fdbb8905d5fabe8c4dca181643dd6e08e0c3f3ffda90f