Pith. sign in

Paper Citation Record · LEDGER

LLM Serving Optimization with Variable Prefill and Decode Lengths

As of 9 August 2026, this Paper Citation Record lists 2 of 2 outbound references and 10 inbound Pith citation observations for arXiv:2508.06133.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.06133 v4

Coverage vector

measured 2 of 2 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T22:59:19.872813Z

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T22:55:55.448797Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T08:39:41.944225Z

Reference resolution

2 of 2 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 638c69ce-eb7c-43c9-be03-16cdedd8f489 · outbound

This paper cites SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models.

LLM Serving Optimization with Variable Prefill and Decode Lengths SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models

Reference 2

Resolution
malformed identifier
no resolver link, observed 2026-08-05T22:59:19.872813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:59:19.872813Z digest=sha256:b0142a72af3db2b96f4fc234f41b6cb77081576c329da68777e84189dfe35bf2

Observation 5321b12d-c7cb-4b99-b21f-f664bdd575c0 · outbound

This paper cites Stochastic Monkeys at Play: Random Augmentations Cheaply Break LLM Safety Alignment.

LLM Serving Optimization with Variable Prefill and Decode Lengths Stochastic Monkeys at Play: Random Augmentations Cheaply Break LLM Safety Alignment

Reference 2024

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T22:59:20.132431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T22:59:19.726224Z digest=sha256:73f506ab96789a9278e14ed6172197ae8e5548950796a6d7bc650286d23f7b6b

Pith citing papers

Observation 8fbed5c4-d278-4e34-9a12-ac32dbf7e135 · inbound

Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints cites this paper.

Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints LLM Serving Optimization with Variable Prefill and Decode Lengths

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-30T02:16:09.189250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T19:44:35.756018Z digest=sha256:be031470e728d6ad720617b37abe16a4fead2722ca37928039c0e85715411c17

Observation 3f4fb693-23e0-43ed-b0ab-89bd60590180 · inbound

SAM Encoder Breach by Adversarial Simplicial Complex Triggers Downstream Model Failures cites this paper.

SAM Encoder Breach by Adversarial Simplicial Complex Triggers Downstream Model Failures LLM Serving Optimization with Variable Prefill and Decode Lengths

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T22:55:55.448797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:55:55.448797Z digest=sha256:90ee049de4856637b36695e6a739676c07825097f4079a9488cd71d1bab2ce9b

Observation de56bfa0-519f-4f09-86af-2d8fa95a00e3 · inbound

Flow-Controlled Scheduling for LLM Inference with Provable Stability Guarantees cites this paper.

Flow-Controlled Scheduling for LLM Inference with Provable Stability Guarantees LLM Serving Optimization with Variable Prefill and Decode Lengths

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T02:16:09.189250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:51:11.343200Z digest=sha256:b6845a5fabe905029950558c573ce3ece3655a5a85150052bda1500e5e7579a7

Observation 48e8820b-e5c1-4798-b0ed-61f1b997a5ab · inbound

Position: LLM Serving Needs Mathematical Optimization and Algorithmic Foundations, Not Just Heuristics cites this paper.

Position: LLM Serving Needs Mathematical Optimization and Algorithmic Foundations, Not Just Heuristics LLM Serving Optimization with Variable Prefill and Decode Lengths

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T02:16:09.189250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T18:53:29.089882Z digest=sha256:4a573acffe6a916cf36e2395f52291ee615cf8d3c0e2d21204382655162a1f50

Observation 580b7487-c84c-4279-956f-9af48ae0fddb · inbound

A Queueing-Theoretic Framework for Stability Analysis of LLM Inference with KV Cache Memory Constraints cites this paper.

A Queueing-Theoretic Framework for Stability Analysis of LLM Inference with KV Cache Memory Constraints LLM Serving Optimization with Variable Prefill and Decode Lengths

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T02:16:09.189250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T17:28:05.492863Z digest=sha256:236256fdaa17a9e3c14ce170acaa5ca95a4d547ceadd4460bfdf9c5bf72802b1

Observation 44518a3f-da22-4fcf-a503-ffab3c8ead8a · inbound

Tackling the Data-Parallel Load Balancing Bottleneck in LLM Serving: Practical Online Routing at Scale cites this paper.

Tackling the Data-Parallel Load Balancing Bottleneck in LLM Serving: Practical Online Routing at Scale LLM Serving Optimization with Variable Prefill and Decode Lengths

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T02:16:09.189250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T05:22:14.122156Z digest=sha256:6143fa593ceb3008c003f51be82c253af1c1a89ddc3df1959c4617207fe7d47c

Observation 1bfc7601-42bb-439c-91e2-0b68510eb68f · inbound

Tackling the Data-Parallel Load Balancing Bottleneck in LLM Serving: Practical Online Routing at Scale cites this paper.

Tackling the Data-Parallel Load Balancing Bottleneck in LLM Serving: Practical Online Routing at Scale LLM Serving Optimization with Variable Prefill and Decode Lengths

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T02:16:09.189250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T00:59:57.639525Z digest=sha256:842d82bc5515ad1b47142e7275cebc57b94a0e871553bb01374482f1e41b39be

Observation 44cf0a7c-fd4b-4bed-bedf-18c5ecf0e7f8 · inbound

Geometry-Aware Online Scheduling for LLM Serving: From Theoretical Bound to System Practice cites this paper.

Geometry-Aware Online Scheduling for LLM Serving: From Theoretical Bound to System Practice LLM Serving Optimization with Variable Prefill and Decode Lengths

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T08:39:41.945441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T11:17:35.860286Z digest=sha256:7ff38e03442a8f7b21a52ae75f35d0b912fe11f09375e6093c17efb98555a1c5

Observation 8e3a61ed-affe-44a7-bf55-0e7f4295ef5c · inbound

General Non-Clairvoyant KV-Cache Scheduling via Regime-Aware Routing cites this paper.

General Non-Clairvoyant KV-Cache Scheduling via Regime-Aware Routing LLM Serving Optimization with Variable Prefill and Decode Lengths

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-13T04:25:48.447055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T04:25:48.447055Z digest=sha256:9f0ba8e2afb21eb30c3a7898cf838fcf3f51bee7390fbf4abe1c44953d8c51c0

Observation a0baf015-af7d-44f8-a8b5-2be04de56ed9 · inbound

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling cites this paper.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling LLM Serving Optimization with Variable Prefill and Decode Lengths

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:03.079660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:03.079660Z digest=sha256:094f0cc4697d4433ef0fdbb8905d5fabe8c4dca181643dd6e08e0c3f3ffda90f