Pith. sign in

Paper Citation Record · LEDGER

Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2401.11181.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.11181 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T18:02:12.366616Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

6
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cd8da1e9-bc32-448d-bd9d-4a7ffbce1000 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 273

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:39:33.469145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:2ff1f7f840321613fb1759eea57a89f95f0efe9a7956f8984469df52d742664c

Observation 7f3734cb-f252-4ab8-8858-7d7df935798f · inbound

Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs cites this paper.

Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T18:02:12.366616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:02:12.366616Z digest=sha256:ceb409f4a551f1ff510cb4dabc00499047dca6fc0bab4ac3dc862372e3058d86

Observation 7c56cfd0-de8d-4781-a1f1-ec29bb3f2ef0 · inbound

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference cites this paper.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T04:32:01.206880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:32:01.206880Z digest=sha256:bff5d9c9358400e86b73a9c7914dbd4c53ff15295ece507a5e5708e818a8e706

Observation 24f5124e-0c28-4768-96ea-f48d73ec1a13 · inbound

HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment cites this paper.

HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 1995

Resolution
unresolved
no resolver link, observed 2026-08-08T11:32:21.510038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:32:21.510038Z digest=sha256:15e0b7fdbec99318dad8f22cefbb7bdb73a447995c6110986c5092fb16d81805

Observation 7c6ccbed-f64c-4d5b-8fd7-4062c9ac36e2 · inbound

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees cites this paper.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.887709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.887709Z digest=sha256:da727c8cd58e2a1a221b16a1e98ba5396a36399488cacdbee45906237f1b0f1c

Observation bf3a698b-6d10-4bad-bedf-03331e1c72c7 · inbound

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization cites this paper.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T18:55:04.330334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:55:04.330334Z digest=sha256:10b449ec30fcd7a759dfdb1fe9b1a5fc7e7944fd2283fb40399f92c96d46564a

Observation d539e8c3-45f1-43ce-a2b0-9593364440c7 · inbound

Beyond the Buzz: A Pragmatic Take on Inference Disaggregation cites this paper.

Beyond the Buzz: A Pragmatic Take on Inference Disaggregation Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:26:32.736179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:32.736179Z digest=sha256:8ec248636e6d4e68285d30d289a43aaae97fb7563571645cc31c7777d7324a7c

Observation 5f148349-ce14-4b64-8106-9106552df42e · inbound

Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving cites this paper.

Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:03:53.051327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:03:53.051327Z digest=sha256:7428f299af2edfdc1b8d4828f0c2e8d82ecf4b92efd085791382b439195372a2

Observation f33f8277-8b6a-4991-b6f4-b43908aa318b · inbound

Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference cites this paper.

Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:36.345043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:36.345043Z digest=sha256:993899c3af087e3db30d76e5935cc363d30acc5cdb2576956a7bdd4bb4e3b805

Observation 1f96ef1a-0a70-4f5b-96ed-3cd4aff33490 · inbound

Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism cites this paper.

Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T20:53:31.939770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:53:31.939770Z digest=sha256:d8faae2b5e54bcf98c94400ab1378346a8af24421c411e19b89bca4962fc815a

Observation 26ad8cfa-5076-4b7b-b8bd-67eb2edd7fbc · inbound

FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters cites this paper.

FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:11:03.617983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T07:07:46.228491Z digest=sha256:7ae15df83e183f38611ff94e7bdfb9c22ebba23e4309453e1c486d7fd1a34d53

Observation 1a795543-9eae-427b-8f11-15ae5f4ad8c2 · inbound

STAR: Decode-Phase Rescheduling for LLM Inference cites this paper.

STAR: Decode-Phase Rescheduling for LLM Inference Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T06:12:25.938366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T06:11:28.120860Z digest=sha256:027965bd2381ea58c9ec849db8aa2fa507404e3c82ce383585a16a103574cfe5

Observation 0af25593-fd88-4ed5-954b-b2bef949192e · inbound

ODMA: On-Demand Memory Allocation Strategy for LLM Serving on LPDDR-Class Accelerators cites this paper.

ODMA: On-Demand Memory Allocation Strategy for LLM Serving on LPDDR-Class Accelerators Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:58:43.005876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T23:54:08.052482Z digest=sha256:f73be0d760e17f8b37a5a2ef3746c0859fe1abe93ac847446fd4c29d292777f9

Observation d20fdeaf-85f2-4752-af64-7f07217c58e6 · inbound

Efficient Multi-round LLM Inference over Disaggregated Serving cites this paper.

Efficient Multi-round LLM Inference over Disaggregated Serving Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 1994

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:44.283021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:44.283021Z digest=sha256:903be669973be844f32c950931c3cc6331be1f158c4a0621f5a458329235be6b

Observation 1ce18c33-284a-4fde-b632-ee06b6f0368e · inbound

Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUs cites this paper.

Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUs Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:01:07.014104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T16:45:39.853764Z digest=sha256:43f1ea49082d7c236d56e2df8a13a14b6a76ea41b479ad97f5c4d7b231d2a7ee

Observation 193c6fde-95d1-4e3c-bdc9-5cfc75a66375 · inbound

VeriCache: Turning Lossy KV Cache into Lossless LLM Inference cites this paper.

VeriCache: Turning Lossy KV Cache into Lossless LLM Inference Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-19T22:12:51.167065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T22:09:15.782104Z digest=sha256:84ecfa8b3791272af8f78fe69d78c5b5af2d1cc3f24475409d20628a701434bf

Observation 93fbeaed-f880-4c22-88d8-a296d194c9da · inbound

AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System cites this paper.

AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:15:17.288802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T03:12:49.028342Z digest=sha256:627be1e426b4a1adc0662eccd7de9766728b11da2adec29ce16cf04d53d9bd85

Observation 34fb4efc-43e1-4098-ae80-ff6ca91927d7 · inbound

Human-Less LLM Serving: Quantifying the Human Tax on Throughput cites this paper.

Human-Less LLM Serving: Quantifying the Human Tax on Throughput Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:45:11.801306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T00:44:55.517266Z digest=sha256:ea16e2995766ed4f2a1babde5433298599f3acbc719cef7b41eac5f56345c61d

Observation 66c5f74e-e132-435a-8a50-d755ca4e86e2 · inbound

Geometry-Aware Online Scheduling for LLM Serving: From Theoretical Bound to System Practice cites this paper.

Geometry-Aware Online Scheduling for LLM Serving: From Theoretical Bound to System Practice Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:39:41.985966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T11:17:35.860286Z digest=sha256:7d7080b4a9cbb564849e2f19515d7317872b3cbecff6e6bbc907e0567ec4984a

Observation 91619bb5-87eb-472e-8b08-6cf731256fcb · inbound

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving cites this paper.

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T14:04:45.341416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T05:37:13.211613Z digest=sha256:f790900ac58dbfcb1b6721d3c6dd0aae128ce28249337dd197f8f1a094239bae

Observation 04af9445-3e67-4eb8-b5ea-8d59b1fe8549 · inbound

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving cites this paper.

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T08:55:35.143870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T07:06:53.318182Z digest=sha256:2ec9428547a0762081550741363429d4d380d9eb9ab6800bbe6c162019465124

Observation 9045a234-d428-4ed8-947a-61dad41f26e5 · inbound

Towards Load-Aware Prefill Deflection for Disaggregated LLM Serving cites this paper.

Towards Load-Aware Prefill Deflection for Disaggregated LLM Serving Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-03T06:07:40.890885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-03T06:05:06.467649Z digest=sha256:52fe34f1dc86fcca98b3b88541f0b9b505b50890e8d0577c110f1ccc1738d0be

Observation ac815cd7-9513-4cb1-b790-b1cb655a2309 · inbound

CoCoScale: Leveraging Layer-wise Scaling to Unlock the Potential of Online LLM Serving cites this paper.

CoCoScale: Leveraging Layer-wise Scaling to Unlock the Potential of Online LLM Serving Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-11T21:08:24.159706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:08:24.159706Z digest=sha256:698083bdd1563d23a4a1fcf57939027f238cd050288697209cd48f6262c18b5a

Observation ebab7f57-b3e2-4028-9376-3b53766a92b9 · inbound

Sangam: Efficiently Serving Diffusion LLMs with the AR Stack cites this paper.

Sangam: Efficiently Serving Diffusion LLMs with the AR Stack Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-11T20:58:04.182764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:58:04.182764Z digest=sha256:b06f286cf24454fa35a235391c698795176f6e31ddc20b73665cda6b4166abd5

Observation 25516814-d561-4c3a-aa0a-af61bd4072d0 · inbound

AutoSLO: Practical Latency SLOs on Cloud Data Warehouses -- Extended Version cites this paper.

AutoSLO: Practical Latency SLOs on Cloud Data Warehouses -- Extended Version Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T03:18:47.408836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T03:18:47.408836Z digest=sha256:1d30033cab0cc1726e01a5052f3008f272138e8c3d7e73417c16ce697bc7cf05

Observation 88f9298f-b660-4f3d-8b42-61c58cd035e9 · inbound

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer cites this paper.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:14.862617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:14.862617Z digest=sha256:012b71470cc0c33bd2f0b5233214ce2a934f7560865d5db539da78dbee2407a0

Observation b49bd378-5512-4d7b-b8f6-08a197292084 · inbound

Rethinking AI Cloud Infrastructure for Agentic Serving Systems with the Aries Experimentation Framework cites this paper.

Rethinking AI Cloud Infrastructure for Agentic Serving Systems with the Aries Experimentation Framework Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T14:14:10.457557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:14:10.457557Z digest=sha256:c3e92771ceeb5bf948a4b9553fe120d6aebbc72223b34febeeb946d0cc814411

Observation 7d2fc877-c3e9-4a5b-9681-8eb28e071b5f · inbound

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling cites this paper.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:22.951059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:22.951059Z digest=sha256:bbd105747f79fe80dab13e1c2ab899a4e09d752ebe15b647a534128b5e399a5a

Observation ea77d351-5772-4605-b59d-d2342eac07e3 · inbound

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure cites this paper.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.353271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.353271Z digest=sha256:ac796e551d8dcbe0557f8de206ace2e7dc37e3465b2773ffc22ce662de4f4c2c