Pith. sign in

Paper Citation Record · LEDGER

Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2401.11181.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.11181 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T18:02:12.366616Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

6
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cd8da1e9-bc32-448d-bd9d-4a7ffbce1000 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 273

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:39:33.469145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:d80fefc5e692192bd84be2384970977f4a5092633fde2fa0e8230d9a48ee5b13

Observation 7f3734cb-f252-4ab8-8858-7d7df935798f · inbound

Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs cites this paper.

Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T18:02:12.366616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:02:12.366616Z digest=sha256:806750c64303b13ac316361417e0715bb4f83dec53dca108642cb8b177bb6ac4

Observation 7c56cfd0-de8d-4781-a1f1-ec29bb3f2ef0 · inbound

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference cites this paper.

HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T04:32:01.206880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:32:01.206880Z digest=sha256:1f39345b5d0081accebce9a58111890a98f92b35c7834d5055774cad8678cae6

Observation 24f5124e-0c28-4768-96ea-f48d73ec1a13 · inbound

HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment cites this paper.

HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 1995

Resolution
unresolved
no resolver link, observed 2026-08-08T11:32:21.510038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:32:21.510038Z digest=sha256:2fd50d82ba99feebff6bef2b970f8729b8a68b843ea336358b563548a417ceb4

Observation 7c6ccbed-f64c-4d5b-8fd7-4062c9ac36e2 · inbound

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees cites this paper.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.887709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.887709Z digest=sha256:9e4c4c0625f6b219b3571b69196b905507c12e8dd5a6711809dfb131f36b8243

Observation bf3a698b-6d10-4bad-bedf-03331e1c72c7 · inbound

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization cites this paper.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T18:55:04.330334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:55:04.330334Z digest=sha256:10b449ec30fcd7a759dfdb1fe9b1a5fc7e7944fd2283fb40399f92c96d46564a

Observation d539e8c3-45f1-43ce-a2b0-9593364440c7 · inbound

Beyond the Buzz: A Pragmatic Take on Inference Disaggregation cites this paper.

Beyond the Buzz: A Pragmatic Take on Inference Disaggregation Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:26:32.736179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:32.736179Z digest=sha256:8ec248636e6d4e68285d30d289a43aaae97fb7563571645cc31c7777d7324a7c

Observation 5f148349-ce14-4b64-8106-9106552df42e · inbound

Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving cites this paper.

Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:03:53.051327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:03:53.051327Z digest=sha256:7428f299af2edfdc1b8d4828f0c2e8d82ecf4b92efd085791382b439195372a2

Observation f33f8277-8b6a-4991-b6f4-b43908aa318b · inbound

Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference cites this paper.

Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:36.345043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:36.345043Z digest=sha256:993899c3af087e3db30d76e5935cc363d30acc5cdb2576956a7bdd4bb4e3b805

Observation 1f96ef1a-0a70-4f5b-96ed-3cd4aff33490 · inbound

Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism cites this paper.

Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T20:53:31.939770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:53:31.939770Z digest=sha256:d8faae2b5e54bcf98c94400ab1378346a8af24421c411e19b89bca4962fc815a

Observation 26ad8cfa-5076-4b7b-b8bd-67eb2edd7fbc · inbound

FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters cites this paper.

FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:11:03.617983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T07:07:46.228491Z digest=sha256:032232181692f2e3bd6d5ae1890609e090611224c931522d39343d333f7748d1

Observation 1a795543-9eae-427b-8f11-15ae5f4ad8c2 · inbound

STAR: Decode-Phase Rescheduling for LLM Inference cites this paper.

STAR: Decode-Phase Rescheduling for LLM Inference Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T06:12:25.938366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T06:11:28.120860Z digest=sha256:8d1744d5f83578e09affb45fe537f7aad954833a91ff50f287dd667024824df0

Observation 0af25593-fd88-4ed5-954b-b2bef949192e · inbound

ODMA: On-Demand Memory Allocation Strategy for LLM Serving on LPDDR-Class Accelerators cites this paper.

ODMA: On-Demand Memory Allocation Strategy for LLM Serving on LPDDR-Class Accelerators Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:58:43.005876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T23:54:08.052482Z digest=sha256:54eb0e35f4c0ed4f560b1e7ded8075f54eb8302a2aff4fb6629890f0101e1b25

Observation d20fdeaf-85f2-4752-af64-7f07217c58e6 · inbound

Efficient Multi-round LLM Inference over Disaggregated Serving cites this paper.

Efficient Multi-round LLM Inference over Disaggregated Serving Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 1994

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:44.283021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:44.283021Z digest=sha256:903be669973be844f32c950931c3cc6331be1f158c4a0621f5a458329235be6b

Observation 1ce18c33-284a-4fde-b632-ee06b6f0368e · inbound

Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUs cites this paper.

Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUs Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:01:07.014104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T16:45:39.853764Z digest=sha256:c9ed0df674b45d542bcfdd96656288aae2a4a260c24b542dc328f8baf7596eb6

Observation 193c6fde-95d1-4e3c-bdc9-5cfc75a66375 · inbound

VeriCache: Turning Lossy KV Cache into Lossless LLM Inference cites this paper.

VeriCache: Turning Lossy KV Cache into Lossless LLM Inference Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-19T22:12:51.167065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T22:09:15.782104Z digest=sha256:882ddab767af7a38cea876ee388eb33306c84e27a6277d7e7a49b1dede7f2d8d

Observation 93fbeaed-f880-4c22-88d8-a296d194c9da · inbound

AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System cites this paper.

AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:15:17.288802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T03:12:49.028342Z digest=sha256:a49a628729acc7e9afb808b68b31537e5d484867f65ff53bee699a7a1a678631

Observation 34fb4efc-43e1-4098-ae80-ff6ca91927d7 · inbound

Human-Less LLM Serving: Quantifying the Human Tax on Throughput cites this paper.

Human-Less LLM Serving: Quantifying the Human Tax on Throughput Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:45:11.801306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T00:44:55.517266Z digest=sha256:ebaf6b2461a67fc3319828035fee3a4b5b2e1e4e62847ef6a29375328cdbf242

Observation 66c5f74e-e132-435a-8a50-d755ca4e86e2 · inbound

Geometry-Aware Online Scheduling for LLM Serving: From Theoretical Bound to System Practice cites this paper.

Geometry-Aware Online Scheduling for LLM Serving: From Theoretical Bound to System Practice Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:39:41.985966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T11:17:35.860286Z digest=sha256:f4234a209643b2895bced1a2f79a1e93d9f814bba61de58d99c37585ea30e6af

Observation 91619bb5-87eb-472e-8b08-6cf731256fcb · inbound

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving cites this paper.

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T14:04:45.341416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T05:37:13.211613Z digest=sha256:be2abf1666efdbec2ded0b244c95574c8e38aa9b25b76d59a91696a9436ba4f4

Observation 04af9445-3e67-4eb8-b5ea-8d59b1fe8549 · inbound

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving cites this paper.

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T08:55:35.143870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T07:06:53.318182Z digest=sha256:23da86efa3a3b3c7cabc11f69220386262dc2b46c78544cb168073c192e67323

Observation 9045a234-d428-4ed8-947a-61dad41f26e5 · inbound

Towards Load-Aware Prefill Deflection for Disaggregated LLM Serving cites this paper.

Towards Load-Aware Prefill Deflection for Disaggregated LLM Serving Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-03T06:07:40.890885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T06:05:06.467649Z digest=sha256:7c1ddfdf467da9b0c40823e24267daace3db343a24e1c7195848fa9fc59f7701

Observation ac815cd7-9513-4cb1-b790-b1cb655a2309 · inbound

CoCoScale: Leveraging Layer-wise Scaling to Unlock the Potential of Online LLM Serving cites this paper.

CoCoScale: Leveraging Layer-wise Scaling to Unlock the Potential of Online LLM Serving Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-11T21:08:24.159706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:08:24.159706Z digest=sha256:698083bdd1563d23a4a1fcf57939027f238cd050288697209cd48f6262c18b5a

Observation ebab7f57-b3e2-4028-9376-3b53766a92b9 · inbound

Sangam: Efficiently Serving Diffusion LLMs with the AR Stack cites this paper.

Sangam: Efficiently Serving Diffusion LLMs with the AR Stack Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-11T20:58:04.182764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:58:04.182764Z digest=sha256:864eb2458b5ef84e97415d3c56008b1586ddf12f2e06c3fa3c60fa599cfe0b6a

Observation 25516814-d561-4c3a-aa0a-af61bd4072d0 · inbound

AutoSLO: Practical Latency SLOs on Cloud Data Warehouses -- Extended Version cites this paper.

AutoSLO: Practical Latency SLOs on Cloud Data Warehouses -- Extended Version Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T03:18:47.408836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T03:18:47.408836Z digest=sha256:1d30033cab0cc1726e01a5052f3008f272138e8c3d7e73417c16ce697bc7cf05

Observation 88f9298f-b660-4f3d-8b42-61c58cd035e9 · inbound

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer cites this paper.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:14.862617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:14.862617Z digest=sha256:db327048d4a31beb38ed992ba841939bc96a6155b08c848c3c9c1e253c8e7aa0

Observation b49bd378-5512-4d7b-b8f6-08a197292084 · inbound

Rethinking AI Cloud Infrastructure for Agentic Serving Systems with the Aries Experimentation Framework cites this paper.

Rethinking AI Cloud Infrastructure for Agentic Serving Systems with the Aries Experimentation Framework Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T14:14:10.457557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:14:10.457557Z digest=sha256:c3e92771ceeb5bf948a4b9553fe120d6aebbc72223b34febeeb946d0cc814411

Observation 7d2fc877-c3e9-4a5b-9681-8eb28e071b5f · inbound

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling cites this paper.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:22.951059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:22.951059Z digest=sha256:bbd105747f79fe80dab13e1c2ab899a4e09d752ebe15b647a534128b5e399a5a

Observation ea77d351-5772-4605-b59d-d2342eac07e3 · inbound

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure cites this paper.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.353271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.353271Z digest=sha256:f35218830ed3fd16beb3563d83479776cffcf77c81229e36d3b069d18e7465f3