Pith. sign in

Paper Citation Record · LEDGER

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage

As of 16 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 7 inbound Pith citation observations for arXiv:2504.19867.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.19867 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:45:34.461813Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T05:50:05.675450Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:57:38.712067Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved24
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 47776202-1c19-4d62-8f11-68028e9f4b96 · outbound

This paper cites Language models are few-shot learners.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Language models are few-shot learners

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:35.169537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:45:34.266574Z digest=sha256:e768bb5b3a1b82ca98653874aadb1789ccd42828318a1eb3adc4b8cd11a93505

Observation 90f40132-4b7b-4c0b-b0b0-a0ae3f23260a · outbound

This paper cites Towards a Human-like Open-Domain Chatbot.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Towards a Human-like Open-Domain Chatbot

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.270931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.270931Z digest=sha256:5c37acf4b4dfcb8dd73ed790ccdfdbc04bb6aba5e0405a6e3c0305128cdd6055

Observation f200d716-a72a-42de-82ab-c576906a338f · outbound

This paper cites Recipes for building an open-domain chatbot.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Recipes for building an open-domain chatbot

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.275409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.275409Z digest=sha256:97439ff4fd72c375f77828b5825c249464c75164436343ba29106ef18938158a

Observation 1e52a746-0e93-483f-9159-863bfba6b574 · outbound

This paper cites CodeBERT: A Pre-Trained Model for Programming and Natural Languages.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage CodeBERT: A Pre-Trained Model for Programming and Natural Languages

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.279420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.279420Z digest=sha256:6e44b71f15980a9c4bc3052905fd386a4f46452ce9615e488a1da1789a96f77f

Observation 9d959efb-140c-4ec9-84e2-620e70a15be0 · outbound

This paper cites IntelliCode Compose: Code Generation Using Transformer.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage IntelliCode Compose: Code Generation Using Transformer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.283941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.283941Z digest=sha256:692f5efa749b893901a786b8456c7d34b9ed8b2209b45b8a1868eff528791863

Observation 5609e358-bb64-4558-bd09-671d4a44dbd5 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.288668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.288668Z digest=sha256:491cf5d50971a88103b651d6f687355867a7045030d0b95e1f708633e273ed1b

Observation b2190ff8-a254-4a92-bd22-e6e080cf52f3 · outbound

This paper cites Training language models to follow instructions with human feedback.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Training language models to follow instructions with human feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.297218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.297218Z digest=sha256:7733d2ad4dc0cef01a947f4a8d1f2cd4cee9a4555fd80dc4da69c7354e788e57

Observation 8a4363b6-5b7c-41bc-9644-f367071331d5 · outbound

This paper cites GPT-4 Technical Report.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage GPT-4 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.301227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.301227Z digest=sha256:e4d6701b6223918ad5e66b761f4ee15c6be78087d1eee7c588f9f9dff040afa6

Observation 4cf210e3-3eee-4bd9-b266-266bc5c4b840 · outbound

This paper cites A Survey on Efficient Inference for Large Language Models.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage A Survey on Efficient Inference for Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.305534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.305534Z digest=sha256:e9ef738e278b1b5ba81556ac5c6cfe74f1caaf876325a4b7f70aa39bd72a93b2

Observation fc028040-1dc2-42bc-b63a-fc1a8eee89de · outbound

This paper cites Orca: A distributed serving system for transformer-based generative models.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Orca: A distributed serving system for transformer-based generative models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:35.158878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:45:34.310388Z digest=sha256:4a98694b70266b4720b57717b09e20f56a8073dcb8d9cdbd4713064c602f8adb

Observation 1ab96d66-c8f0-46d6-aaf3-ab7c9acf3998 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Efficient memory management for large language model serving with pagedattention

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:35.148157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:45:34.314009Z digest=sha256:892735009c0073189d672506f61177ee0f2c1c87673a3932646c2540e9f4f194

Observation fe41a937-246d-4b2e-9c57-3ab601119849 · outbound

This paper cites Fastertransformer: About transformer related optimization, including bert, gpt.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Fastertransformer: About transformer related optimization, including bert, gpt

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:35.137285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:45:34.317680Z digest=sha256:20490f8cb01e32ac8149f57d378d320ce9f90070694b922d48dc764c6a081152

Observation 8153ace8-1a71-4ff4-8aeb-3f593da7c1b1 · outbound

This paper cites SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.321242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.321242Z digest=sha256:938279d920c2222e70b1b74baa398ddf9458ec61d1e4a0b51e90bea420787c01

Observation 2f0a985d-a116-40b9-8d50-f3648ef4cc2f · outbound

This paper cites DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.325191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.325191Z digest=sha256:9cf2bd0365e2f29ab21be573bc623fb56c3faa02266a73382b3929d1e7e22e37

Observation e8fa9981-97ac-4d03-9c8a-c10d2a3bf1ce · outbound

This paper cites Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.328992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.328992Z digest=sha256:ce0b52210a8ec9fea422bcaea55aadc9a2356dad08cf49d05583972191a2cca3

Observation 5aceb0f4-8fe5-42df-aa2d-d5328c90a820 · outbound

This paper cites Splitwise: Efficient generative llm inference using phase splitting.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Splitwise: Efficient generative llm inference using phase splitting

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:35.124854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:45:34.332824Z digest=sha256:d733365593c65e61342b46ba007b37858cf2ed50c07494bba0c508880914981b

Observation 53417fe7-6d49-414e-b39d-1ecb77c9aad1 · outbound

This paper cites Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:35.112281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:45:34.336304Z digest=sha256:f349f63df70450ff0ad5588c2e00ddbefa2d887523286ab8fbc04ef889202cb6

Observation 605ab4f6-e493-4c55-8681-8fe3fc915257 · outbound

This paper cites Exegpt: Constraint-aware resource scheduling for llm inference.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Exegpt: Constraint-aware resource scheduling for llm inference

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:35.099556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:45:34.339718Z digest=sha256:deb69fd1cd2a1914a514d9faa2458963a4ff555ac3865f0e7d9faf187d097794

Observation 6d18cfe1-2ed3-468e-a366-805219e82cad · outbound

This paper cites Mooncake: A kvcache-centric disaggregated architecture for llm serving.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Mooncake: A kvcache-centric disaggregated architecture for llm serving

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:35.085362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:45:34.342931Z digest=sha256:c61c92fa7bf583530d08fca763d5b79f2df66654be24b2daeb925b166b9b9216

Observation 2f0d0fd9-ae99-4432-9916-015faae932ce · outbound

This paper cites Nvlink and nvlink switch, 2024.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Nvlink and nvlink switch, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:35.072104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:45:34.346272Z digest=sha256:39c7d2ea67f72ad10d0714a97ce74c011a3f1d54126a29a1d71bf18b9f6cfae8

Observation 19a7e791-1d27-4c19-b003-5b9ce4d540e3 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage LLaMA: Open and Efficient Foundation Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.349353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.349353Z digest=sha256:b5d055ea20786007f7a6c0695b0090f37c94dab679a37701768cb6468fa77b82

Observation 61a5b929-9972-443c-a059-0060fc856a34 · outbound

This paper cites Attention is all you need.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Attention is all you need

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:35.058629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:45:34.352750Z digest=sha256:a5510eedf38a5faf1c90600cd66b1d5692eb8e9bd4e2c9b9503a7526f6957269

Observation a8f64235-01cc-41b8-83f2-45ac5d5ac413 · outbound

This paper cites an unresolved cited work.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.356125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.356125Z digest=sha256:b06c0bb961a269467fa74dd19d6e852b57c31bca7edafb9c747821812d9ee8a6

Observation 025766aa-b84a-4062-961a-d9611f355298 · outbound

This paper cites Gulavani, Alexey Tumanov, , and Ramachandran Ramjee.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Gulavani, Alexey Tumanov, , and Ramachandran Ramjee

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:35.037896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:45:34.359353Z digest=sha256:471c8fa0fcf4e37612ca0b2b56ab10ca69dbf1e32acb27b3009e3039c2001334

Observation 787b0509-0962-4de6-a495-60f5a6655c8e · outbound

This paper cites Loongserve: Efficiently serving long-context large language models with elastic sequence parallelism.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Loongserve: Efficiently serving long-context large language models with elastic sequence parallelism

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:35.025999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:45:34.362826Z digest=sha256:9bb0f8dfa7b995cc0f8610b0ef80fa71e094e08f8587c0247675e17f1c5da69e

Observation 74289d5d-3bcd-448a-a81d-db02c871eaec · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage SGLang: Efficient Execution of Structured Language Model Programs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.366590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.366590Z digest=sha256:7a6d0329eaf2845ac885bf132f0cdb47668d1a90cc99ceaca9f44c50032581b0

Observation 82e0b917-d1d3-449f-b5cc-63c673e43468 · outbound

This paper cites Optimizing inference on large language models with nvidia tensorrt-llm, now publicly available.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Optimizing inference on large language models with nvidia tensorrt-llm, now publicly available

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:35.014350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:45:34.370833Z digest=sha256:fca65e343af9432ac94b12d23c9084be5ad116e8840a52ba3bd2fc197b341040

Observation 49ce9b54-90b2-4672-9260-5257d1ed35d5 · outbound

This paper cites Lmdeploy is a toolkit for compressing, deploying, and serving llms.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Lmdeploy is a toolkit for compressing, deploying, and serving llms

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:34.891645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:45:34.374025Z digest=sha256:ffc5bb23003e3f2fa05c7fa72634a1dcbeeeab8b20fe5dbe066c1f7c80c1fd06

Observation af0b6ff2-100a-484a-b98e-e1f283babc91 · outbound

This paper cites MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.377394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.377394Z digest=sha256:7e158bc63e4e39e180dd8a33cdc06db6054000c105554efcfce8ce77243c4e0b

Observation 6c86c13c-95ee-470e-9758-15b75b97e875 · outbound

This paper cites Fast Distributed Inference Serving for Large Language Models.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Fast Distributed Inference Serving for Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.381122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.381122Z digest=sha256:e418065d532ce4af0823f855370d937c6748fbe83b0466f3e92407e51c44e2b7

Observation 3703adc7-a210-4eb9-9430-454b655b1f06 · outbound

This paper cites Gonzalez, and Ion Stoica.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Gonzalez, and Ion Stoica

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:34.880190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:45:34.384669Z digest=sha256:564b2198f975da6ea75f9baf7c1b5bdca1bcd993262b8ef4daef2e91b99cd4cd

Observation 12fe33a9-2ade-4f83-ae34-b7a892e9bb2b · outbound

This paper cites Efficient LLM Scheduling by Learning to Rank.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Efficient LLM Scheduling by Learning to Rank

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.388256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.388256Z digest=sha256:e8da1fdd70110251df1bb2be23e2b25b929e52310e82953895cc649c6a3079d2

Observation 9adcf7f8-bb58-409d-857e-a029ef0afc5d · outbound

This paper cites NanoFlow: Towards Optimal Large Language Model Serving Throughput.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage NanoFlow: Towards Optimal Large Language Model Serving Throughput

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.392125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.392125Z digest=sha256:5a0f16674d908980a8ddb9018fc71fec9ea0a3afe2ec2935a69ba96ae9160378

Observation fb5fe41f-b203-4c5d-8b3b-4ca9ff7ddfeb · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:34.868461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:45:34.396280Z digest=sha256:b7bf3291ce34d3a226c60495e110965043bf5b8ba500488c3b246d098bf26654

Observation 0bbe9b5e-f846-4403-8b22-2c96df9ca148 · outbound

This paper cites Flashdecoding++: Faster large language model inference on gpus.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Flashdecoding++: Faster large language model inference on gpus

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:34.856980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:45:34.400476Z digest=sha256:b32d941b10b975732db087b14fd5b3710a8647e72d844296c097403b73cc8431

Observation 71241fdc-d701-47f7-b3b0-1a1919dc4262 · outbound

This paper cites Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:34.846885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:45:34.405100Z digest=sha256:5e5e3750925b3c375728bc9fe3370e587a17c7e500c7c03f0e9889f3ff8b9920

Observation 0e18ebe6-3235-46da-9089-1787c765d9d3 · outbound

This paper cites Nvidia mps, 2022.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Nvidia mps, 2022

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:34.832646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:45:34.408628Z digest=sha256:a8433fc58f795723b503d0f8011a3aa5146d359e4988118f3e9b3ec5dd6f55e0

Observation 1b218a1e-ec8d-4f71-9fac-6b5116b08159 · outbound

This paper cites Inter-process communication.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Inter-process communication

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:34.819752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:45:34.412068Z digest=sha256:1f18f43f5915c331935d616be63132583984d0fea74d9812096ed1f13106a5df

Observation a1e3dcc0-d14c-4ac3-8596-67cceaa20d7b · outbound

This paper cites Fundamentals of queueing theory, volume.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Fundamentals of queueing theory, volume

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:34.807656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:45:34.415688Z digest=sha256:258c6c6764808a76e2cfb4e3c2cdc5e1f52939dd1183c66527efece097a1e208

Observation 32c97de5-8db4-4ec2-aaf5-ab40b1ca800f · outbound

This paper cites Sharegpt.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Sharegpt

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:34.784756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:45:34.423215Z digest=sha256:32dc0150bc26dcb87c611622074e7bc3fd0b21e5f67f1429e70336d8ea06bd61

Observation 5fff933f-4759-4a93-b4ec-aca5dcb6429d · outbound

This paper cites Nvidia dynamo: A datacenter scale distributed inference serving framework.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Nvidia dynamo: A datacenter scale distributed inference serving framework

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:34.774210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:45:34.426768Z digest=sha256:b1a866d6edf296e61077cc285bd3158498ec81ef582c028096e936673a1bd6b2

Observation 3b079a6a-c177-4b00-a078-dcecb026a0dc · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.430466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.430466Z digest=sha256:d4f7e83bb87649b0ee2e9a2d71bc1ddc71fba5a1af6729c70773b3804705676c

Observation 50bb15ee-baaf-4d75-a44c-bc24ac1745a2 · outbound

This paper cites The Llama 3 Herd of Models.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage The Llama 3 Herd of Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.433792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.433792Z digest=sha256:78c29706aee28237543c2bbb3ba37eb714047ab37d81ff99a7dc9643c0a6839e

Observation ce818738-dcb8-41e4-8fe0-d87c25ae686b · outbound

This paper cites Cuda toolkit.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Cuda toolkit

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:34.763472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:45:34.437398Z digest=sha256:b13da50da3b309a5a68e94123efc92574984b101c270d653714f5fb10a4484c8

Observation 15cd71d6-93a4-42bc-ab0a-8f4e04ee8476 · outbound

This paper cites Efficient large message broadcast using nccl and cuda-aware mpi for deep learning.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Efficient large message broadcast using nccl and cuda-aware mpi for deep learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:34.753873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:45:34.440731Z digest=sha256:410dfffc7110223e43284968246b1176b7b028663c9154ac45911112a6edb9a5

Observation 9b40d1aa-b6cb-4fd6-a00a-c2d648a2434c · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Pytorch: An imperative style, high-performance deep learning library

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.443972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.443972Z digest=sha256:8624c6b23c8f1e1eda8f0a88fe0afc5fc53805c789d93dc1e499e8bc049e4eac

Observation 77bb8015-bfae-4dc5-9d35-b7e2466b5ae5 · outbound

This paper cites Introducing meta llama 3: The most capable openly available llm to date, April 2024.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Introducing meta llama 3: The most capable openly available llm to date, April 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:34.737330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:45:34.447264Z digest=sha256:41cc5d70d4dcd503ad3946fafa25eb8b623e81edc9029881bfdeb6f2a6d16a8b

Observation cecaecdf-2c08-498a-a75b-57767c79f289 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.450930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.450930Z digest=sha256:067b0ef360d8be17eee4d657822f80962217c50b86d6288a1f421981411144f2

Observation c630acf1-ec20-4f47-8cb6-5a26ad0f6a86 · outbound

This paper cites DeepSeek-V3 Technical Report.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage DeepSeek-V3 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.454577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.454577Z digest=sha256:b8790ceb7db430d1c303f91be5b6d96067698aaa7f65e9139bc0d986c4cb80ce

Observation dea17199-d56d-4780-a656-f998876d9994 · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.458370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.458370Z digest=sha256:5e748ab2d861ff552ada39e453b01df6c4cae4a442a4ced3ee9e5fb72479b5e6

Observation 321e6f93-1aec-497f-9af9-4b890bdf8f8c · outbound

This paper cites Math-500.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Math-500

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:34.725530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:45:34.461813Z digest=sha256:32e20a341443a783f5b8b15e992a5b6d9e8fede7347305cb6f72a45b468e6898

Observation 59b81587-048a-473c-a81b-0722d232172f · outbound

This paper cites an unresolved cited work.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Unresolved cited work

Reference 399

Resolution
parse uncertain
raw_fallback, observed 2026-08-16T05:45:34.795365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:45:34.419634Z digest=sha256:352a54434d9481e1403803923dcb607ebd17741f67955950abf5fbebab489ba1

Pith citing papers

Observation fd7d91d0-44e7-4401-95ab-fb53c88276eb · inbound

Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving cites this paper.

Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:03:52.902518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:03:52.902518Z digest=sha256:88d8a57003b97bb539070941457f5f7f53c5885fbb812d5579f3405fa8b48e38

Observation 633d21dc-2e09-4afa-a53e-f7dfe3ee0580 · inbound

DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing cites this paper.

DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T23:39:06.362657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:39:06.362657Z digest=sha256:8b37a645022f0945dd27af99ab566d8a6bd13888a681738f15d502d3b5a6fe3c

Observation e2ba28d7-9b2c-4bf3-9bb3-e753026a5fc8 · inbound

Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities cites this paper.

Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:16:10.262035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-08T09:45:57.201837Z digest=sha256:674fadffc6ca610f63bd4ebc86eed9d4a471b74aa05a77b8b2597090f5ee4651

Observation e3cf6b4c-9744-43e2-b0f6-bc0407270db5 · inbound

KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving cites this paper.

KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-14T17:42:31.180388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-14T17:42:06.870534Z digest=sha256:37315cb161060d3c78c82771d9ba69cfbdc4a86d869829a2aa21a6f07481fbf4

Observation 249612ad-faab-49d3-b2e2-9277affae70a · inbound

FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location cites this paper.

FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T10:46:52.308744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T04:57:08.746551Z digest=sha256:ddce9566b8a5a1459f2aa5b97e044561bb1e2deab431fd61c09a911b8d758cc1

Observation 478a67f2-2dab-4a49-9434-6844a04c21d7 · inbound

Prefilling-dLLM: Predictive Prefilling for Long-Context Inference in Diffusion Language Models cites this paper.

Prefilling-dLLM: Predictive Prefilling for Long-Context Inference in Diffusion Language Models semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:57:38.713663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-27T13:30:19.620692Z digest=sha256:a355968b345d3093a3ba1da741d49aa1d152aca6c968bc60c6f0140dbfa63bb9

Observation 3c53029e-3559-4473-98af-3781038d5753 · inbound

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving cites this paper.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T05:50:05.675450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T05:50:05.675450Z digest=sha256:324c895080c8e40c8480e4072647cc5cd14a5641e77c7b95cd4ccef59b257a5e