Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:45:34.461813Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 7 inbound Pith citation observations for arXiv:2504.19867.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:45:34.461813Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-14T05:50:05.675450Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T04:57:38.712067Z
52 of 52 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 47776202-1c19-4d62-8f11-68028e9f4b96 · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Language models are few-shot learners
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 90f40132-4b7b-4c0b-b0b0-a0ae3f23260a · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Towards a Human-like Open-Domain Chatbot
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f200d716-a72a-42de-82ab-c576906a338f · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Recipes for building an open-domain chatbot
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e52a746-0e93-483f-9159-863bfba6b574 · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage CodeBERT: A Pre-Trained Model for Programming and Natural Languages
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d959efb-140c-4ec9-84e2-620e70a15be0 · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage IntelliCode Compose: Code Generation Using Transformer
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5609e358-bb64-4558-bd09-671d4a44dbd5 · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Sparks of Artificial General Intelligence: Early experiments with GPT-4
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2190ff8-a254-4a92-bd22-e6e080cf52f3 · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Training language models to follow instructions with human feedback
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a4363b6-5b7c-41bc-9644-f367071331d5 · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage GPT-4 Technical Report
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cf210e3-3eee-4bd9-b266-266bc5c4b840 · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage A Survey on Efficient Inference for Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc028040-1dc2-42bc-b63a-fc1a8eee89de · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Orca: A distributed serving system for transformer-based generative models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1ab96d66-c8f0-46d6-aaf3-ab7c9acf3998 · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Efficient memory management for large language model serving with pagedattention
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fe41a937-246d-4b2e-9c57-3ab601119849 · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Fastertransformer: About transformer related optimization, including bert, gpt
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8153ace8-1a71-4ff4-8aeb-3f593da7c1b1 · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f0a985d-a116-40b9-8d50-f3648ef4cc2f · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8fa9981-97ac-4d03-9c8a-c10d2a3bf1ce · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5aceb0f4-8fe5-42df-aa2d-d5328c90a820 · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Splitwise: Efficient generative llm inference using phase splitting
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 53417fe7-6d49-414e-b39d-1ecb77c9aad1 · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 605ab4f6-e493-4c55-8681-8fe3fc915257 · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Exegpt: Constraint-aware resource scheduling for llm inference
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6d18cfe1-2ed3-468e-a366-805219e82cad · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Mooncake: A kvcache-centric disaggregated architecture for llm serving
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2f0d0fd9-ae99-4432-9916-015faae932ce · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Nvlink and nvlink switch, 2024
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 19a7e791-1d27-4c19-b003-5b9ce4d540e3 · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage LLaMA: Open and Efficient Foundation Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61a5b929-9972-443c-a059-0060fc856a34 · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Attention is all you need
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a8f64235-01cc-41b8-83f2-45ac5d5ac413 · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 025766aa-b84a-4062-961a-d9611f355298 · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Gulavani, Alexey Tumanov, , and Ramachandran Ramjee
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 787b0509-0962-4de6-a495-60f5a6655c8e · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Loongserve: Efficiently serving long-context large language models with elastic sequence parallelism
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 74289d5d-3bcd-448a-a81d-db02c871eaec · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage SGLang: Efficient Execution of Structured Language Model Programs
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82e0b917-d1d3-449f-b5cc-63c673e43468 · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Optimizing inference on large language models with nvidia tensorrt-llm, now publicly available
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 49ce9b54-90b2-4672-9260-5257d1ed35d5 · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Lmdeploy is a toolkit for compressing, deploying, and serving llms
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation af0b6ff2-100a-484a-b98e-e1f283babc91 · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c86c13c-95ee-470e-9758-15b75b97e875 · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Fast Distributed Inference Serving for Large Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3703adc7-a210-4eb9-9430-454b655b1f06 · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Gonzalez, and Ion Stoica
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 12fe33a9-2ade-4f83-ae34-b7a892e9bb2b · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Efficient LLM Scheduling by Learning to Rank
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9adcf7f8-bb58-409d-857e-a029ef0afc5d · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage NanoFlow: Towards Optimal Large Language Model Serving Throughput
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb5fe41f-b203-4c5d-8b3b-4ca9ff7ddfeb · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Flashattention: Fast and memory-efficient exact attention with io-awareness
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0bbe9b5e-f846-4403-8b22-2c96df9ca148 · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Flashdecoding++: Faster large language model inference on gpus
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 71241fdc-d701-47f7-b3b0-1a1919dc4262 · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0e18ebe6-3235-46da-9089-1787c765d9d3 · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Nvidia mps, 2022
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1b218a1e-ec8d-4f71-9fac-6b5116b08159 · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Inter-process communication
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a1e3dcc0-d14c-4ac3-8596-67cceaa20d7b · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Fundamentals of queueing theory, volume
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 32c97de5-8db4-4ec2-aaf5-ab40b1ca800f · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Sharegpt
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5fff933f-4759-4a93-b4ec-aca5dcb6429d · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Nvidia dynamo: A datacenter scale distributed inference serving framework
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3b079a6a-c177-4b00-a078-dcecb026a0dc · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50bb15ee-baaf-4d75-a44c-bc24ac1745a2 · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage The Llama 3 Herd of Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce818738-dcb8-41e4-8fe0-d87c25ae686b · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Cuda toolkit
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 15cd71d6-93a4-42bc-ab0a-8f4e04ee8476 · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Efficient large message broadcast using nccl and cuda-aware mpi for deep learning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9b40d1aa-b6cb-4fd6-a00a-c2d648a2434c · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Pytorch: An imperative style, high-performance deep learning library
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77bb8015-bfae-4dc5-9d35-b7e2466b5ae5 · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Introducing meta llama 3: The most capable openly available llm to date, April 2024
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation cecaecdf-2c08-498a-a75b-57767c79f289 · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c630acf1-ec20-4f47-8cb6-5a26ad0f6a86 · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage DeepSeek-V3 Technical Report
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dea17199-d56d-4780-a656-f998876d9994 · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 321e6f93-1aec-497f-9af9-4b890bdf8f8c · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Math-500
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 59b81587-048a-473c-a81b-0722d232172f · outbound
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Unresolved cited work
Reference 399
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fd7d91d0-44e7-4401-95ab-fb53c88276eb · inbound
Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 633d21dc-2e09-4afa-a53e-f7dfe3ee0580 · inbound
DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2ba28d7-9b2c-4bf3-9bb3-e753026a5fc8 · inbound
Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e3cf6b4c-9744-43e2-b0f6-bc0407270db5 · inbound
KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 249612ad-faab-49d3-b2e2-9277affae70a · inbound
FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 478a67f2-2dab-4a49-9434-6844a04c21d7 · inbound
Prefilling-dLLM: Predictive Prefilling for Long-Context Inference in Diffusion Language Models semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3c53029e-3559-4473-98af-3781038d5753 · inbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.