Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:27:17.414952Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 3 inbound Pith citation observations for arXiv:2504.15720.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:27:17.414952Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-21T08:39:31.911497Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-21T08:39:53.273252Z
56 of 56 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7340987e-080f-449a-9a66-6ac0cdbe4ecc · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference https://github.com/NVIDIA/ FasterTransformer, 2019
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6137fcf8-0bf5-47cf-a180-f172b67436dc · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference https://grpc.io, 2021
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bfe8ab4c-5883-4dae-bdd5-b7c2a033375e · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference https://github.com/intel/ Multi-llms-Chatbot-CloudNative-LangChain , 2022
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7cb0b10c-5c72-42c7-820a-23ae4df85cd2 · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference https://docs.nvidia.com/ deploy/mps/index.html, 2022
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f8e2f591-8374-4098-a8a6-4f439e3d9f51 · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference https://sharegpt.com/, 2023
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0e2ffc21-970c-4021-ad09-c809c8a3fd76 · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference https://github.com/NVIDIA/ TensorRT-LLM, 2023
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 30ee97f0-454d-4498-a683-085927b9d4cf · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference https://github.com/ huggingface/text-generation-inference, 2023
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3f08bd48-9ae9-407b-8dae-d6a497397c42 · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference GPT-4 Technical Report
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86a07eb9-9c87-4048-8481-745792a8a018 · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d05bb618-29ef-4a9e-a683-2da7b9daee7d · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference pfabric: Minimal near-optimal datacenter transport
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 45ceae2e-00af-42c3-8b54-fa97901524c5 · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 74150bb5-c2f8-4f96-b4b7-3924b8695457 · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a423f4c-22be-4540-b5c4-3b861345263b · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Language models are few-shot learners
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 12a6f59a-544e-4085-804c-2856e0574a86 · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Round-robin syn- chronization: Mitigating communication bottlenecks in parameter servers
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 108fd049-8043-4ddb-8568-ed22aa21b78f · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Evaluating Large Language Models Trained on Code
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ae1a0f0-9cc4-4839-9fe7-67b4416419a4 · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Gonzalez, Ion Sto- ica, and Eric P
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ebed2ad7-3604-4734-ab14-18343632a9ba · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Fu, Stefano Ermon, Atri Rudra, and Christopher Ré
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4aca20fb-7ba0-4821-a637-cfb9847e78cf · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Muxserve: Flexible spatial-temporal multiplex- ing for multiple llm serving
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 64b2057c-11b8-4c18-9e0e-77cd9ccde5b3 · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference The Llama 3 Herd of Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 964531ed-3e3d-4bc5-a3d5-2105cfc0a74d · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Turbotransformers: an efficient gpu serving system for transformer models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 561bcb60-52e4-4e30-8242-52a98e2b0249 · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Elasticflow: An elastic server- less training platform for distributed deep learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ca31ddfd-51f7-4a7e-8a9a-c8f5cfa0e6e4 · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Microsecond-scale preemption for concurrent gpu-accelerated dnn inferences
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d2c85ef4-c9e7-4393-b965-2be97d35296d · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac7f61a9-07fa-4315-b3e7-bb5ad79f5fec · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Flashdecod- ing++: Faster large language model inference with asyn- chronization, flat gemm optimization, and heuristics
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e2a3d4bc-076a-42b9-bf77-696f2ed63a25 · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7719e71-a5ac-471e-8d7e-3d7c26213d34 · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Gpipe: Effi- cient training of giant neural networks using pipeline parallelism
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ecc4599a-b499-46bf-b337-5d25da035521 · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Shinjuku: Preemptive scheduling for µsecond-scale tail latency
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 290c8da1-53d9-4015-94fb-2ea69faa371a · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Reducing activation recomputation in large transformer models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 13a05f4b-25dc-4ac6-9aaa-230fb6bc8c73 · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Gonza- lez, Hao Zhang, and Ion Stoica
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 15d34d5a-454c-40b0-a7d4-a61e1350e5fc · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Sequence parallelism: Long sequence training from system perspective
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 993a94d8-7221-492f-b46e-c519cac68129 · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Competition-level code generation with alphacode
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3b3d3043-2aeb-4426-aba0-bfea757543cd · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Alpaserve: Sta- tistical multiplexing with model parallelism for deep learning serving
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7af6f457-8d82-43b7-a81d-c976d297061c · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Terapipe: Token-level pipeline parallelism for training large-scale language models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 24f35d1e-0f34-4aa3-8a44-83907a6554a9 · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Zico: Efficient gpu memory sharing for concurrent dnn training
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 10393a52-3a95-41f3-9c7c-3d58ce127f58 · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Pipedream: Gen- eralized pipeline parallelism for dnn training
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1c598833-6275-46a8-a4b4-14eaddbebf7b · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Splitwise: Efficient generative llm inference using phase splitting
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 680c6e61-5bc9-4070-9d7e-59aeacb5e1e8 · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Efficiently scal- ing transformer inference
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a6758c99-b124-4416-8159-45e92a8712b3 · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Zero: Memory optimizations toward train- ing trillion parameter models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 418eeff4-2bb2-44d8-8635-3d86d76ec71a · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Serverless in the wild: Characterizing and optimizing the serverless workload at a large cloud provider
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 28abbef3-c935-4056-99c5-e7b03e64a3c1 · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 686fb88e-7559-4346-b6b9-d5b9663b7f76 · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Llm-planner: Few-shot grounded planning for embodied agents with large language models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b6110508-98e0-496b-b71e-c1b844c56dbd · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Dynamollm: Designing llm inference clusters for performance and energy efficiency
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6e4fa56-6336-48b4-851b-aa1407346ef2 · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Llumnix: Dynamic scheduling for large language model serving
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 178d2813-11b0-4bd9-9b3e-677066d2c086 · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80e70e61-a666-41bf-8b35-1fd428345ac4 · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Finding Optimal Policy for Queueing Models: New Parameterization
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 303f878a-d1b0-453a-adb6-f66e547d35bb · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Dynamic scheduling with convex delay costs: The generalized c| mu rule
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e519926c-c0b5-4239-a092-d6b098a22f3f · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference LightSeq: A High Performance Inference Library for Transformers
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4124632a-afae-4454-b462-ec55e7c25d5c · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c26e7288-8978-4be1-a3cf-b83a1a225c51 · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38ed9563-9588-4085-b534-e9e980d893e1 · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Fast Distributed Inference Serving for Large Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 635dc22d-7d7d-4672-964a-1ad137728cf6 · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference dLoRA: Dynamically orches- trating requests and adapters for LoRA LLM serving
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8fbf48e9-723a-482d-87ef-10e42c1b0b70 · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Antman: Dynamic scaling on gpu clusters for deep learning
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation cf8516fd-f1f6-4a2b-bc5d-37625e1bc7ee · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Orca: A distributed serving system for Transformer-Based generative mod- els
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0f5f2a6b-ec18-4eee-ba3a-ac7b892d32a3 · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference OPT: Open Pre-trained Transformer Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 316f5ec4-ee82-46ad-8360-c5da5ec6dd62 · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Multi- resource interleaving for deep learning training
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f1052953-1145-48e7-b62c-76bc95815f9f · outbound
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Dist- serve: Disaggregating prefill and decoding for goodput- optimized large language model serving
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 592b20dd-ebed-4ee1-aa86-bea085db47fc · inbound
Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c5985ef2-03a5-4a08-9244-cc03e0a45768 · inbound
ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e5689416-4417-4ade-b42a-f2ea026da362 · inbound
ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.