Pith. sign in

Paper Citation Record · LEDGER

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference

As of 19 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 3 inbound Pith citation observations for arXiv:2504.15720.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.15720 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:27:17.414952Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-21T08:39:31.911497Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T08:39:53.273252Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact1
  • verified fuzzy40
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7340987e-080f-449a-9a66-6ac0cdbe4ecc · outbound

This paper cites https://github.com/NVIDIA/ FasterTransformer, 2019.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference https://github.com/NVIDIA/ FasterTransformer, 2019

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.437702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.165095Z digest=sha256:f8058097f3507722829122fff65e433f7d420f9d1f182671756a7f865e5c16b0

Observation 6137fcf8-0bf5-47cf-a180-f172b67436dc · outbound

This paper cites https://grpc.io, 2021.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference https://grpc.io, 2021

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.422918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.170496Z digest=sha256:0cb048b35adb5aaa882f616d1f306cb4790c736fcc97b7fbd1baad676759421f

Observation bfe8ab4c-5883-4dae-bdd5-b7c2a033375e · outbound

This paper cites https://github.com/intel/ Multi-llms-Chatbot-CloudNative-LangChain , 2022.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference https://github.com/intel/ Multi-llms-Chatbot-CloudNative-LangChain , 2022

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.407689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.174815Z digest=sha256:0ff6e4ff60e9b1f3cb921e400b04abe3c4553958bea2a9b441bc3ee3f5210a1f

Observation 7cb0b10c-5c72-42c7-820a-23ae4df85cd2 · outbound

This paper cites https://docs.nvidia.com/ deploy/mps/index.html, 2022.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference https://docs.nvidia.com/ deploy/mps/index.html, 2022

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.392475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.179061Z digest=sha256:688be522c16106bdefa2cba92f49e9045308f0c107a23e01802c2eec0781bda9

Observation f8e2f591-8374-4098-a8a6-4f439e3d9f51 · outbound

This paper cites https://sharegpt.com/, 2023.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference https://sharegpt.com/, 2023

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.377688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.183565Z digest=sha256:f4df0891971d3f6574e721c8327acda3d0986a9ea201653ca03236050eb054a0

Observation 0e2ffc21-970c-4021-ad09-c809c8a3fd76 · outbound

This paper cites https://github.com/NVIDIA/ TensorRT-LLM, 2023.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference https://github.com/NVIDIA/ TensorRT-LLM, 2023

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.362918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.188222Z digest=sha256:c9b2626843b956604dd34d73ea2a71a75b9411208870f751c54c81f66c9f9c04

Observation 30ee97f0-454d-4498-a683-085927b9d4cf · outbound

This paper cites https://github.com/ huggingface/text-generation-inference, 2023.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference https://github.com/ huggingface/text-generation-inference, 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.348008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.193938Z digest=sha256:3b5cd0465926e6bc53d31303b63c2850d1acfe9ef8ff76cc8466a594ff36c7ff

Observation 3f08bd48-9ae9-407b-8dae-d6a497397c42 · outbound

This paper cites GPT-4 Technical Report.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference GPT-4 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:17.199332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:17.199332Z digest=sha256:2794c8331cbfc4e427ef83b9ef2f89741007cee4fbd5afb3d8185f195e163311

Observation 86a07eb9-9c87-4048-8481-745792a8a018 · outbound

This paper cites SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:17.203985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:17.203985Z digest=sha256:901d7d69a01d6c60c0af0a3e1dba69452cd137863c9cd6bf5a1315f84ad61646

Observation d05bb618-29ef-4a9e-a683-2da7b9daee7d · outbound

This paper cites pfabric: Minimal near-optimal datacenter transport.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference pfabric: Minimal near-optimal datacenter transport

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.332789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.208802Z digest=sha256:4af93b8e302f98581cfd2d51a8a54412ec5b06e9a55c2cce0daf502f15f1c782

Observation 45ceae2e-00af-42c3-8b54-fa97901524c5 · outbound

This paper cites Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.317402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.213120Z digest=sha256:86bb0876465ca73115f119dc92ebf558c458fd88c8a17742a70eae773d38e943

Observation 74150bb5-c2f8-4f96-b4b7-3924b8695457 · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:17.217764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:17.217764Z digest=sha256:f7fea4ad18e33547a8b3244abeae6689c450cf753e462639323cbd3b8b329219

Observation 8a423f4c-22be-4540-b5c4-3b861345263b · outbound

This paper cites Language models are few-shot learners.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Language models are few-shot learners

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.302016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.222805Z digest=sha256:aaee7007a1dfce73ce3ddf9d8001d013b846efd68f9fea201f1655ca2b2bc50c

Observation 12a6f59a-544e-4085-804c-2856e0574a86 · outbound

This paper cites Round-robin syn- chronization: Mitigating communication bottlenecks in parameter servers.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Round-robin syn- chronization: Mitigating communication bottlenecks in parameter servers

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.287280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.227789Z digest=sha256:aa32578b1d55c840290620296d9174eb1ca1a6dd46052295c3d176ad711397f2

Observation 108fd049-8043-4ddb-8568-ed22aa21b78f · outbound

This paper cites Evaluating Large Language Models Trained on Code.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Evaluating Large Language Models Trained on Code

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:17.232360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:17.232360Z digest=sha256:24b68409b1833b1f65ddf4c5ced610bf2133bae8e32af54d7b7d934a976e60b8

Observation 6ae1a0f0-9cc4-4839-9fe7-67b4416419a4 · outbound

This paper cites Gonzalez, Ion Sto- ica, and Eric P.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Gonzalez, Ion Sto- ica, and Eric P

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.273477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.236934Z digest=sha256:07c8ef17b3ccec7bbc7e9c618490cd88917ad6845a7bf184bdd4b6dfccf529ab

Observation ebed2ad7-3604-4734-ab14-18343632a9ba · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher Ré.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.260324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.241661Z digest=sha256:e583c26959820d453e15b1fd2857718f9ac38a69b8030c41fab1217a2ee5b4c4

Observation 4aca20fb-7ba0-4821-a637-cfb9847e78cf · outbound

This paper cites Muxserve: Flexible spatial-temporal multiplex- ing for multiple llm serving.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Muxserve: Flexible spatial-temporal multiplex- ing for multiple llm serving

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.245593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.246315Z digest=sha256:00b6c474329c44b12f00243b0a4e5b0ec163cfe2fd6e7b9d3c189f493dec9c44

Observation 64b2057c-11b8-4c18-9e0e-77cd9ccde5b3 · outbound

This paper cites The Llama 3 Herd of Models.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference The Llama 3 Herd of Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:17.251005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:17.251005Z digest=sha256:1e8a90389fe80868a3d696c4ed18220ae9d765546a5cb2b708e11a923249d210

Observation 964531ed-3e3d-4bc5-a3d5-2105cfc0a74d · outbound

This paper cites Turbotransformers: an efficient gpu serving system for transformer models.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Turbotransformers: an efficient gpu serving system for transformer models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.230830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.255630Z digest=sha256:6f62fa8911dbedcc27645ffd70a004a9958409389cc3149740a86533091b9118

Observation 561bcb60-52e4-4e30-8242-52a98e2b0249 · outbound

This paper cites Elasticflow: An elastic server- less training platform for distributed deep learning.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Elasticflow: An elastic server- less training platform for distributed deep learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.215724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.259691Z digest=sha256:30ad64d8aeed0b76df769318f482e5ff5dc9a5b19c8b7a7d5f39087a387d161e

Observation ca31ddfd-51f7-4a7e-8a9a-c8f5cfa0e6e4 · outbound

This paper cites Microsecond-scale preemption for concurrent gpu-accelerated dnn inferences.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Microsecond-scale preemption for concurrent gpu-accelerated dnn inferences

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.199737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.263861Z digest=sha256:5c4f4da753eb26b945a9570b52e47b2d1d1d27c6d004016fe07ba25e314bc01e

Observation d2c85ef4-c9e7-4393-b965-2be97d35296d · outbound

This paper cites DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:17.268072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:17.268072Z digest=sha256:b01d3aaf2699db97721fe83bafa8a87ddf9211e706315345c4434f6fdc398f8a

Observation ac7f61a9-07fa-4315-b3e7-bb5ad79f5fec · outbound

This paper cites Flashdecod- ing++: Faster large language model inference with asyn- chronization, flat gemm optimization, and heuristics.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Flashdecod- ing++: Faster large language model inference with asyn- chronization, flat gemm optimization, and heuristics

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.184577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.272830Z digest=sha256:23c7d6ee8ecd3b517b7bc922c0d9d6efa95b4b829ac18c60caeeafea69b5b8a0

Observation e2a3d4bc-076a-42b9-bf77-696f2ed63a25 · outbound

This paper cites Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:17.277341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:17.277341Z digest=sha256:75db73f7453cddf9686dfa49126927e32215f7ce8a9223c6c274c38f1c7e9c98

Observation c7719e71-a5ac-471e-8d7e-3d7c26213d34 · outbound

This paper cites Gpipe: Effi- cient training of giant neural networks using pipeline parallelism.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Gpipe: Effi- cient training of giant neural networks using pipeline parallelism

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.168562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.281992Z digest=sha256:63b303db0604070d651a47780a935e71f6099b68b561e433140b7e705d05f738

Observation ecc4599a-b499-46bf-b337-5d25da035521 · outbound

This paper cites Shinjuku: Preemptive scheduling for µsecond-scale tail latency.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Shinjuku: Preemptive scheduling for µsecond-scale tail latency

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.153806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.286445Z digest=sha256:87029f21c91b9e94f6b52a01db52bc53facaf8ed96956b386d964f7027e39bd2

Observation 290c8da1-53d9-4015-94fb-2ea69faa371a · outbound

This paper cites Reducing activation recomputation in large transformer models.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Reducing activation recomputation in large transformer models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.138872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.291196Z digest=sha256:32c1626ebe43d7703850c20f91cda31103f1fc48a61e1e618892ebe1a88fce7f

Observation 13a05f4b-25dc-4ac6-9aaa-230fb6bc8c73 · outbound

This paper cites Gonza- lez, Hao Zhang, and Ion Stoica.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Gonza- lez, Hao Zhang, and Ion Stoica

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.123732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.295780Z digest=sha256:741a4677c59100472c9ed7cf12bbba4c2fcf94fbc58372bcb3c82a8e63a50299

Observation 15d34d5a-454c-40b0-a7d4-a61e1350e5fc · outbound

This paper cites Sequence parallelism: Long sequence training from system perspective.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Sequence parallelism: Long sequence training from system perspective

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.109417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.300240Z digest=sha256:7e2d3dc4d4467f3a5d2c7c8606c576a2560521116c2d74dcda379585b6db3361

Observation 993a94d8-7221-492f-b46e-c519cac68129 · outbound

This paper cites Competition-level code generation with alphacode.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Competition-level code generation with alphacode

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.094156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.304850Z digest=sha256:ed846e70d47d3bfe947f27d6e5be9b5f0bbfdbf4155fd0b8afd6ccbc5f164801

Observation 3b3d3043-2aeb-4426-aba0-bfea757543cd · outbound

This paper cites Alpaserve: Sta- tistical multiplexing with model parallelism for deep learning serving.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Alpaserve: Sta- tistical multiplexing with model parallelism for deep learning serving

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.079789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.309266Z digest=sha256:bfad3c946c47de0ddb3e969ef18fbd26100bae7b138f31b3bbc699a9cd35a15d

Observation 7af6f457-8d82-43b7-a81d-c976d297061c · outbound

This paper cites Terapipe: Token-level pipeline parallelism for training large-scale language models.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Terapipe: Token-level pipeline parallelism for training large-scale language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.064680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.313460Z digest=sha256:ea8654f67ac86442e84f46ba2146b6ff2bf0af8f030f9ccd3d6b9de676906a4c

Observation 24f35d1e-0f34-4aa3-8a44-83907a6554a9 · outbound

This paper cites Zico: Efficient gpu memory sharing for concurrent dnn training.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Zico: Efficient gpu memory sharing for concurrent dnn training

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.048956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.317817Z digest=sha256:f6274ac9868410f6c2b36aafd086ef10d054c7d9f08421e1fc5c3403384c9dc8

Observation 10393a52-3a95-41f3-9c7c-3d58ce127f58 · outbound

This paper cites Pipedream: Gen- eralized pipeline parallelism for dnn training.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Pipedream: Gen- eralized pipeline parallelism for dnn training

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.033732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.322245Z digest=sha256:742803ad8804199b28dcdd4b20636ac453dbd94a1ddddcfe9c1d3b8303330a1e

Observation 1c598833-6275-46a8-a4b4-14eaddbebf7b · outbound

This paper cites Splitwise: Efficient generative llm inference using phase splitting.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Splitwise: Efficient generative llm inference using phase splitting

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.018424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.326827Z digest=sha256:4692ca3e7e63cf52affae3540027859906452747e3e5b38380ab4eac41dff143

Observation 680c6e61-5bc9-4070-9d7e-59aeacb5e1e8 · outbound

This paper cites Efficiently scal- ing transformer inference.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Efficiently scal- ing transformer inference

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.002764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.331272Z digest=sha256:d2c88f463488621cf7ad26bf78effb08c80770141965f12a6e287ed70bb4f7df

Observation a6758c99-b124-4416-8159-45e92a8712b3 · outbound

This paper cites Zero: Memory optimizations toward train- ing trillion parameter models.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Zero: Memory optimizations toward train- ing trillion parameter models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:17.987739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.335903Z digest=sha256:bf6ade3e70579e25455f76a3b18580b2cdd65b9f2a1a5c19aeaf9308601bae82

Observation 418eeff4-2bb2-44d8-8635-3d86d76ec71a · outbound

This paper cites Serverless in the wild: Characterizing and optimizing the serverless workload at a large cloud provider.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Serverless in the wild: Characterizing and optimizing the serverless workload at a large cloud provider

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:17.972262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.340171Z digest=sha256:a402c1347d7f72763c7d6e56f7ae14d4bd8607c907855f2bfc976e764e613f16

Observation 28abbef3-c935-4056-99c5-e7b03e64a3c1 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:17.344593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:17.344593Z digest=sha256:43ce2fb5502b249eee60757ae8a766101f0f0eb89f2fa28fab7a69255e1f1872

Observation 686fb88e-7559-4346-b6b9-d5b9663b7f76 · outbound

This paper cites Llm-planner: Few-shot grounded planning for embodied agents with large language models.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Llm-planner: Few-shot grounded planning for embodied agents with large language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:17.956997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.349303Z digest=sha256:1480c3432111efee54c26242eebc3ecf6d6b5a30294bc9a296a30dfc477b780f

Observation b6110508-98e0-496b-b71e-c1b844c56dbd · outbound

This paper cites Dynamollm: Designing llm inference clusters for performance and energy efficiency.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Dynamollm: Designing llm inference clusters for performance and energy efficiency

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:17.353326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:17.353326Z digest=sha256:f41e1000efb9b94200bfad507a295bfe6534a96dbfb37e9a100c40c1eb11bf4b

Observation a6e4fa56-6336-48b4-851b-aa1407346ef2 · outbound

This paper cites Llumnix: Dynamic scheduling for large language model serving.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Llumnix: Dynamic scheduling for large language model serving

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:17.941685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.357268Z digest=sha256:daa8e9861abb22f84cbd47fe222cfc9d03921880f94c59944e72bdce811ea98b

Observation 178d2813-11b0-4bd9-9b3e-677066d2c086 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:17.361613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:17.361613Z digest=sha256:569fc9cdcaca64c94ece9b60993fd1c35ada1d8254ef3f9b320dc18e2b15f4dc

Observation 80e70e61-a666-41bf-8b35-1fd428345ac4 · outbound

This paper cites Finding Optimal Policy for Queueing Models: New Parameterization.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Finding Optimal Policy for Queueing Models: New Parameterization

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:27:17.534781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.365401Z digest=sha256:4f87d3744fee1fbf2c36058bb0ba6c19179791afaf5d598be7cab6a6e0d10649

Observation 303f878a-d1b0-453a-adb6-f66e547d35bb · outbound

This paper cites Dynamic scheduling with convex delay costs: The generalized c| mu rule.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Dynamic scheduling with convex delay costs: The generalized c| mu rule

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:17.924984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.369609Z digest=sha256:838fe4a49a4f3d8913908037eceeef6d34364e3b0fe878cfec5f58dc1e44644e

Observation e519926c-c0b5-4239-a092-d6b098a22f3f · outbound

This paper cites LightSeq: A High Performance Inference Library for Transformers.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference LightSeq: A High Performance Inference Library for Transformers

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:17.373953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:17.373953Z digest=sha256:06edadbc53c9e04dabe875d0af84966059863beffea8ca036cb8e1a33bbd1a99

Observation 4124632a-afae-4454-b462-ec55e7c25d5c · outbound

This paper cites BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:17.378835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:17.378835Z digest=sha256:3f6ebcfe28fff820347ad8ce5f5db723f409e14adf1d0f1611c6a5a6cebb0ad2

Observation c26e7288-8978-4be1-a3cf-b83a1a225c51 · outbound

This paper cites LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:17.383632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:17.383632Z digest=sha256:64e1715ff84c38d3f5bcf71c1bfe516fb4e73e61c88d08b7cc1ea90ef8f5965b

Observation 38ed9563-9588-4085-b534-e9e980d893e1 · outbound

This paper cites Fast Distributed Inference Serving for Large Language Models.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Fast Distributed Inference Serving for Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:17.388272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:17.388272Z digest=sha256:1fe867aaef4b399b26a8f676703825c1e2c6c7a90923f916a419dd3a61f96309

Observation 635dc22d-7d7d-4672-964a-1ad137728cf6 · outbound

This paper cites dLoRA: Dynamically orches- trating requests and adapters for LoRA LLM serving.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference dLoRA: Dynamically orches- trating requests and adapters for LoRA LLM serving

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:17.907122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.392963Z digest=sha256:53429c223506ed16880dc53e801d2363e11a36c9ffa582a4c2a7542c37d75780

Observation 8fbf48e9-723a-482d-87ef-10e42c1b0b70 · outbound

This paper cites Antman: Dynamic scaling on gpu clusters for deep learning.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Antman: Dynamic scaling on gpu clusters for deep learning

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:17.892653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.397544Z digest=sha256:e0507c0e48415b9e53c7bd46a0299020557c89c3c916867008bb1d3a8ca50009

Observation cf8516fd-f1f6-4a2b-bc5d-37625e1bc7ee · outbound

This paper cites Orca: A distributed serving system for Transformer-Based generative mod- els.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Orca: A distributed serving system for Transformer-Based generative mod- els

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:17.878113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.401771Z digest=sha256:0a5b0439a13062e95a399df6b48998c77b3d4b4a1bf58ddc5231820ccd577ed1

Observation 0f5f2a6b-ec18-4eee-ba3a-ac7b892d32a3 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference OPT: Open Pre-trained Transformer Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:17.406179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:17.406179Z digest=sha256:4b090cecd820654ee07b986273920fdfe455b01f169a59c8f78cac2eacb28b41

Observation 316f5ec4-ee82-46ad-8360-c5da5ec6dd62 · outbound

This paper cites Multi- resource interleaving for deep learning training.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Multi- resource interleaving for deep learning training

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:17.863787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.410400Z digest=sha256:e3874b5bf9ba19533f0bc75450e5c03b84941df9a7c890083cd8021e5ca07f23

Observation f1052953-1145-48e7-b62c-76bc95815f9f · outbound

This paper cites Dist- serve: Disaggregating prefill and decoding for goodput- optimized large language model serving.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Dist- serve: Disaggregating prefill and decoding for goodput- optimized large language model serving

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:17.847784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.414952Z digest=sha256:39fb124472c00e0f12611e38fb5b8707ee4ab360518b82941475f645341a955e

Pith citing papers

Observation 592b20dd-ebed-4ee1-aa86-bea085db47fc · inbound

Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines cites this paper.

Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:59:03.501623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T09:52:40.057014Z digest=sha256:038fdcc7206419aa6c973655c0bc886632f2cda76def5e725c95cd65ce03e978

Observation c5985ef2-03a5-4a08-9244-cc03e0a45768 · inbound

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL cites this paper.

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:31:16.469748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T05:14:14.168753Z digest=sha256:0ff4e0e340aa1424052aea78d699909e456697cd098c0cc5e52e562c946b880e

Observation e5689416-4417-4ade-b42a-f2ea026da362 · inbound

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL cites this paper.

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:39:53.275321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T08:39:31.911497Z digest=sha256:a1cb5776e6a36c35cf00ac0af0a4e5b4f948215ee84785729b2e1609fb98c80b