Pith. sign in

Paper Citation Record · LEDGER

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference

As of 19 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 3 inbound Pith citation observations for arXiv:2504.15720.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.15720 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:27:17.414952Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-21T08:39:31.911497Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T08:39:53.273252Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact1
  • verified fuzzy40
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7340987e-080f-449a-9a66-6ac0cdbe4ecc · outbound

This paper cites https://github.com/NVIDIA/ FasterTransformer, 2019.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference https://github.com/NVIDIA/ FasterTransformer, 2019

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.437702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.165095Z digest=sha256:374a60f493a60a06421bcfdd8369414c997c6b56a9e5924dcdc52aa649a1438e

Observation 6137fcf8-0bf5-47cf-a180-f172b67436dc · outbound

This paper cites https://grpc.io, 2021.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference https://grpc.io, 2021

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.422918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.170496Z digest=sha256:e83f279abe29d7e4da35ef639c808007e727f1d426925abcaccfb98954ee8167

Observation bfe8ab4c-5883-4dae-bdd5-b7c2a033375e · outbound

This paper cites https://github.com/intel/ Multi-llms-Chatbot-CloudNative-LangChain , 2022.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference https://github.com/intel/ Multi-llms-Chatbot-CloudNative-LangChain , 2022

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.407689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.174815Z digest=sha256:9f8fde7d55fa6b3fb9b578a7534d77d92b2e59fcc39879d78d0f9ffc1a31fb8a

Observation 7cb0b10c-5c72-42c7-820a-23ae4df85cd2 · outbound

This paper cites https://docs.nvidia.com/ deploy/mps/index.html, 2022.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference https://docs.nvidia.com/ deploy/mps/index.html, 2022

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.392475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.179061Z digest=sha256:58a1c942448bf26e81c85bc6b55f071b2203166ff3fd78b4222285dc3b808d7f

Observation f8e2f591-8374-4098-a8a6-4f439e3d9f51 · outbound

This paper cites https://sharegpt.com/, 2023.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference https://sharegpt.com/, 2023

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.377688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.183565Z digest=sha256:4e2f52217718fd31fe85c7045924012905767ce14aca45bfca08f25730f54ac4

Observation 0e2ffc21-970c-4021-ad09-c809c8a3fd76 · outbound

This paper cites https://github.com/NVIDIA/ TensorRT-LLM, 2023.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference https://github.com/NVIDIA/ TensorRT-LLM, 2023

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.362918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.188222Z digest=sha256:79c5842bad6e8c089e17b27c37e55558b7b2a98529a383776b884330454fba31

Observation 30ee97f0-454d-4498-a683-085927b9d4cf · outbound

This paper cites https://github.com/ huggingface/text-generation-inference, 2023.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference https://github.com/ huggingface/text-generation-inference, 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.348008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.193938Z digest=sha256:53da3c959777344a5a03352375924004552ba6fbf78ef24e0095c7f1a7dbcb6b

Observation 3f08bd48-9ae9-407b-8dae-d6a497397c42 · outbound

This paper cites GPT-4 Technical Report.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference GPT-4 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:17.199332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:17.199332Z digest=sha256:b7c9ea115e67fe3d8d965fae359f20432579a7b0a5a96b3ef19d8672e6667dea

Observation 86a07eb9-9c87-4048-8481-745792a8a018 · outbound

This paper cites SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:17.203985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:17.203985Z digest=sha256:fe20b7f3063daa1d28163a3482c90258751d3e8a1bf2526eb41446da279db8c9

Observation d05bb618-29ef-4a9e-a683-2da7b9daee7d · outbound

This paper cites pfabric: Minimal near-optimal datacenter transport.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference pfabric: Minimal near-optimal datacenter transport

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.332789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.208802Z digest=sha256:f0526dbe5c86f116ece8aaef70543b1fd22960e039253797195b4138e96accbc

Observation 45ceae2e-00af-42c3-8b54-fa97901524c5 · outbound

This paper cites Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.317402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.213120Z digest=sha256:6429c9d72558cd27a5f03ccd580400562889fa2c223e7e8f56f055e23c1c878e

Observation 74150bb5-c2f8-4f96-b4b7-3924b8695457 · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:17.217764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:17.217764Z digest=sha256:575b1a150d09a241df9f9d5ca02a92f48eebfad0f10f3529974b1ad6b00c66a4

Observation 8a423f4c-22be-4540-b5c4-3b861345263b · outbound

This paper cites Language models are few-shot learners.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Language models are few-shot learners

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.302016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.222805Z digest=sha256:f9233e6b974005fe0f49cced4d4e349b20fb28ef873f09805387dd4d6132f387

Observation 12a6f59a-544e-4085-804c-2856e0574a86 · outbound

This paper cites Round-robin syn- chronization: Mitigating communication bottlenecks in parameter servers.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Round-robin syn- chronization: Mitigating communication bottlenecks in parameter servers

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.287280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.227789Z digest=sha256:55c68e4d78ec4841cd70c3f2777a5950940c7280cab19247fcee51c15b13836b

Observation 108fd049-8043-4ddb-8568-ed22aa21b78f · outbound

This paper cites Evaluating Large Language Models Trained on Code.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Evaluating Large Language Models Trained on Code

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:17.232360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:17.232360Z digest=sha256:1e05c037bdc3dd30a67c1543b1a8afb5124c99cd75b3e9ec66940a6701539623

Observation 6ae1a0f0-9cc4-4839-9fe7-67b4416419a4 · outbound

This paper cites Gonzalez, Ion Sto- ica, and Eric P.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Gonzalez, Ion Sto- ica, and Eric P

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.273477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.236934Z digest=sha256:3ba22b3e65a5e7e253fff3ba960eddb1671fa4ff1eb916b6fe54a4aba72d1297

Observation ebed2ad7-3604-4734-ab14-18343632a9ba · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher Ré.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.260324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.241661Z digest=sha256:85960574a8f65c5969653316d677cb7396b9ecfab4a1f844e4c626c112cf0b1c

Observation 4aca20fb-7ba0-4821-a637-cfb9847e78cf · outbound

This paper cites Muxserve: Flexible spatial-temporal multiplex- ing for multiple llm serving.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Muxserve: Flexible spatial-temporal multiplex- ing for multiple llm serving

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.245593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.246315Z digest=sha256:283e08a6d856451aeec07c654a37b28cf072dfb9380e337640136cdfde3d636c

Observation 64b2057c-11b8-4c18-9e0e-77cd9ccde5b3 · outbound

This paper cites The Llama 3 Herd of Models.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference The Llama 3 Herd of Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:17.251005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:17.251005Z digest=sha256:489e6fbfadf8ad153125f9bd882a3f91d0c8728eefee4da00f39ba0d030866ba

Observation 964531ed-3e3d-4bc5-a3d5-2105cfc0a74d · outbound

This paper cites Turbotransformers: an efficient gpu serving system for transformer models.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Turbotransformers: an efficient gpu serving system for transformer models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.230830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.255630Z digest=sha256:9a89b70981fab5b1260e932c3fa63637c832ad0cf9bca3f3bafc97402d55a94e

Observation 561bcb60-52e4-4e30-8242-52a98e2b0249 · outbound

This paper cites Elasticflow: An elastic server- less training platform for distributed deep learning.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Elasticflow: An elastic server- less training platform for distributed deep learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.215724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.259691Z digest=sha256:668e13fa42ca0473a1f6656f15c53f4adc0c06cba085d4146374607393b6b2b4

Observation ca31ddfd-51f7-4a7e-8a9a-c8f5cfa0e6e4 · outbound

This paper cites Microsecond-scale preemption for concurrent gpu-accelerated dnn inferences.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Microsecond-scale preemption for concurrent gpu-accelerated dnn inferences

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.199737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.263861Z digest=sha256:c9ec2b21c4da4492c149947266bf8aa06b3d563f2bc9e070bf82d365ee852c84

Observation d2c85ef4-c9e7-4393-b965-2be97d35296d · outbound

This paper cites DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:17.268072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:17.268072Z digest=sha256:0fbda5ca86007f57b46114a1c0bb8dcfbc9fba70d229f147fc85bdf54c6877d8

Observation ac7f61a9-07fa-4315-b3e7-bb5ad79f5fec · outbound

This paper cites Flashdecod- ing++: Faster large language model inference with asyn- chronization, flat gemm optimization, and heuristics.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Flashdecod- ing++: Faster large language model inference with asyn- chronization, flat gemm optimization, and heuristics

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.184577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.272830Z digest=sha256:3431055e3a054e68473bdfd7f8f972733573b8921746d7badc4b9e53188a1615

Observation e2a3d4bc-076a-42b9-bf77-696f2ed63a25 · outbound

This paper cites Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:17.277341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:17.277341Z digest=sha256:c430b2c70ad5efb4aae74c8e574fbe879b523ac334e26e15a37cc0628d897540

Observation c7719e71-a5ac-471e-8d7e-3d7c26213d34 · outbound

This paper cites Gpipe: Effi- cient training of giant neural networks using pipeline parallelism.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Gpipe: Effi- cient training of giant neural networks using pipeline parallelism

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.168562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.281992Z digest=sha256:314b37b0614054e1f99f7f650e610ddae16bf3dc1c14a52a6aaacba0fba04ca7

Observation ecc4599a-b499-46bf-b337-5d25da035521 · outbound

This paper cites Shinjuku: Preemptive scheduling for µsecond-scale tail latency.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Shinjuku: Preemptive scheduling for µsecond-scale tail latency

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.153806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.286445Z digest=sha256:03b988fa1a242e3efe7e844e9ce4f74415b8c0ed69aa0bb4c22b6fefbe113059

Observation 290c8da1-53d9-4015-94fb-2ea69faa371a · outbound

This paper cites Reducing activation recomputation in large transformer models.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Reducing activation recomputation in large transformer models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.138872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.291196Z digest=sha256:115a9cd7203e8fa44ee15a878857b7dd5997a83ddb8ef8a8c80a7d417db32e83

Observation 13a05f4b-25dc-4ac6-9aaa-230fb6bc8c73 · outbound

This paper cites Gonza- lez, Hao Zhang, and Ion Stoica.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Gonza- lez, Hao Zhang, and Ion Stoica

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.123732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.295780Z digest=sha256:79844994b41d459d03cedab62afe9652735aec29e5fb53230e9d82cbf4271980

Observation 15d34d5a-454c-40b0-a7d4-a61e1350e5fc · outbound

This paper cites Sequence parallelism: Long sequence training from system perspective.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Sequence parallelism: Long sequence training from system perspective

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.109417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.300240Z digest=sha256:41a7d9505713a4a599bea06f1c23da819565fc6d3727a2808c0bdb2d0e9cfa01

Observation 993a94d8-7221-492f-b46e-c519cac68129 · outbound

This paper cites Competition-level code generation with alphacode.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Competition-level code generation with alphacode

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.094156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.304850Z digest=sha256:4ee04a893d6b71316a97f5c76cb65f6bcbf891a6aec2091f31979e27b76ebc58

Observation 3b3d3043-2aeb-4426-aba0-bfea757543cd · outbound

This paper cites Alpaserve: Sta- tistical multiplexing with model parallelism for deep learning serving.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Alpaserve: Sta- tistical multiplexing with model parallelism for deep learning serving

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.079789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.309266Z digest=sha256:1bd5529690a7a8f6b4b22499672dd845321ad2e0e64eed9598b3dffbba91b5ca

Observation 7af6f457-8d82-43b7-a81d-c976d297061c · outbound

This paper cites Terapipe: Token-level pipeline parallelism for training large-scale language models.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Terapipe: Token-level pipeline parallelism for training large-scale language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.064680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.313460Z digest=sha256:57ca2a14f19737e28b04a95d162200753838ccaa60e26aa9b99e1c5ec5b81917

Observation 24f35d1e-0f34-4aa3-8a44-83907a6554a9 · outbound

This paper cites Zico: Efficient gpu memory sharing for concurrent dnn training.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Zico: Efficient gpu memory sharing for concurrent dnn training

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.048956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.317817Z digest=sha256:402d9d3dd083810b53cdba9973c624faa08b3e02ad31bb0218f9b5091863b035

Observation 10393a52-3a95-41f3-9c7c-3d58ce127f58 · outbound

This paper cites Pipedream: Gen- eralized pipeline parallelism for dnn training.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Pipedream: Gen- eralized pipeline parallelism for dnn training

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.033732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.322245Z digest=sha256:2e1359dc4958c2bb70fc1ce3e385138e95acf19f0caab2146c814a06b45f0c24

Observation 1c598833-6275-46a8-a4b4-14eaddbebf7b · outbound

This paper cites Splitwise: Efficient generative llm inference using phase splitting.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Splitwise: Efficient generative llm inference using phase splitting

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.018424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.326827Z digest=sha256:34fb6f32de0e4a3891de7de3896a661efcfc09b5f68877309b84bc8ef95a6c99

Observation 680c6e61-5bc9-4070-9d7e-59aeacb5e1e8 · outbound

This paper cites Efficiently scal- ing transformer inference.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Efficiently scal- ing transformer inference

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:18.002764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.331272Z digest=sha256:a0f777438e655eb59a0c44fd59563ed3783508e07ae771677fff14c8392deca1

Observation a6758c99-b124-4416-8159-45e92a8712b3 · outbound

This paper cites Zero: Memory optimizations toward train- ing trillion parameter models.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Zero: Memory optimizations toward train- ing trillion parameter models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:17.987739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.335903Z digest=sha256:be53d461fa1443551af74911b73bea2a02a9456615e358bb90872f4544599302

Observation 418eeff4-2bb2-44d8-8635-3d86d76ec71a · outbound

This paper cites Serverless in the wild: Characterizing and optimizing the serverless workload at a large cloud provider.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Serverless in the wild: Characterizing and optimizing the serverless workload at a large cloud provider

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:17.972262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.340171Z digest=sha256:d91b2fa79fafd7caba6b111be2de3b51d8e39784fb57efa28f608b06134f8077

Observation 28abbef3-c935-4056-99c5-e7b03e64a3c1 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:17.344593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:17.344593Z digest=sha256:a337ba8931e22636fed02f11ffacfc73fd47b0520688cf20fcf56e01b45bb84a

Observation 686fb88e-7559-4346-b6b9-d5b9663b7f76 · outbound

This paper cites Llm-planner: Few-shot grounded planning for embodied agents with large language models.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Llm-planner: Few-shot grounded planning for embodied agents with large language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:17.956997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.349303Z digest=sha256:4eb81023dc9ba758add4f009c03c775201794cd8460ee1299313b138a0eab7ed

Observation b6110508-98e0-496b-b71e-c1b844c56dbd · outbound

This paper cites Dynamollm: Designing llm inference clusters for performance and energy efficiency.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Dynamollm: Designing llm inference clusters for performance and energy efficiency

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:17.353326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:17.353326Z digest=sha256:87c2ff8547467a245e997ffda2acec6293cd95368b81c272d4e9df2e3cf9bb44

Observation a6e4fa56-6336-48b4-851b-aa1407346ef2 · outbound

This paper cites Llumnix: Dynamic scheduling for large language model serving.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Llumnix: Dynamic scheduling for large language model serving

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:17.941685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.357268Z digest=sha256:c080ecd88bb29913249d8163a2ab0bbac29e0ab8c412362e91b7b066638b0e67

Observation 178d2813-11b0-4bd9-9b3e-677066d2c086 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:17.361613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:17.361613Z digest=sha256:3d5a51d1539a530eb2f718ae7765c7a601bec9794406ee28117b6c5e5863b170

Observation 80e70e61-a666-41bf-8b35-1fd428345ac4 · outbound

This paper cites Finding Optimal Policy for Queueing Models: New Parameterization.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Finding Optimal Policy for Queueing Models: New Parameterization

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:27:17.534781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.365401Z digest=sha256:0ad3f5c7f9eaa1522dae0d7263bffc2cee401491f818fbf8f5fe2673c3d90102

Observation 303f878a-d1b0-453a-adb6-f66e547d35bb · outbound

This paper cites Dynamic scheduling with convex delay costs: The generalized c| mu rule.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Dynamic scheduling with convex delay costs: The generalized c| mu rule

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:17.924984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.369609Z digest=sha256:74faf0e8ec6d172e8ba745fe449cb88664bb580282f1dc5122a8c1c4ff59acc3

Observation e519926c-c0b5-4239-a092-d6b098a22f3f · outbound

This paper cites LightSeq: A High Performance Inference Library for Transformers.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference LightSeq: A High Performance Inference Library for Transformers

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:17.373953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:17.373953Z digest=sha256:0c47d81841e6faa9dd79b7ca206529fdf57c4830320cd5771347f8144c60b4f5

Observation 4124632a-afae-4454-b462-ec55e7c25d5c · outbound

This paper cites BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:17.378835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:17.378835Z digest=sha256:09c2a5c76bf59a377c049a01293a78f033243640b53fde66887c4808fed39b4e

Observation c26e7288-8978-4be1-a3cf-b83a1a225c51 · outbound

This paper cites LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:17.383632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:17.383632Z digest=sha256:5d89f89ca26c42d4ec4718e97b8d89c93935b01c8912f68301b3a865ede5b40a

Observation 38ed9563-9588-4085-b534-e9e980d893e1 · outbound

This paper cites Fast Distributed Inference Serving for Large Language Models.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Fast Distributed Inference Serving for Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:17.388272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:17.388272Z digest=sha256:c6fda4292ff8e38db8162a342d44eacbd640489a4b0266e4ce259a0096946194

Observation 635dc22d-7d7d-4672-964a-1ad137728cf6 · outbound

This paper cites dLoRA: Dynamically orches- trating requests and adapters for LoRA LLM serving.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference dLoRA: Dynamically orches- trating requests and adapters for LoRA LLM serving

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:17.907122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.392963Z digest=sha256:b39bac6d8f64d46c85f169f942d205d181e4ef691f8160d78e81b3f63c6f4385

Observation 8fbf48e9-723a-482d-87ef-10e42c1b0b70 · outbound

This paper cites Antman: Dynamic scaling on gpu clusters for deep learning.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Antman: Dynamic scaling on gpu clusters for deep learning

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:17.892653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.397544Z digest=sha256:8587803f288562ac74df8132f34e25e2b8bf100db0d8a2722e0a1e6b3719fe84

Observation cf8516fd-f1f6-4a2b-bc5d-37625e1bc7ee · outbound

This paper cites Orca: A distributed serving system for Transformer-Based generative mod- els.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Orca: A distributed serving system for Transformer-Based generative mod- els

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:17.878113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.401771Z digest=sha256:626d3ce27c0d3f8bc23812a662dbec54f3c7c3774d0185d6e0f02302cab9b039

Observation 0f5f2a6b-ec18-4eee-ba3a-ac7b892d32a3 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference OPT: Open Pre-trained Transformer Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:17.406179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:17.406179Z digest=sha256:5d54306984cb03c07daec7c148f6e36774dbc1eee5cec58c8ae6e692df3d555a

Observation 316f5ec4-ee82-46ad-8360-c5da5ec6dd62 · outbound

This paper cites Multi- resource interleaving for deep learning training.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Multi- resource interleaving for deep learning training

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:17.863787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.410400Z digest=sha256:3628b10c5c0336663403ba903d143e5086c1084a46ddf432e9fb321e6e418c6b

Observation f1052953-1145-48e7-b62c-76bc95815f9f · outbound

This paper cites Dist- serve: Disaggregating prefill and decoding for goodput- optimized large language model serving.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference Dist- serve: Disaggregating prefill and decoding for goodput- optimized large language model serving

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:17.847784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:27:17.414952Z digest=sha256:536fe9d34a355463790ea1faa8f426fa7d86ebdbb9ae43c22a7958a7a56b7394

Pith citing papers

Observation 592b20dd-ebed-4ee1-aa86-bea085db47fc · inbound

Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines cites this paper.

Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:59:03.501623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T09:52:40.057014Z digest=sha256:ecc9d812f1ab8dcd5f721d45be22de5bb3982c3a3981d551f9634501cb6689f9

Observation c5985ef2-03a5-4a08-9244-cc03e0a45768 · inbound

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL cites this paper.

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:31:16.469748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T05:14:14.168753Z digest=sha256:0522c525d8d295f2b38ca2db70e33d435abb387f45328f828e30361c353873a0

Observation e5689416-4417-4ade-b42a-f2ea026da362 · inbound

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL cites this paper.

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:39:53.275321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T08:39:31.911497Z digest=sha256:2fee22f2378365eaad3629f31aa21af5d3ea235945467809825807c722359ee5