Pith. sign in

Paper Citation Record · LEDGER

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures

As of 18 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2506.05871.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05871 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:18:24.943745Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T20:35:48.828322Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9a84f048-fc01-47c7-bdd5-71fb2a72c78c · outbound

This paper cites Gqa: Training generalized multi-query transformer models from multi-head checkpoints, 2023.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Gqa: Training generalized multi-query transformer models from multi-head checkpoints, 2023

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:21.975462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:21.975462Z digest=sha256:9952ff08fcb2faba1a34660dd5a04b315a8effc30c858be40a901434aafc93b0

Observation b76c699d-04ab-4a89-9df4-bb88f9d7a139 · outbound

This paper cites How continuous batching enables 23x throughput in LLM inference while reducing p50 latency.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures How continuous batching enables 23x throughput in LLM inference while reducing p50 latency

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:32.117041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:22.047032Z digest=sha256:9f34c8beda4253810da60bc5e81f6eb0f4e2eac9f4825949bfa7570fcb917f75

Observation f116b98f-bafb-43fe-a337-bb60bf1d77e5 · outbound

This paper cites an unresolved cited work.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:18:31.863421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:22.111097Z digest=sha256:5e409a0afb842b8b94c8d19202042d7ace5d2d2ee55974ff6466f4c59e25d016

Observation d3569a65-42ef-4074-a853-8884f11aae61 · outbound

This paper cites an unresolved cited work.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:22.166883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:22.166883Z digest=sha256:f608fc8fe5e68625a990b23aaf4e7f591025466192ed4d4ac4462a30271690d3

Observation 4ffec0d6-76ce-4efd-934c-30698fb7ab4c · outbound

This paper cites Throughput is not all you need: Maximizing goodput in llm serving using prefill-decode disaggregation.https://hao-ai- lab.github.io/blogs/distserve/, 2024.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Throughput is not all you need: Maximizing goodput in llm serving using prefill-decode disaggregation.https://hao-ai- lab.github.io/blogs/distserve/, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:31.602003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:22.238130Z digest=sha256:93a09cee5bb680b29db82192ec40b996ec8fb3c1cc24f15a3bb7ae8e942f8be6

Observation 3a7fe4fb-4914-4440-b118-d8163dfad57a · outbound

This paper cites an unresolved cited work.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:18:31.356513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:22.310379Z digest=sha256:ada60a8bb2215e6ddd97e66d6745f8f4d9972a1fdd4ba3da3491516626fe6e3b

Observation f8c79ad7-811b-472b-ab2e-bb55a4b7f564 · outbound

This paper cites an unresolved cited work.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:18:31.052695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:22.376684Z digest=sha256:7198664042a5bccbab4a091ff2d00137146114a036b82dc1e63119a56b89d8db

Observation 32076b1c-5f97-41ab-ab8d-614d6ea9393c · outbound

This paper cites an unresolved cited work.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:18:30.817164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:22.458150Z digest=sha256:28819149ea576ac3c8b18675c6fbcebd2bf544a4af9bf05fcb11c593e4b3e0dc

Observation 24b96557-fb56-44a3-95ff-7b54f2e5471a · outbound

This paper cites Sigmoid-weighted linear units for neural network function approximation in reinforcement learning, 2017.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Sigmoid-weighted linear units for neural network function approximation in reinforcement learning, 2017

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:22.540447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:22.540447Z digest=sha256:0c4c53665f5f4c2854fae8fc31835bfdf6cd97b09e844e16cfcdcbe3cf816732

Observation c16b42c3-e6aa-402a-918d-71d2ccaafcb1 · outbound

This paper cites Text generation inference.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Text generation inference

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:30.599741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:22.605821Z digest=sha256:96c45da8ca009022ff656e9b2e6e0e3231cdd0f186ce6cf695a99e3eeeafd3cd

Observation 6e795329-54e9-4350-9bd8-073f3c64a412 · outbound

This paper cites Low latency rnn inference with cellular batching.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Low latency rnn inference with cellular batching

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:30.273233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:22.666922Z digest=sha256:6cebe6642f6b84717c21db786315e9ab1db9dce23b3bbfbcef7054025394c0c2

Observation d676dfcc-085d-457a-8234-0659cbf90478 · outbound

This paper cites Getting started with CUDA graphs.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Getting started with CUDA graphs

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:30.035124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:22.721275Z digest=sha256:3a30345e5261c10482d2347507f6fdb83ea2d8eb0ddb9d94c6917a047e6fe7d4

Observation 86bb2b9d-0fd9-496c-b9fb-290bcd0678d7 · outbound

This paper cites Shortle, James M.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Shortle, James M

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:29.734840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:22.782207Z digest=sha256:e2492e86bf6fa072c7d64b8071b06f25ec3ae68f7a4d33904c7c0b0fb056298e

Observation 0aee096f-153b-446d-b9b7-33c76b19bbb1 · outbound

This paper cites Pipedream: Fast and efficient pipeline parallel dnn training, 2018.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Pipedream: Fast and efficient pipeline parallel dnn training, 2018

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:29.478952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:22.842703Z digest=sha256:aba1e2959add2168ca1a6cedb8fa635017eab6618b1f9ddbd451b4588debf1f6

Observation f32f8b3b-45e4-4fcc-8ad3-536d46c7d48a · outbound

This paper cites Inference without interference: Disaggregate llm inference for mixed downstream workloads, 2024.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Inference without interference: Disaggregate llm inference for mixed downstream workloads, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:22.918596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:22.918596Z digest=sha256:9b4765cb823975c39fe66f381513d0e3a9f37ba51a103b8618f02b468732759f

Observation fc0b7dfd-3473-4204-aaa8-8d8ad4ed900b · outbound

This paper cites Le, Yonghui Wu, and Zhifeng Chen.GPipe: efficient training of giant neural networks using pipeline parallelism.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Le, Yonghui Wu, and Zhifeng Chen.GPipe: efficient training of giant neural networks using pipeline parallelism

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:29.186315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:23.004203Z digest=sha256:853a83e91040a532d551d65a4112c7097d9578b941e42d7ab643a5c0a16dc359

Observation 6a9ad0c6-582b-4c8c-bf15-b812ea35e28e · outbound

This paper cites an unresolved cited work.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:23.064579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:23.064579Z digest=sha256:24f9517c4c111de537a82125127bcafa3afbeaa04a8142e62384e7e4fa10abb2

Observation 145f5443-f828-4631-87b0-8ea6a3de16bb · outbound

This paper cites Efficient memory management for large language model serving with PagedAttention.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Efficient memory management for large language model serving with PagedAttention

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:29.105821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:23.122706Z digest=sha256:7eab905cc385b9ced92386b22a1c0c11e42b92ddc2d42ff98e6a69e576d09878

Observation 0d23d8a8-73ad-40d6-91da-9b8aedc2cdec · outbound

This paper cites Transformers KV caching explained.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Transformers KV caching explained

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:29.006858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:23.182303Z digest=sha256:b49a797f8e6a04649337e2efd7967765bc1e61b06da2592bdb867d8bffbc92d0

Observation 45caf8d9-9714-4566-b9fa-0db1dcdb3cea · outbound

This paper cites Sequence parallelism: Long sequence training from system perspective, 2022.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Sequence parallelism: Long sequence training from system perspective, 2022

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:28.761856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:23.303096Z digest=sha256:4bcc50cc52a296d87c1ab6ca57a1ca54ab00ea228951af37b90792ce36755200

Observation 64e64901-3b20-4764-b6a6-651291e110b3 · outbound

This paper cites Llama 3.2: Revolutionizing edge AI and vision with open, customizable models.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Llama 3.2: Revolutionizing edge AI and vision with open, customizable models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:28.549453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:23.353733Z digest=sha256:f1ebd3a21e0411627d4db5e0a19ac26dbcbda8b47c579dec829e8e7796169359

Observation e8618f5a-0dd0-46b8-a1b4-6c6368e7b5cd · outbound

This paper cites NVIDIA TensorRT-LLM.https: //docs.nvidia.com/tensorrt-llm/index.html.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures NVIDIA TensorRT-LLM.https: //docs.nvidia.com/tensorrt-llm/index.html

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:28.335252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:23.468448Z digest=sha256:31c40a9ae1febc722cc3b7654a39f91bc70c7a66110b6f982f858ff53c696907

Observation bc6c61ce-9f85-444a-8f49-441f6dd4c07c · outbound

This paper cites OpenAI o3-mini.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures OpenAI o3-mini

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:28.084574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:23.581221Z digest=sha256:ef571cd22b0d8dd6f9e91e56417abd21ac8308b04661c8e12db3e0851fb01ae8

Observation 294d1086-4bb9-4ba4-8247-257ff3a68a55 · outbound

This paper cites Splitwise: Efficient generative llm inference using phase splitting, 2024.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Splitwise: Efficient generative llm inference using phase splitting, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:27.810127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:23.701320Z digest=sha256:c1accfeec83cd81b57a393f899f7cdd6d13deb7f665cede27cc0cbb359f0a6d6

Observation dd690d7d-cf99-41df-8990-1f63a9f18590 · outbound

This paper cites Mooncake: A KVCache-centric disaggregated architecture for LLM serving, 2024.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Mooncake: A KVCache-centric disaggregated architecture for LLM serving, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:27.600512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:23.849100Z digest=sha256:cd595b0fa0b16b2d48ca3947c9259adf6eb8094a3ef3f442dab4805bd5455c7d

Observation fa45246a-90e3-447d-98ee-280bf8e2edb9 · outbound

This paper cites Focus: For tech giants, AI like Bing and Bard poses billion-dollar search problem.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Focus: For tech giants, AI like Bing and Bard poses billion-dollar search problem

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:27.412217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:23.901389Z digest=sha256:bc1cccadbf81ff3eba5e0424f902ff1752a3a7cea530d5ff3e3a93bca62f601d

Observation a06f4a7f-9185-4cf2-9fd8-08725cf7b59b · outbound

This paper cites an unresolved cited work.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:18:27.145596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:23.942676Z digest=sha256:e5d58a9b7319ef4fcfd75e241d700e6a016aaf941bc8c82133f0fd79ecbd6726

Observation 23909fa7-09e8-4b89-96af-7ccdb9dbc366 · outbound

This paper cites Megatron-lm: Training multi-billion parameter language models using model parallelism, 2020.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Megatron-lm: Training multi-billion parameter language models using model parallelism, 2020

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:23.999552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:23.999552Z digest=sha256:c999ce964bde264ed6ec0919bbc91f9ba6a15c71db1d8801853b202788a646e8

Observation c7662fe8-a85d-4e52-a49c-f766c9634da4 · outbound

This paper cites Gomez, Łukasz Kaiser, and Illia Polosukhin.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Gomez, Łukasz Kaiser, and Illia Polosukhin

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:26.742702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:24.118959Z digest=sha256:d5a25244e6b0502b73d3df092d211dd9b16236228a91ec1d42e3510a63d13410

Observation ed9c0fd2-6578-4842-a611-4a965f03cead · outbound

This paper cites https://docs.vllm.ai/en/v0.4.2/index.html.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures https://docs.vllm.ai/en/v0.4.2/index.html

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:26.408769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:24.229841Z digest=sha256:14d8a96defd816c7dd5de7c794b83aa52a910b9431173ac93278229ce4fec554

Observation 2f0c0e4d-6004-450f-8805-292f75ab23b8 · outbound

This paper cites https://github.com/vllm-project/vllm-ascend.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures https://github.com/vllm-project/vllm-ascend

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:26.114753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:24.346122Z digest=sha256:ec051fe1b9693aad90f12ae756e17693167df6789e67a2840f1ffef002ff9e58

Observation 248364bd-e2c4-427e-ac45-4f940f6b129f · outbound

This paper cites Simai: Unifying architecture design and performance tuning for large-scale large language model training with scalability and precision.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Simai: Unifying architecture design and performance tuning for large-scale large language model training with scalability and precision

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:25.835717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:24.449900Z digest=sha256:119d9691049f5a6ebfa5afaa4bbb70adb011586adc83a85d98a7e1368637ab5b

Observation 86364566-a7b8-4d13-b770-c3fc07d52163 · outbound

This paper cites Roofline: an insightful visual performance model for multicore architectures.Commun.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Roofline: an insightful visual performance model for multicore architectures.Commun

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:25.636873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:24.579266Z digest=sha256:fa0ede0eb646382e78b57a7f75e9db8edca00bf88752dbc614baf1c8ece47c53

Observation 9ac34d5f-de27-4130-a262-bbfaf2fe842d · outbound

This paper cites Orca: A distributed serving system for Transformer-Based generative models.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Orca: A distributed serving system for Transformer-Based generative models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:25.376580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:24.712230Z digest=sha256:1fbd13fc014de37eb29dda6176105071c447733a7078f8e125f930089bb11f9e

Observation 9e1804fc-072c-44a9-b496-b7a853b79dec · outbound

This paper cites Root mean square layer normalization, 2019.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Root mean square layer normalization, 2019

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:24.846042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:24.846042Z digest=sha256:e542efc81370109b0aeb354a4521d7da924d53f3bb0051698faa99e2e0795756

Observation a0d35056-e58c-43c2-81eb-367868a0e7fd · outbound

This paper cites DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:25.146940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:18:24.943745Z digest=sha256:59694243362e8175165141946ba2239bd3196b9e9bf7964ba30cb864ef4d881e

Pith citing papers

Observation 23e4ae32-5d5a-4207-aa2e-5810f04e42cb · inbound

Simple is Better: Multiplication May Be All You Need for LLM Request Scheduling cites this paper.

Simple is Better: Multiplication May Be All You Need for LLM Request Scheduling BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T20:35:48.828322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T20:35:48.828322Z digest=sha256:0a73b063a4f5b9e83312fbf383a04fd68cdc1be51c699a52b5637ada1deace6a