Pith. sign in

Paper Citation Record · LEDGER

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures

As of 19 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2506.05871.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05871 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:18:24.943745Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T20:35:48.828322Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9a84f048-fc01-47c7-bdd5-71fb2a72c78c · outbound

This paper cites Gqa: Training generalized multi-query transformer models from multi-head checkpoints, 2023.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Gqa: Training generalized multi-query transformer models from multi-head checkpoints, 2023

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:21.975462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:21.975462Z digest=sha256:9952ff08fcb2faba1a34660dd5a04b315a8effc30c858be40a901434aafc93b0

Observation b76c699d-04ab-4a89-9df4-bb88f9d7a139 · outbound

This paper cites How continuous batching enables 23x throughput in LLM inference while reducing p50 latency.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures How continuous batching enables 23x throughput in LLM inference while reducing p50 latency

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:32.117041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:22.047032Z digest=sha256:00a5fbc8c6e7b02d65c66f8c9ee3422368f409afc45aee8c2057026e2888e49f

Observation f116b98f-bafb-43fe-a337-bb60bf1d77e5 · outbound

This paper cites an unresolved cited work.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:18:31.863421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:22.111097Z digest=sha256:b67e6ed5cfb3e7ade702aa458bcd037e4013329c82f685211e4aec3398af1a15

Observation d3569a65-42ef-4074-a853-8884f11aae61 · outbound

This paper cites an unresolved cited work.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:22.166883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:22.166883Z digest=sha256:f608fc8fe5e68625a990b23aaf4e7f591025466192ed4d4ac4462a30271690d3

Observation 4ffec0d6-76ce-4efd-934c-30698fb7ab4c · outbound

This paper cites Throughput is not all you need: Maximizing goodput in llm serving using prefill-decode disaggregation.https://hao-ai- lab.github.io/blogs/distserve/, 2024.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Throughput is not all you need: Maximizing goodput in llm serving using prefill-decode disaggregation.https://hao-ai- lab.github.io/blogs/distserve/, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:31.602003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:22.238130Z digest=sha256:52f5b6689c046a473617b6b1dc1619128774e955ef658a6d10360f26caff51ca

Observation 3a7fe4fb-4914-4440-b118-d8163dfad57a · outbound

This paper cites an unresolved cited work.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:18:31.356513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:22.310379Z digest=sha256:575dd5445d4b9efb23de16df5cd9966abcba243300cbea832426f103225c4747

Observation f8c79ad7-811b-472b-ab2e-bb55a4b7f564 · outbound

This paper cites an unresolved cited work.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:18:31.052695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:22.376684Z digest=sha256:cf04a7fc92bfb7b675f717b2772713389048690bd3cd1fd7b1a63eeba1bc4092

Observation 32076b1c-5f97-41ab-ab8d-614d6ea9393c · outbound

This paper cites an unresolved cited work.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:18:30.817164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:22.458150Z digest=sha256:95d1a06a3acd88d56496b287953ae0dbc6dfc1d86846b80d268e6dea8c505af8

Observation 24b96557-fb56-44a3-95ff-7b54f2e5471a · outbound

This paper cites Sigmoid-weighted linear units for neural network function approximation in reinforcement learning, 2017.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Sigmoid-weighted linear units for neural network function approximation in reinforcement learning, 2017

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:22.540447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:22.540447Z digest=sha256:0c4c53665f5f4c2854fae8fc31835bfdf6cd97b09e844e16cfcdcbe3cf816732

Observation c16b42c3-e6aa-402a-918d-71d2ccaafcb1 · outbound

This paper cites Text generation inference.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Text generation inference

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:30.599741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:22.605821Z digest=sha256:6c6b2730474649d210fcd094630ee3680eb3e86c1335b176b0ec71a293a5fdcc

Observation 6e795329-54e9-4350-9bd8-073f3c64a412 · outbound

This paper cites Low latency rnn inference with cellular batching.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Low latency rnn inference with cellular batching

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:30.273233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:22.666922Z digest=sha256:c90661b58de8b8bee532aebae9338f7a00d2e28245eaac3c49bb4eaf6f6a3332

Observation d676dfcc-085d-457a-8234-0659cbf90478 · outbound

This paper cites Getting started with CUDA graphs.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Getting started with CUDA graphs

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:30.035124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:22.721275Z digest=sha256:164c82f1f37b7c7d0048d17f3caf4f008f5b53fd24039dbbe9865acd109f1c5c

Observation 86bb2b9d-0fd9-496c-b9fb-290bcd0678d7 · outbound

This paper cites Shortle, James M.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Shortle, James M

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:29.734840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:22.782207Z digest=sha256:f0faf9acc1af6ad0d00edd3b923b2ff5814ff7f2bd3265cba577dc3cb9438307

Observation 0aee096f-153b-446d-b9b7-33c76b19bbb1 · outbound

This paper cites Pipedream: Fast and efficient pipeline parallel dnn training, 2018.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Pipedream: Fast and efficient pipeline parallel dnn training, 2018

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:29.478952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:22.842703Z digest=sha256:93f05ee94a7d7dfbe3d9ed934ed9ab06f38f619b50aebe9935113ae2790b978c

Observation f32f8b3b-45e4-4fcc-8ad3-536d46c7d48a · outbound

This paper cites Inference without interference: Disaggregate llm inference for mixed downstream workloads, 2024.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Inference without interference: Disaggregate llm inference for mixed downstream workloads, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:22.918596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:22.918596Z digest=sha256:9b4765cb823975c39fe66f381513d0e3a9f37ba51a103b8618f02b468732759f

Observation fc0b7dfd-3473-4204-aaa8-8d8ad4ed900b · outbound

This paper cites Le, Yonghui Wu, and Zhifeng Chen.GPipe: efficient training of giant neural networks using pipeline parallelism.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Le, Yonghui Wu, and Zhifeng Chen.GPipe: efficient training of giant neural networks using pipeline parallelism

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:29.186315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:23.004203Z digest=sha256:76f9d24426a74d425b9f64e5d9565ec031e18b310364c2bd439ee32e978f1911

Observation 6a9ad0c6-582b-4c8c-bf15-b812ea35e28e · outbound

This paper cites an unresolved cited work.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:23.064579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:23.064579Z digest=sha256:24f9517c4c111de537a82125127bcafa3afbeaa04a8142e62384e7e4fa10abb2

Observation 145f5443-f828-4631-87b0-8ea6a3de16bb · outbound

This paper cites Efficient memory management for large language model serving with PagedAttention.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Efficient memory management for large language model serving with PagedAttention

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:29.105821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:23.122706Z digest=sha256:c90a9bf9ed8f897c7d944a45227faa72c5d8f6935a3ef8aef52bf30f059f7f75

Observation 0d23d8a8-73ad-40d6-91da-9b8aedc2cdec · outbound

This paper cites Transformers KV caching explained.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Transformers KV caching explained

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:29.006858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:23.182303Z digest=sha256:580a3bdf635f082cc8b56f6bcbb4747e769934890017cc738526981130e4b1f2

Observation 45caf8d9-9714-4566-b9fa-0db1dcdb3cea · outbound

This paper cites Sequence parallelism: Long sequence training from system perspective, 2022.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Sequence parallelism: Long sequence training from system perspective, 2022

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:28.761856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:23.303096Z digest=sha256:d44369418b1e7410f3f1a096c00cc184f63b0783d39954e21d55da78dcb8ce66

Observation 64e64901-3b20-4764-b6a6-651291e110b3 · outbound

This paper cites Llama 3.2: Revolutionizing edge AI and vision with open, customizable models.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Llama 3.2: Revolutionizing edge AI and vision with open, customizable models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:28.549453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:23.353733Z digest=sha256:e147f1a73bf25102b4a1fe584a9d67c6cbc5ea36af5c457bf1f34b43d7d81257

Observation e8618f5a-0dd0-46b8-a1b4-6c6368e7b5cd · outbound

This paper cites NVIDIA TensorRT-LLM.https: //docs.nvidia.com/tensorrt-llm/index.html.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures NVIDIA TensorRT-LLM.https: //docs.nvidia.com/tensorrt-llm/index.html

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:28.335252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:23.468448Z digest=sha256:a1e87319948108295b31255869ebebc1a0c30e45c40f07171035300d2471282a

Observation bc6c61ce-9f85-444a-8f49-441f6dd4c07c · outbound

This paper cites OpenAI o3-mini.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures OpenAI o3-mini

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:28.084574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:23.581221Z digest=sha256:a07ee23a872e0d91cd716681a57f5d3017d0ebf0c1f2da26dbf4fbfa685c477e

Observation 294d1086-4bb9-4ba4-8247-257ff3a68a55 · outbound

This paper cites Splitwise: Efficient generative llm inference using phase splitting, 2024.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Splitwise: Efficient generative llm inference using phase splitting, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:27.810127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:23.701320Z digest=sha256:5e529deafc1e9d20c5edb25194af7280bece0550e84a6929f9018c2172b4250b

Observation dd690d7d-cf99-41df-8990-1f63a9f18590 · outbound

This paper cites Mooncake: A KVCache-centric disaggregated architecture for LLM serving, 2024.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Mooncake: A KVCache-centric disaggregated architecture for LLM serving, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:27.600512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:23.849100Z digest=sha256:47aabb00febfee2fe0dd06cafe387774fee62357305d11bdaed871cc222aa85f

Observation fa45246a-90e3-447d-98ee-280bf8e2edb9 · outbound

This paper cites Focus: For tech giants, AI like Bing and Bard poses billion-dollar search problem.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Focus: For tech giants, AI like Bing and Bard poses billion-dollar search problem

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:27.412217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:23.901389Z digest=sha256:8dabb21948c8bbb53dd8f7886afd45ff1302cd6e2d70fa7f92a013e98f355c81

Observation a06f4a7f-9185-4cf2-9fd8-08725cf7b59b · outbound

This paper cites an unresolved cited work.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:18:27.145596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:23.942676Z digest=sha256:bfae29319e286211492b5169cc773ee545d158fa976ec1a4b68e1601ec4f19b8

Observation 23909fa7-09e8-4b89-96af-7ccdb9dbc366 · outbound

This paper cites Megatron-lm: Training multi-billion parameter language models using model parallelism, 2020.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Megatron-lm: Training multi-billion parameter language models using model parallelism, 2020

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:23.999552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:23.999552Z digest=sha256:c999ce964bde264ed6ec0919bbc91f9ba6a15c71db1d8801853b202788a646e8

Observation c7662fe8-a85d-4e52-a49c-f766c9634da4 · outbound

This paper cites Gomez, Łukasz Kaiser, and Illia Polosukhin.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Gomez, Łukasz Kaiser, and Illia Polosukhin

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:26.742702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:24.118959Z digest=sha256:9550232970a0a8e11f28d48d4fdeb04e6c8c6bcef80ccd8581927e6d5fb5a12c

Observation ed9c0fd2-6578-4842-a611-4a965f03cead · outbound

This paper cites https://docs.vllm.ai/en/v0.4.2/index.html.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures https://docs.vllm.ai/en/v0.4.2/index.html

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:26.408769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:24.229841Z digest=sha256:7b0668ad877ef1d7a98bcbaf12e649899306040a4ace5b0c88ad22deac872358

Observation 2f0c0e4d-6004-450f-8805-292f75ab23b8 · outbound

This paper cites https://github.com/vllm-project/vllm-ascend.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures https://github.com/vllm-project/vllm-ascend

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:26.114753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:24.346122Z digest=sha256:9f5ace13adda88ed64dcbb68ba185418182fd20c1da9d524735b212b960f5871

Observation 248364bd-e2c4-427e-ac45-4f940f6b129f · outbound

This paper cites Simai: Unifying architecture design and performance tuning for large-scale large language model training with scalability and precision.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Simai: Unifying architecture design and performance tuning for large-scale large language model training with scalability and precision

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:25.835717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:24.449900Z digest=sha256:649f072376240ccb7632e7131f502dd9de137484acc0e09d4b2ed8a75e850f83

Observation 86364566-a7b8-4d13-b770-c3fc07d52163 · outbound

This paper cites Roofline: an insightful visual performance model for multicore architectures.Commun.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Roofline: an insightful visual performance model for multicore architectures.Commun

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:25.636873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:24.579266Z digest=sha256:37f4a3a6f62678329eeaec63443f809443f973cc9b068363c522999f84be0627

Observation 9ac34d5f-de27-4130-a262-bbfaf2fe842d · outbound

This paper cites Orca: A distributed serving system for Transformer-Based generative models.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Orca: A distributed serving system for Transformer-Based generative models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:25.376580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:24.712230Z digest=sha256:6761ecae01f9323e80fc2a327286435952198e282972f4490ee948ec347b9fc9

Observation 9e1804fc-072c-44a9-b496-b7a853b79dec · outbound

This paper cites Root mean square layer normalization, 2019.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Root mean square layer normalization, 2019

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:24.846042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:24.846042Z digest=sha256:e542efc81370109b0aeb354a4521d7da924d53f3bb0051698faa99e2e0795756

Observation a0d35056-e58c-43c2-81eb-367868a0e7fd · outbound

This paper cites DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:25.146940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:18:24.943745Z digest=sha256:f76306d8874e26df24390d939b29d7b086da64563261e0cda424d3db4de2273c

Pith citing papers

Observation 23e4ae32-5d5a-4207-aa2e-5810f04e42cb · inbound

Simple is Better: Multiplication May Be All You Need for LLM Request Scheduling cites this paper.

Simple is Better: Multiplication May Be All You Need for LLM Request Scheduling BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T20:35:48.828322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T20:35:48.828322Z digest=sha256:69dfd16bf11cb55ae7c158d51ec104b78d3022d741b212f2b8238c2769f37528