Pith. sign in

Paper Citation Record · LEDGER

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling

As of 10 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2507.18006.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.18006 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:45:20.821661Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

64 of 64 outbound references displayed

  • verified exact0
  • verified fuzzy43
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c2e070f6-fd83-452e-81fd-310b7c0dd7c0 · outbound

This paper cites GPT-4 Technical Report.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.562559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.562559Z digest=sha256:5c9c466c89e61f44392eea5042473042fafb870f53a179b27b087b9344c008cd

Observation f5dd5571-5204-4e44-b7d8-0eccd1211acf · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.567366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.567366Z digest=sha256:281b44659420c913dd17f42be5027b970c6f71809ad3521555c48dfcff423147

Observation 603b8f94-dccd-40a7-8181-7e40f6adac5b · outbound

This paper cites DeepSeek-V3 Technical Report.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling DeepSeek-V3 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.571852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.571852Z digest=sha256:e9f264e19072fa7163ec704529a30a9e40abb71eef1bed61e185b1af9d1ba464

Observation 738dbaeb-a711-43db-bfa1-62f073dc36c5 · outbound

This paper cites Black-Box Tuning for Language-Model-as-a-Service.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Black-Box Tuning for Language-Model-as-a-Service

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.576394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.576394Z digest=sha256:1b53b39d3dee1048493eb4ed63e1a28bd1a9898ab5199efeb3853e9a6f9e40d7

Observation db8eb68e-5004-45fc-b8fa-c030f11af1f7 · outbound

This paper cites Xing, Hao Zhang, Joseph E.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Xing, Hao Zhang, Joseph E

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:28.556477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.581123Z digest=sha256:666dd7a9ea5028b4c2fe0f8c4bfdf81817f36d4f1e67b7b0e47dd80079ae4e46

Observation 55b91946-2e24-42ce-8a21-b0f42ff41d2e · outbound

This paper cites A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.585569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.585569Z digest=sha256:7a90cbba60a13ae375ca7cc6f666c78cc78d00d4f5f5dfe909ee9e348e771241

Observation 71922fdf-ee1e-4418-99a7-d8fa77b1303b · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Evaluating Large Language Models Trained on Code

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.591384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.591384Z digest=sha256:468c80eaa4b6795840bd957fa4abfe37daf16f65cf215d1e6b3454dd422ca57c

Observation bdf5a873-497e-407e-862f-77e370601bbc · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Code Llama: Open Foundation Models for Code

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.595686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.595686Z digest=sha256:012948758b6ca98f970c76d21a2cbd56a9c40815c00b78681b73485b8b59a642

Observation 51fe50b3-27fa-4952-8078-3a7759ba61d3 · outbound

This paper cites Accessed: Apr.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Accessed: Apr

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:28.337931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.600056Z digest=sha256:53e3b2e752096612999dd931a7fdca5940283bdcdf7319de237bdd7b629b073c

Observation b7a24f66-cd92-485f-a4af-09779ae31ac7 · outbound

This paper cites Accessed: Apr.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Accessed: Apr

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:28.135926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.603833Z digest=sha256:3251d787a125c79eec13245caf63838806b9e1b4b5fc9ec94adb6fbb92125e75

Observation d5d67e8e-aa31-40b0-85f3-1718fcff4046 · outbound

This paper cites Smoothquant: Accurate and efficient post-training quantization for large language models, 2024.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Smoothquant: Accurate and efficient post-training quantization for large language models, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.607908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.607908Z digest=sha256:b5c09b6ab2f85a7f98268747bb42f2713391f5dd3b07616fb36f7576faa91a6f

Observation 8ed255ca-b436-478e-be84-cba993f9765f · outbound

This paper cites Plug-and-play: An efficient post-training pruningmethodforlargelanguagemodels.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Plug-and-play: An efficient post-training pruningmethodforlargelanguagemodels

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:27.983499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.612417Z digest=sha256:f9c1a11208bd7ca1a5975fb852971e04ac8b2acd886273d7c741f86af21fff40

Observation 226107fc-4b94-4798-8697-ebada942d07b · outbound

This paper cites Awq: Activation-aware weight quantization for llm compression and acceleration, 2024.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Awq: Activation-aware weight quantization for llm compression and acceleration, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:27.805979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.616704Z digest=sha256:72d8c540341dfe75f59dad9aa2b269005ef46f3fbe31a22afbaccad9a56a29e7

Observation bbac1862-4e2a-4f20-9b8f-9b4d6315ab69 · outbound

This paper cites Zico Kolter.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Zico Kolter

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:27.665893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.620795Z digest=sha256:55f581c3bdcc440ea36c049de2ee457bfa96a00f782b8b518e44aa41a22d4540

Observation 7f6173b9-8bbc-4803-a0ec-d755b9c17f1b · outbound

This paper cites Multiplexing dynamic deep learning workloads with slo-awareness in gpu clusters.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Multiplexing dynamic deep learning workloads with slo-awareness in gpu clusters

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:27.515390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.624755Z digest=sha256:e929731a0b0c6e783df6e2c7ef89ad00574e159c5229c319cf2323ee23d44770

Observation 43514fe2-a55b-48e3-862c-9f96811117c8 · outbound

This paper cites Cloudnativesim: A toolkit for modeling and simulation of cloud-native applications.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Cloudnativesim: A toolkit for modeling and simulation of cloud-native applications

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:27.346301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.628509Z digest=sha256:653fca6e37d266be59eac3ed212a94b85479d542c01467c9462039edfe32d85b

Observation ab4ad36b-d4ea-4784-8c2b-96f02774f583 · outbound

This paper cites Llminaflash:Efficientlargelanguagemodelinferencewith limited memory, 2024.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Llminaflash:Efficientlargelanguagemodelinferencewith limited memory, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:27.162975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.632443Z digest=sha256:002a15b510503dbed62298cdce868be28165fc0a9e12bc15536fda6a042b910d

Observation 59bf1481-178e-4fc5-b5d8-ae5b96ba2e8a · outbound

This paper cites Spotserve: Serving generative large language models on preemptible instances.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Spotserve: Serving generative large language models on preemptible instances

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.636501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.636501Z digest=sha256:2084f8dd182e774796f2e2f563132c1ba438d63f950468f40106d091043d85cb

Observation 0736df2e-d30d-49e9-8f47-458134c404ed · outbound

This paper cites Llumnix:Dynamicschedulingforlargelanguage model serving.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Llumnix:Dynamicschedulingforlargelanguage model serving

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:26.979547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.640277Z digest=sha256:6d5fff29db2b773d36d6eb44fcaf4177d2f56059d89fda0ef33d112ce6eea145

Observation 559ffcaf-15c7-4c90-8d29-7712367d1853 · outbound

This paper cites Serving heterogeneous machine learning models on multi-gpu servers with spatio-temporal sharing.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Serving heterogeneous machine learning models on multi-gpu servers with spatio-temporal sharing

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:26.787021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.643981Z digest=sha256:cc7dd04a64e3ab92815024e1bc867975cd350c163fdb11bb005b3866797e7005

Observation 9c86c373-b4bc-422f-b8be-c2e46e1f2c0b · outbound

This paper cites Inferline:latency-aware provisioningandscalingforpredictionservingpipelines.In Proceedings of the 11th ACM Symposium on Cloud Computing, pages 477–491, 2020.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Inferline:latency-aware provisioningandscalingforpredictionservingpipelines.In Proceedings of the 11th ACM Symposium on Cloud Computing, pages 477–491, 2020

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:26.604327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.647609Z digest=sha256:a55fc74a97023529f9d4157d1655cb89a71a6133f27f842fa7d48f47ef91e4f6

Observation c8937796-4580-4d32-a58c-fee8349597d3 · outbound

This paper cites Optimizing llm inference throughput via memory-aware and sla-constrained dynamic batching, 2025.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Optimizing llm inference throughput via memory-aware and sla-constrained dynamic batching, 2025

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:26.431926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.651829Z digest=sha256:23726ffe577c73f863dde620eb6a8e4e947298a42423b541c85e98b6ab44cbad

Observation ae8538a9-5dcf-4c02-88d2-141892950ff8 · outbound

This paper cites Alloystack: A library operating system for serverless workflow applications.Pro- ceedings of the Twentieth European Conference on Computer Systems, 2025.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Alloystack: A library operating system for serverless workflow applications.Pro- ceedings of the Twentieth European Conference on Computer Systems, 2025

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:26.270323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.655748Z digest=sha256:e3336890c24b467571b4fe60c8adbdd10ae98fb31906ac9cca6c04d42cf6fce4

Observation 0777050f-f079-44e5-8789-ffd80508e5ea · outbound

This paper cites Lora-flow: Dynamic lora fusion for large lan- guage models in generative tasks.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Lora-flow: Dynamic lora fusion for large lan- guage models in generative tasks

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:26.052390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.660020Z digest=sha256:50dc49e7b440ba2472397f3e61224d8ebf8a8b0bc5f99d7fcec3f607d985f2a4

Observation f1344fe9-90df-4ba6-8d48-ab40f74d133f · outbound

This paper cites Accessed: Apr.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Accessed: Apr

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:25.830080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.663827Z digest=sha256:ba1c317975d447588be5125b464a46591e961a3124e491da29ef123a1efb4ea2

Observation 01a69b05-5d9b-463a-b295-4217399f77da · outbound

This paper cites an unresolved cited work.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:45:25.642539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.667723Z digest=sha256:620b5aa857dbea764aaf1ff56faf4262364b22948f1c4ab4b575f73008bb0147

Observation 2213fac0-1631-4eec-bdc7-3978d02e6221 · outbound

This paper cites an unresolved cited work.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:45:25.486001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.675885Z digest=sha256:a66abe86d87aaf48150ab735373fd882c0c4960c4b64ef976d2178a4502c4226

Observation 1421a709-5b0d-4168-99c6-cff475463e4b · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.681562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.681562Z digest=sha256:b2a2ff2a44a48e4e2bb0566cb2d0fd39d71d633f1621f7fb8665c7760b984721

Observation 25fa635e-8597-4a2c-a108-282095a2f577 · outbound

This paper cites Llm inference serving: Survey of recent advances and opportunities, 2024.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Llm inference serving: Survey of recent advances and opportunities, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:25.265670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.686206Z digest=sha256:0024b11ffe673910e9f7d6c9bc5141c61137677b1b347db9790df60400847cee

Observation 85001b70-bf8c-41bd-9932-9291d297961a · outbound

This paper cites Optimizing mixture-of-experts inference time combining model deployment and communication scheduling, 2024.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Optimizing mixture-of-experts inference time combining model deployment and communication scheduling, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:25.101053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.690334Z digest=sha256:a39fd46d5000d9f5b0817fb82257d561df95bf95012c72008b449b8f5bd23ca7

Observation 3df35fa0-3e6a-4d50-95ce-866af36fc334 · outbound

This paper cites Mooncake: A kvcache-centric disag- gregated architecture for llm serving, 2024.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Mooncake: A kvcache-centric disag- gregated architecture for llm serving, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:24.909654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.694330Z digest=sha256:0f2879c88a67cc0a4622df7ecffb97615ad0fd71aed2dc46f338a1d4455df0a1

Observation ec14dcb2-cd43-4691-8eb5-08930a6dfee8 · outbound

This paper cites Spinfer: Leveraging low-level sparsity for efficient large language model inference on gpus.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Spinfer: Leveraging low-level sparsity for efficient large language model inference on gpus

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:24.699198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.698281Z digest=sha256:f7dd0be2c98ece528abae6ed3bb224f15fb075a1385a94543cb396be486e26fb

Observation 1ecc3806-31e9-491d-ae80-e0550faa2b0b · outbound

This paper cites H2o:Heavy-hitteroracleforefficientgenerativeinference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling H2o:Heavy-hitteroracleforefficientgenerativeinference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:24.511967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.702463Z digest=sha256:b8a97dfa4c07119eec9f0361209a9d9421068bac5e651001c3fb8f6b25fbae89

Observation 76faaf30-a734-4153-8442-4a54223cfa77 · outbound

This paper cites Splitwise: Efficient gen- erative llm inference using phase splitting.2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA), pages 118–132, 2023.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Splitwise: Efficient gen- erative llm inference using phase splitting.2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA), pages 118–132, 2023

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:24.286745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.706607Z digest=sha256:58947d0b41f42b58acf9c152d0b3ded7777615b5f53c32d2b6ddf7b6212d2a4e

Observation c4c725c7-5a59-4d32-8bf4-3e95d7168ab7 · outbound

This paper cites an unresolved cited work.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:45:24.086241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.710774Z digest=sha256:01186563e3a5d8239e17fa9fa1110567faa91b6f1e3d6c63e2cd0e7781460c01

Observation 43ce76af-0cca-4fae-bae4-1514640d0afa · outbound

This paper cites Inference without interference: Disaggregate llm inference for mixed downstream workloads, 2024.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Inference without interference: Disaggregate llm inference for mixed downstream workloads, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:23.843614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.714736Z digest=sha256:7c5f8251a8c1d0a0699123b3d5ec47e0df798b8e1b7bc02e30c2a955cbdafe83

Observation 904791a1-6169-4a39-a9f4-18cb5f46ec30 · outbound

This paper cites Dynamollm:Designingllminferenceclustersforperformance and energy efficiency, 2024.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Dynamollm:Designingllminferenceclustersforperformance and energy efficiency, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:23.669168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.719014Z digest=sha256:8a6387ef2005e32bd11654102cb20eda461abe63607d648ca4acc5f5e022e65d

Observation d9ebb7f2-f091-404d-8285-19b0415de7b8 · outbound

This paper cites Skyserve:Servingaimodels across regions and clouds with spot instances.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Skyserve:Servingaimodels across regions and clouds with spot instances

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:23.437689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.722857Z digest=sha256:8f047130edfba51a7cb216531e0b984474c6ae63fb743a85860617980a8b9bd0

Observation d650dde7-70e7-4800-9354-556790f414ee · outbound

This paper cites Towardsefficientandreliablellmserving: A real-world workload study.arXiv preprint arXiv:2402.XXXXX, 2024.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Towardsefficientandreliablellmserving: A real-world workload study.arXiv preprint arXiv:2402.XXXXX, 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:23.227733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.727165Z digest=sha256:02a3c9826d20ef7eba3a096a4afedc6bb4cbb5a78ec00bd05a88f6ae35b5d919

Observation 50e2209c-e0bc-4164-972b-9d186e868a2e · outbound

This paper cites Usher: Holistic interference avoidance for resource optimized ML inference.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Usher: Holistic interference avoidance for resource optimized ML inference

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:23.012903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.731400Z digest=sha256:04d9783cc51990ed88cf479719cffbfc022638b9852c4d32b6dc75443b56a4a7

Observation 71fe8728-f510-4dde-8c3a-93b4cea88868 · outbound

This paper cites an unresolved cited work.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:45:22.779244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.735534Z digest=sha256:81d217243a7ae1df18c3bba6b06f6d461123da560c658c6c77f437c0b079ffd9

Observation 515dbeab-5865-4a1e-89d1-4b2dbc2a8595 · outbound

This paper cites Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:22.599134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.739514Z digest=sha256:38c773a524a95bbfb04bb7cb3e821caa93aa29fce0600b8a8b8428b3bab47610

Observation b17279ee-733a-4137-87dc-7c2105163f89 · outbound

This paper cites Alpaserve: Statistical multiplexing with model parallelism for deep learning serving.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Alpaserve: Statistical multiplexing with model parallelism for deep learning serving

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:22.387183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.743488Z digest=sha256:bd25e3a571e169801d638305178c54b348db8f89d3097428bab1dcbde34f5bee

Observation 3db551ec-4c4a-449c-8a13-15ea022f6fdf · outbound

This paper cites an unresolved cited work.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:45:22.205088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.747191Z digest=sha256:23be11cf6a33c53c6c7fddb5d24cddb56fcc520f2273e218fa18fdf639e62762

Observation 8e952f21-82e8-4dfb-aeab-26bb3c1c3e81 · outbound

This paper cites Yadwadkar, and Christos Kozyrakis.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Yadwadkar, and Christos Kozyrakis

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.999977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.750754Z digest=sha256:0638241e7c4d49dc39e58e1738cd614ed76826d844f230ca96416b73506903d6

Observation 29feca7a-4ba7-4d1c-870a-f553e4fdb36f · outbound

This paper cites Reducing Activation Recomputation in Large Transformer Models.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Reducing Activation Recomputation in Large Transformer Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.758677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.758677Z digest=sha256:9d7883cf89acde22b9e574435286052abb4a0b9739d4297556ca04f8cfd51f10

Observation 6192e0ab-e576-407f-938f-4a21570f9f37 · outbound

This paper cites On parallel processing systems: Amdahl’s law generalized and some results on optimal design.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling On parallel processing systems: Amdahl’s law generalized and some results on optimal design

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.610721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.762745Z digest=sha256:04354b371feae69b60872188594128aa65672ff6e891cc3d62e3b82f933ad917

Observation 3607326f-b68b-41b0-bfc0-29a9cb5085ec · outbound

This paper cites xformers: A modular and hackable transformer modelling library.https://github.com/facebookresearch/ xformers, 2022.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling xformers: A modular and hackable transformer modelling library.https://github.com/facebookresearch/ xformers, 2022

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.412262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.766581Z digest=sha256:19f9946fac3058498d2e19a1ef915cdb765a8a86d0a5931233077cabd7c32c41

Observation b995e9ff-346e-43c5-a422-8101a97dc13e · outbound

This paper cites https://developer.nvidia.com/management-library-nvml, 2025.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling https://developer.nvidia.com/management-library-nvml, 2025

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.313583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.770284Z digest=sha256:37b70d4a7a55d8ef650c88022e0c3fd3087ff5b01b2944117a5bfb9a2f3fe9de

Observation 33dca7d2-896f-4afa-a787-f6a83496e92f · outbound

This paper cites Orca: A distributed serving system for transformer- based generative models.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Orca: A distributed serving system for transformer- based generative models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.283738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.774358Z digest=sha256:65527b4e597190c9a6a4b485d2d01a19397f9d15e89efa204286c25905bd6280

Observation 080e1029-b4cb-43b7-9733-66fb8c5041d6 · outbound

This paper cites Uellm: A unified and efficient approach for large language model inference serving.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Uellm: A unified and efficient approach for large language model inference serving

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.245867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.778185Z digest=sha256:1e64b095b3bf7de317ba8b4b288f003eb79e91e76102dbb141af2af24fa289f7

Observation 85ae5a38-9a43-400b-9e93-4dcf3748b8cf · outbound

This paper cites Stanford alpaca: An instruction- following llama model.https://github.com/tatsu-lab/stanford_alpaca,.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Stanford alpaca: An instruction- following llama model.https://github.com/tatsu-lab/stanford_alpaca,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.202598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.781963Z digest=sha256:d92f9aacd5e2e1bedb818c77a18449741bdcc3898ed09d57148f73dadae8d060

Observation f8ce4997-8690-4e79-92bb-6d436cbc5305 · outbound

This paper cites Mepipe: Democratizing llm training with memory- efficient slice-level pipeline scheduling on cost-effective accelerators.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Mepipe: Democratizing llm training with memory- efficient slice-level pipeline scheduling on cost-effective accelerators

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.113047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.789659Z digest=sha256:3752f660979042649619ea4ad07cdfce038f59e3f7b81c5398b197a4cdcf8461

Observation d44c4132-fcb8-4e9b-b5eb-fe849f84db51 · outbound

This paper cites Le, and Z.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Le, and Z

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.078885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.793446Z digest=sha256:7cfbfe6793f37d47340ae5cf8d4d3b94f166027cd6ee7aa7d076cacaac338dcc

Observation ab48421c-b07c-4781-905c-e160553b1ac6 · outbound

This paper cites Mist:Efficientdistributedtrainingoflargelanguagemodelsviamemory- parallelism co-optimization.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Mist:Efficientdistributedtrainingoflargelanguagemodelsviamemory- parallelism co-optimization

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.055530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.797367Z digest=sha256:bc9805c2e5b6133b8c4467f7c005d5a1f9386c1df5994194a89c44253c3b7445

Observation b2b93895-5df9-4f87-91b4-4b5a78aab060 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.801131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.801131Z digest=sha256:57f3ad62affa74e26a786fcb3f42558fa6625cb05f646bb0c4663d22def501b9

Observation 848a515f-29f3-418e-8124-111d46d06335 · outbound

This paper cites Alpa: Automating inter and intra- operator parallelism for distributed deep learning.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Alpa: Automating inter and intra- operator parallelism for distributed deep learning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.041953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.805316Z digest=sha256:8c5c0a1d8ed4f835d1840e852cdfcd4a6b95e48c480cdb2cd4bc2dcd6051f86a

Observation a7567b5a-d5ea-4aff-aee2-d1c5c27566fc · outbound

This paper cites Fast state restoration in llm serving with hcache.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Fast state restoration in llm serving with hcache

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.028193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.809631Z digest=sha256:806dfe0815b6e20c294778cf37323b18dfd9edb51d452e5e0c899b7271d3c10b

Observation 5b2ca156-3b79-4afc-8b43-a3107e2d25aa · outbound

This paper cites Fast and live model auto scaling with o(1) host caching, 2024.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Fast and live model auto scaling with o(1) host caching, 2024

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.013948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.813723Z digest=sha256:f4d330ae1dd9d619e360eb067be75b2fbee23d093454ec89fa2a0ff73ef91d4c

Observation 106593b2-91f9-4854-87af-93f5faf2cc46 · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.817827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.817827Z digest=sha256:621cef3df2c5adf2a4ce1e44478794d4b86ad67c55a3abfd21e438b80e517ec2

Observation d07250cf-5f05-4c41-8fdb-e4ecfcde3f15 · outbound

This paper cites Infinigen: Efficientgenerativeinferenceoflargelanguagemodelswithdynamickv cache management.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Infinigen: Efficientgenerativeinferenceoflargelanguagemodelswithdynamickv cache management

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:20.991136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.821661Z digest=sha256:0ab339b69ac8e68aaa9f5c3128ce5244098e32f5458552d03a407b246fc70cc6

Observation bb71e73a-2187-431a-8428-e8c3931b8909 · outbound

This paper cites an unresolved cited work.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Unresolved cited work

Reference 411

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:45:21.824444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.754810Z digest=sha256:9ab816abcb909fb704955a6982e18e8e2042f9d589e4f5346370d570286700f3

Observation a18aea7a-6abb-454a-a786-988f05df980b · outbound

This paper cites an unresolved cited work.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Unresolved cited work

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.671778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.671778Z digest=sha256:b6a439e49c254cc89d7859374f99f9c6911ac9339c60388f89c9eb5ac33f4868

Observation 2ef497c9-742a-44aa-9990-dbfd9f1c122c · outbound

This paper cites an unresolved cited work.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:45:21.165178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:45:20.785956Z digest=sha256:3afa44e1d75d6574a9c2c832fd79f7f3aebb389a130a7a5490f97767a02af138

Pith citing papers

No inbound Pith citation observations are available.