Pith. sign in

Paper Citation Record · LEDGER

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling

As of 8 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2507.18006.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.18006 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:45:20.821661Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

64 of 64 outbound references displayed

  • verified exact0
  • verified fuzzy43
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c2e070f6-fd83-452e-81fd-310b7c0dd7c0 · outbound

This paper cites GPT-4 Technical Report.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.562559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.562559Z digest=sha256:7aeef47e7061858b34bdc199eed145d8e071c7c7d445d3454ee2f26e5e85fbd2

Observation f5dd5571-5204-4e44-b7d8-0eccd1211acf · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.567366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.567366Z digest=sha256:d54fa9b2425b6f2716d37e45065ea923d163b2b055728d6012edb269bc3aa3ae

Observation 603b8f94-dccd-40a7-8181-7e40f6adac5b · outbound

This paper cites DeepSeek-V3 Technical Report.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling DeepSeek-V3 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.571852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.571852Z digest=sha256:487b5855468e7710febdb1451c78ce6bc7530051583283ec8c7fb908b4d6abad

Observation 738dbaeb-a711-43db-bfa1-62f073dc36c5 · outbound

This paper cites Black-Box Tuning for Language-Model-as-a-Service.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Black-Box Tuning for Language-Model-as-a-Service

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.576394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.576394Z digest=sha256:449c9829f1f6ed58c006b8a697bfcaf565b6d245f39a5a7e4eb313165990609d

Observation db8eb68e-5004-45fc-b8fa-c030f11af1f7 · outbound

This paper cites Xing, Hao Zhang, Joseph E.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Xing, Hao Zhang, Joseph E

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:28.556477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.581123Z digest=sha256:d72943eb342835b5263528b2c2042bff57dd189fdb1ba1906f5ef5b9350d75a9

Observation 55b91946-2e24-42ce-8a21-b0f42ff41d2e · outbound

This paper cites A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.585569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.585569Z digest=sha256:11ce254d140e499b5d082d3fac8c894352e85eb474e875de3fb24deb504dbcae

Observation 71922fdf-ee1e-4418-99a7-d8fa77b1303b · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Evaluating Large Language Models Trained on Code

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.591384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.591384Z digest=sha256:4ddbf14464ac453ba7d4783e21d9ceabe4cf4573863e579a4ec45ef2bb942a70

Observation bdf5a873-497e-407e-862f-77e370601bbc · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Code Llama: Open Foundation Models for Code

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.595686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.595686Z digest=sha256:000853dc6c36bc25be99a97f29a49a30a41d6810059b891e90a20e4714cf7560

Observation 51fe50b3-27fa-4952-8078-3a7759ba61d3 · outbound

This paper cites Accessed: Apr.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Accessed: Apr

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:28.337931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.600056Z digest=sha256:0566da5c6219e701797401df436adab8364ead8cab72043d88b878e2a43bd1f4

Observation b7a24f66-cd92-485f-a4af-09779ae31ac7 · outbound

This paper cites Accessed: Apr.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Accessed: Apr

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:28.135926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.603833Z digest=sha256:e749fb3a6aacae20188b30117a998e174866aa437349671738c5d631d5662961

Observation d5d67e8e-aa31-40b0-85f3-1718fcff4046 · outbound

This paper cites Smoothquant: Accurate and efficient post-training quantization for large language models, 2024.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Smoothquant: Accurate and efficient post-training quantization for large language models, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.607908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.607908Z digest=sha256:73a82f237ff000d5053c8276d7df8683a79a4f625bf6f839366ca4a0519009a3

Observation 8ed255ca-b436-478e-be84-cba993f9765f · outbound

This paper cites Plug-and-play: An efficient post-training pruningmethodforlargelanguagemodels.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Plug-and-play: An efficient post-training pruningmethodforlargelanguagemodels

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:27.983499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.612417Z digest=sha256:5332fb99adaf1a47967963887aaeddab1cdcd26f22532d1e17597711f21e31b6

Observation 226107fc-4b94-4798-8697-ebada942d07b · outbound

This paper cites Awq: Activation-aware weight quantization for llm compression and acceleration, 2024.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Awq: Activation-aware weight quantization for llm compression and acceleration, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:27.805979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.616704Z digest=sha256:13ed33cf14bb5a9e08f4e0726ebccf545aa21770de46e06fd70812d857b194ee

Observation bbac1862-4e2a-4f20-9b8f-9b4d6315ab69 · outbound

This paper cites Zico Kolter.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Zico Kolter

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:27.665893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.620795Z digest=sha256:3b2fe77b7a1ab44d05b9809e723ca07303641177e8b296faaaa3c22d31c7bd62

Observation 7f6173b9-8bbc-4803-a0ec-d755b9c17f1b · outbound

This paper cites Multiplexing dynamic deep learning workloads with slo-awareness in gpu clusters.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Multiplexing dynamic deep learning workloads with slo-awareness in gpu clusters

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:27.515390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.624755Z digest=sha256:8c5ddedd610a27c391afd24512ee0e422184db026376948ba37de2a85089ef9f

Observation 43514fe2-a55b-48e3-862c-9f96811117c8 · outbound

This paper cites Cloudnativesim: A toolkit for modeling and simulation of cloud-native applications.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Cloudnativesim: A toolkit for modeling and simulation of cloud-native applications

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:27.346301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.628509Z digest=sha256:bbc1df282c3da0ef97a527c605845808e8aac6d275b396399c97e375c62b409f

Observation ab4ad36b-d4ea-4784-8c2b-96f02774f583 · outbound

This paper cites Llminaflash:Efficientlargelanguagemodelinferencewith limited memory, 2024.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Llminaflash:Efficientlargelanguagemodelinferencewith limited memory, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:27.162975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.632443Z digest=sha256:e2ecc27c6484f3ee1f0f24ee5434555366f2e617356a191fbe88e346723d830d

Observation 59bf1481-178e-4fc5-b5d8-ae5b96ba2e8a · outbound

This paper cites Spotserve: Serving generative large language models on preemptible instances.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Spotserve: Serving generative large language models on preemptible instances

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.636501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.636501Z digest=sha256:1fb830f4fc83198f1337e0e37fd3cda01adf6af2b038a18a8f1ebebff363422e

Observation 0736df2e-d30d-49e9-8f47-458134c404ed · outbound

This paper cites Llumnix:Dynamicschedulingforlargelanguage model serving.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Llumnix:Dynamicschedulingforlargelanguage model serving

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:26.979547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.640277Z digest=sha256:c0aa8b4ca478912d7579c29a75b66401a0f887680e9ec6e1cea8107f953b5335

Observation 559ffcaf-15c7-4c90-8d29-7712367d1853 · outbound

This paper cites Serving heterogeneous machine learning models on multi-gpu servers with spatio-temporal sharing.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Serving heterogeneous machine learning models on multi-gpu servers with spatio-temporal sharing

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:26.787021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.643981Z digest=sha256:e57553cdf0ddb01f305bdaefbfce2c4a287a68d6ff6f8d8d0ecbb04d017dec7a

Observation 9c86c373-b4bc-422f-b8be-c2e46e1f2c0b · outbound

This paper cites Inferline:latency-aware provisioningandscalingforpredictionservingpipelines.In Proceedings of the 11th ACM Symposium on Cloud Computing, pages 477–491, 2020.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Inferline:latency-aware provisioningandscalingforpredictionservingpipelines.In Proceedings of the 11th ACM Symposium on Cloud Computing, pages 477–491, 2020

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:26.604327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.647609Z digest=sha256:295f0d70f61b23c03e9c8570d360d9cedaeaa33a65de8b6f1ac872301e852e09

Observation c8937796-4580-4d32-a58c-fee8349597d3 · outbound

This paper cites Optimizing llm inference throughput via memory-aware and sla-constrained dynamic batching, 2025.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Optimizing llm inference throughput via memory-aware and sla-constrained dynamic batching, 2025

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:26.431926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.651829Z digest=sha256:39bb0fe283ce5a29f3ed85c9625e7c2a30d42b18402b0f6b05c63556ebce9ef6

Observation ae8538a9-5dcf-4c02-88d2-141892950ff8 · outbound

This paper cites Alloystack: A library operating system for serverless workflow applications.Pro- ceedings of the Twentieth European Conference on Computer Systems, 2025.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Alloystack: A library operating system for serverless workflow applications.Pro- ceedings of the Twentieth European Conference on Computer Systems, 2025

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:26.270323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.655748Z digest=sha256:2880d8f6e68bb6e3a56591a43880a2888cb6b987d8733d0887fdf04d6173d03c

Observation 0777050f-f079-44e5-8789-ffd80508e5ea · outbound

This paper cites Lora-flow: Dynamic lora fusion for large lan- guage models in generative tasks.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Lora-flow: Dynamic lora fusion for large lan- guage models in generative tasks

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:26.052390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.660020Z digest=sha256:10fd007e747cd4066ff03fdef7612d6ac31ab43ab9f2fbb9bb6ff8e9d14d8876

Observation f1344fe9-90df-4ba6-8d48-ab40f74d133f · outbound

This paper cites Accessed: Apr.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Accessed: Apr

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:25.830080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.663827Z digest=sha256:70f8cf46c8d1cd7ac850b302c56c983203f9bc0d62e2e777f463b36dbc5e2dd1

Observation 01a69b05-5d9b-463a-b295-4217399f77da · outbound

This paper cites an unresolved cited work.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:45:25.642539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.667723Z digest=sha256:8c7a3840267dde3733167b57eae9bb2d88e4c683205270d690cc94556cbb2ce0

Observation 2213fac0-1631-4eec-bdc7-3978d02e6221 · outbound

This paper cites an unresolved cited work.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:45:25.486001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.675885Z digest=sha256:19fa5f2c1f0227d9250c4434baf7e6b37f24ac0236b6bdfbd0639cdf3588c95a

Observation 1421a709-5b0d-4168-99c6-cff475463e4b · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.681562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.681562Z digest=sha256:ea5ad86aba0c6fc7b78cebb371cd32c5ecab3d15b311241cc5ecd306f6fa4fe0

Observation 25fa635e-8597-4a2c-a108-282095a2f577 · outbound

This paper cites Llm inference serving: Survey of recent advances and opportunities, 2024.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Llm inference serving: Survey of recent advances and opportunities, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:25.265670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.686206Z digest=sha256:e19e22c16d5b39abf79a5af0692b4f968d5a15cf724b96ddc942a56b56adad6c

Observation 85001b70-bf8c-41bd-9932-9291d297961a · outbound

This paper cites Optimizing mixture-of-experts inference time combining model deployment and communication scheduling, 2024.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Optimizing mixture-of-experts inference time combining model deployment and communication scheduling, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:25.101053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.690334Z digest=sha256:e95aacff0dd453c71e4d4167f853c7ecc06196f0acc53c5dad8eaa877e0a9e03

Observation 3df35fa0-3e6a-4d50-95ce-866af36fc334 · outbound

This paper cites Mooncake: A kvcache-centric disag- gregated architecture for llm serving, 2024.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Mooncake: A kvcache-centric disag- gregated architecture for llm serving, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:24.909654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.694330Z digest=sha256:f0d11f761ed4b17823ed70439f2173515013008702e6a73f087740c19c6ba0e4

Observation ec14dcb2-cd43-4691-8eb5-08930a6dfee8 · outbound

This paper cites Spinfer: Leveraging low-level sparsity for efficient large language model inference on gpus.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Spinfer: Leveraging low-level sparsity for efficient large language model inference on gpus

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:24.699198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.698281Z digest=sha256:1e0b6982ff6e90d33bc8d83046772d82b3d2bab2b5ac577a6d456b89cf0c9bcc

Observation 1ecc3806-31e9-491d-ae80-e0550faa2b0b · outbound

This paper cites H2o:Heavy-hitteroracleforefficientgenerativeinference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling H2o:Heavy-hitteroracleforefficientgenerativeinference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:24.511967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.702463Z digest=sha256:4133fb93d5c428e7bd88b8462bd82aab0e3609a57987c2da4a581c061c2f15b6

Observation 76faaf30-a734-4153-8442-4a54223cfa77 · outbound

This paper cites Splitwise: Efficient gen- erative llm inference using phase splitting.2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA), pages 118–132, 2023.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Splitwise: Efficient gen- erative llm inference using phase splitting.2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA), pages 118–132, 2023

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:24.286745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.706607Z digest=sha256:7cad6c38b6918ccf4dec0fd6634763e9d54d49b418d30dc1117a39482af1482a

Observation c4c725c7-5a59-4d32-8bf4-3e95d7168ab7 · outbound

This paper cites an unresolved cited work.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:45:24.086241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.710774Z digest=sha256:6a8901e182004c7764641ea94bbead97260ebe3fafe59cdf4115b625ccd71a68

Observation 43ce76af-0cca-4fae-bae4-1514640d0afa · outbound

This paper cites Inference without interference: Disaggregate llm inference for mixed downstream workloads, 2024.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Inference without interference: Disaggregate llm inference for mixed downstream workloads, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:23.843614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.714736Z digest=sha256:db418c2cb86cfbbf81d1aa56a4ce94ed96544e54f540b2e2c06e9abb44ff4f14

Observation 904791a1-6169-4a39-a9f4-18cb5f46ec30 · outbound

This paper cites Dynamollm:Designingllminferenceclustersforperformance and energy efficiency, 2024.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Dynamollm:Designingllminferenceclustersforperformance and energy efficiency, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:23.669168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.719014Z digest=sha256:5665afdd118aaf5184aa855621da64f64760f23459d64329c73a7653f987aace

Observation d9ebb7f2-f091-404d-8285-19b0415de7b8 · outbound

This paper cites Skyserve:Servingaimodels across regions and clouds with spot instances.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Skyserve:Servingaimodels across regions and clouds with spot instances

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:23.437689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.722857Z digest=sha256:94530f1cbc82e63ad540e44dcae6671fd18fbdb6b3690ec0284b4c59c9961f15

Observation d650dde7-70e7-4800-9354-556790f414ee · outbound

This paper cites Towardsefficientandreliablellmserving: A real-world workload study.arXiv preprint arXiv:2402.XXXXX, 2024.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Towardsefficientandreliablellmserving: A real-world workload study.arXiv preprint arXiv:2402.XXXXX, 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:23.227733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.727165Z digest=sha256:2f0521d8844072b23c11a1327bf55b23d848414fde5faa0b0c5a1694883b30c5

Observation 50e2209c-e0bc-4164-972b-9d186e868a2e · outbound

This paper cites Usher: Holistic interference avoidance for resource optimized ML inference.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Usher: Holistic interference avoidance for resource optimized ML inference

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:23.012903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.731400Z digest=sha256:cb0daea22006077754807a91481ae31b4677c12c62a51dde3e8a0d8429559a9f

Observation 71fe8728-f510-4dde-8c3a-93b4cea88868 · outbound

This paper cites an unresolved cited work.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:45:22.779244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.735534Z digest=sha256:5d3fdaca2dcfb882c29422cc1a2d6337de811eba486cbdd8f822d50d5859ee2c

Observation 515dbeab-5865-4a1e-89d1-4b2dbc2a8595 · outbound

This paper cites Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:22.599134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.739514Z digest=sha256:1774eba36f7eaa6b8aa1008f25cbce3de48a566abf7df84eb201c97029c6e78f

Observation b17279ee-733a-4137-87dc-7c2105163f89 · outbound

This paper cites Alpaserve: Statistical multiplexing with model parallelism for deep learning serving.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Alpaserve: Statistical multiplexing with model parallelism for deep learning serving

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:22.387183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.743488Z digest=sha256:bf124db347ea480ca236073f03061b0fa37d42545a3cc4cf3c0a3ca3a90c0bea

Observation 3db551ec-4c4a-449c-8a13-15ea022f6fdf · outbound

This paper cites an unresolved cited work.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:45:22.205088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.747191Z digest=sha256:97e5f2d7a6dd3ebb7f68b2566de7e82457cb1288703ef1937bba2cd3506cab59

Observation 8e952f21-82e8-4dfb-aeab-26bb3c1c3e81 · outbound

This paper cites Yadwadkar, and Christos Kozyrakis.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Yadwadkar, and Christos Kozyrakis

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.999977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.750754Z digest=sha256:a94b224b657b2211f0848bde665115636417bec2ff3bdaf4c9a7663ff5b0a714

Observation 29feca7a-4ba7-4d1c-870a-f553e4fdb36f · outbound

This paper cites Reducing Activation Recomputation in Large Transformer Models.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Reducing Activation Recomputation in Large Transformer Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.758677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.758677Z digest=sha256:654bd22e5bf2917cd6eac416829ffb382b50ee83353f0b70dfb43a3308b7aa04

Observation 6192e0ab-e576-407f-938f-4a21570f9f37 · outbound

This paper cites On parallel processing systems: Amdahl’s law generalized and some results on optimal design.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling On parallel processing systems: Amdahl’s law generalized and some results on optimal design

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.610721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.762745Z digest=sha256:120943577d29e8332412574c5780aa094c7519dd99ef4ad37714fdf7672ba69b

Observation 3607326f-b68b-41b0-bfc0-29a9cb5085ec · outbound

This paper cites xformers: A modular and hackable transformer modelling library.https://github.com/facebookresearch/ xformers, 2022.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling xformers: A modular and hackable transformer modelling library.https://github.com/facebookresearch/ xformers, 2022

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.412262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.766581Z digest=sha256:fc58f2670c5a3ba144e9eaae58caac88fb86fa5788e34fb8b36fa098adc64244

Observation b995e9ff-346e-43c5-a422-8101a97dc13e · outbound

This paper cites https://developer.nvidia.com/management-library-nvml, 2025.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling https://developer.nvidia.com/management-library-nvml, 2025

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.313583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.770284Z digest=sha256:e10f5277cbde4da950ed11b66f660921178af32d1805c2b791362c05744a29da

Observation 33dca7d2-896f-4afa-a787-f6a83496e92f · outbound

This paper cites Orca: A distributed serving system for transformer- based generative models.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Orca: A distributed serving system for transformer- based generative models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.283738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.774358Z digest=sha256:77eaff8d54584baaf0c19fb9c59912c9cdf8f033219fee34dfc51494249578d5

Observation 080e1029-b4cb-43b7-9733-66fb8c5041d6 · outbound

This paper cites Uellm: A unified and efficient approach for large language model inference serving.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Uellm: A unified and efficient approach for large language model inference serving

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.245867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.778185Z digest=sha256:c245895f7e8df1ecd68cdccf40d8493d0d9e15eab50d8477fd7ba59d69f2500f

Observation 85ae5a38-9a43-400b-9e93-4dcf3748b8cf · outbound

This paper cites Stanford alpaca: An instruction- following llama model.https://github.com/tatsu-lab/stanford_alpaca,.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Stanford alpaca: An instruction- following llama model.https://github.com/tatsu-lab/stanford_alpaca,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.202598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.781963Z digest=sha256:fbaa35e4c6bbc9a0dce09eced225b853600392594898296cbd1da464c651a746

Observation f8ce4997-8690-4e79-92bb-6d436cbc5305 · outbound

This paper cites Mepipe: Democratizing llm training with memory- efficient slice-level pipeline scheduling on cost-effective accelerators.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Mepipe: Democratizing llm training with memory- efficient slice-level pipeline scheduling on cost-effective accelerators

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.113047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.789659Z digest=sha256:bc79fce5f9cd5c3a353f05e19b4754bb800f44c0087e0101efa846ca7d6ed72b

Observation d44c4132-fcb8-4e9b-b5eb-fe849f84db51 · outbound

This paper cites Le, and Z.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Le, and Z

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.078885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.793446Z digest=sha256:223f4733e2d31fdb7638a387997091e5f1039374e8e892296d85d666619528a6

Observation ab48421c-b07c-4781-905c-e160553b1ac6 · outbound

This paper cites Mist:Efficientdistributedtrainingoflargelanguagemodelsviamemory- parallelism co-optimization.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Mist:Efficientdistributedtrainingoflargelanguagemodelsviamemory- parallelism co-optimization

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.055530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.797367Z digest=sha256:7e4d380c073513cf9af7b6d2e91f96a7847f0bef3b3bf800157913e868adcff1

Observation b2b93895-5df9-4f87-91b4-4b5a78aab060 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.801131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.801131Z digest=sha256:c92e1cd7244c913d30085855a7abbc36313b551f46e94ef7e50ba2ded7244cc1

Observation 848a515f-29f3-418e-8124-111d46d06335 · outbound

This paper cites Alpa: Automating inter and intra- operator parallelism for distributed deep learning.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Alpa: Automating inter and intra- operator parallelism for distributed deep learning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.041953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.805316Z digest=sha256:2d6f04b648c99b60deec047b42b10541517eaf2bb4a48a597cb5714f92adde5f

Observation a7567b5a-d5ea-4aff-aee2-d1c5c27566fc · outbound

This paper cites Fast state restoration in llm serving with hcache.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Fast state restoration in llm serving with hcache

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.028193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.809631Z digest=sha256:857a307fe42d2d5c4717c8644d789b38dbe37db0bc5f4a7613e237f9ec38dd17

Observation 5b2ca156-3b79-4afc-8b43-a3107e2d25aa · outbound

This paper cites Fast and live model auto scaling with o(1) host caching, 2024.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Fast and live model auto scaling with o(1) host caching, 2024

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.013948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.813723Z digest=sha256:812ee433696108212562dab30c791f0d41764df9107a67e68af8a18efc070db3

Observation 106593b2-91f9-4854-87af-93f5faf2cc46 · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.817827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.817827Z digest=sha256:d6924bb170054686d52cda02c076c3260415198d8c16fc7fca54295378d07955

Observation d07250cf-5f05-4c41-8fdb-e4ecfcde3f15 · outbound

This paper cites Infinigen: Efficientgenerativeinferenceoflargelanguagemodelswithdynamickv cache management.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Infinigen: Efficientgenerativeinferenceoflargelanguagemodelswithdynamickv cache management

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:20.991136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.821661Z digest=sha256:9820dc34262924bbff91760ec25116c8ca91ddc16b6a6a0e4d12400327aab3fb

Observation bb71e73a-2187-431a-8428-e8c3931b8909 · outbound

This paper cites an unresolved cited work.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Unresolved cited work

Reference 411

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:45:21.824444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.754810Z digest=sha256:f24ed7454cc6589a16c33664dbc3ba64bfdc233c6fa99bddb20377a2d839e55d

Observation a18aea7a-6abb-454a-a786-988f05df980b · outbound

This paper cites an unresolved cited work.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Unresolved cited work

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.671778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.671778Z digest=sha256:85c33241e8d9e9daf362cf195f5351daa95293eedd0c83513753ebe9a2032e6a

Observation 2ef497c9-742a-44aa-9990-dbfd9f1c122c · outbound

This paper cites an unresolved cited work.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:45:21.165178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:45:20.785956Z digest=sha256:fb160dabbca363d4094f51159a4f73fc10b0bcc76d9be5812bf69b2a90e128e8

Pith citing papers

No inbound Pith citation observations are available.