Pith. sign in

Paper Citation Record · LEDGER

BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2401.17644.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.17644 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:24:02.464919Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-07T20:34:09.759421Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation decb8d3c-e512-4ab7-a628-a57eb2b26359 · inbound

Topology-aware Preemptive Scheduling for Co-located LLM Workloads cites this paper.

Topology-aware Preemptive Scheduling for Co-located LLM Workloads BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T18:28:01.438721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:28:01.438721Z digest=sha256:b70f4b22947642de3aa152ad7604e432fb42d932dfb6b2b022dfa2290584966d

Observation a1a7915f-97aa-48e9-9234-8724b843e290 · inbound

An Empirical Characterization of Outages and Incidents in Public Services for Large Language Models cites this paper.

An Empirical Characterization of Outages and Incidents in Public Services for Large Language Models BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T17:14:06.597467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:14:06.597467Z digest=sha256:95d0ae736359521185b86344cf753a13e60e99fbed6e2564b9a55289b697d952

Observation 1b3aaba0-7d6d-4784-beb9-01d826be92c8 · inbound

Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs cites this paper.

Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T18:02:12.494628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:02:12.494628Z digest=sha256:2f3b0926082c1e632c92ec0cb9c629a5cba98c414c936a37dfc4cb6f15546b39

Observation 71049746-53cc-4c9a-9454-e5f0143eed05 · inbound

SecPE: Secure Prompt Ensembling for Private and Robust Large Language Models cites this paper.

SecPE: Secure Prompt Ensembling for Private and Robust Large Language Models BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:28.834585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:28.834585Z digest=sha256:6b59b65f5ca2cd38cb49956e2cb5d96cf280a9491f1a148f42249e1c2a8cc28f

Observation 5d64ee6f-3385-4f1c-b491-5cf482f51275 · inbound

RAGDoll: Efficient Offloading-based Online RAG System on a Single GPU cites this paper.

RAGDoll: Efficient Offloading-based Online RAG System on a Single GPU BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T12:24:02.464919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:24:02.464919Z digest=sha256:4496539362d3e3a1e2ebba64364bb84e7c6d2ece4aad42e639ab3b2b56c1113c

Observation 4124632a-afae-4454-b462-ec55e7c25d5c · inbound

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference cites this paper.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:17.378835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:17.378835Z digest=sha256:3f6ebcfe28fff820347ad8ce5f5db723f409e14adf1d0f1611c6a5a6cebb0ad2

Observation 0acdf2cc-a8b7-4153-9ee0-25fa153b28c9 · inbound

MuTRAP: Multi-trigger Trojans Attacking Robot Task Planning Systems cites this paper.

MuTRAP: Multi-trigger Trojans Attacking Robot Task Planning Systems BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T10:55:33.413030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:55:33.413030Z digest=sha256:0b8c2243bd597971427d4c2e3df6ffa08bf732443189f5940e8eef68846469fc

Observation 8ed1a345-b4ac-400f-bee9-f55e2e36dfd2 · inbound

Energy Considerations of Large Language Model Inference and Efficiency Optimizations cites this paper.

Energy Considerations of Large Language Model Inference and Efficiency Optimizations BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T10:37:51.489709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:37:51.489709Z digest=sha256:b62a70d0d01399caf7efdda68afe4c0bb0fa297e0f5adc66c1f88a21f4007be9

Observation 83f8b1b0-af64-4547-ae1d-9e31630141f0 · inbound

ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production cites this paper.

ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-22T15:44:58.154778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T15:42:05.266854Z digest=sha256:e4b9b587aeec5ead39e2d148aefacee1b1f93ac8cd806655395930d2f1a9010a

Observation b98b6759-a8c0-47e6-adb8-b97827578fab · inbound

Past-Future Scheduler for LLM Serving under SLA Guarantees cites this paper.

Past-Future Scheduler for LLM Serving under SLA Guarantees BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:20.940919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:20.940919Z digest=sha256:6939bb8b130f912ab3342dd123c396c7df3bf5cdd771a4679cc3314ffb3c2059

Observation c8f02fcc-0387-4b39-9431-998ed626a0a2 · inbound

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts cites this paper.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.499001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.499001Z digest=sha256:dc1e25802ebc9eba0794513b2ee942f1f38b032879a07ce866345d76f5ad6938

Observation c132c0cb-e3cd-44e3-8c06-33ebda30f0fa · inbound

Rethinking Caching for LLM Serving Systems: Beyond Traditional Heuristics cites this paper.

Rethinking Caching for LLM Serving Systems: Beyond Traditional Heuristics BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T16:21:09.370531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:21:09.370531Z digest=sha256:24e6a79087e8387e17842c57fec55c4e3ce38e091d4682b2ba098a8c234123cb

Observation cf06d08f-ec08-47af-8b66-4e8f0b46a461 · inbound

Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism cites this paper.

Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T20:53:33.420816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:53:33.420816Z digest=sha256:562b9b86036c58a6b563c15f6b8b4ef51495fa95b74eb1a97063fb39c6f1a238

Observation de068e43-80e3-4e2f-a387-5f7c1b53a9c6 · inbound

WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving cites this paper.

WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:24:51.184801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T12:22:56.972437Z digest=sha256:397e6dd065d1c919bbb8689dc0d8a7081aeaea8a217a6eb9da644e0599fda096

Observation 3b320243-1c8e-436f-8bc9-20baaf6d93fa · inbound

Analytical Provisioning for Attention-FFN Disaggregated LLM Serving under Stochastic Workloads cites this paper.

Analytical Provisioning for Attention-FFN Disaggregated LLM Serving under Stochastic Workloads BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:57:43.041309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T09:53:13.057587Z digest=sha256:f4586b02017544c910f9d0626fca07c82728b37297af47238c744df6d74e9dca

Observation 5fb16b47-dbc8-4e6b-9f14-009f7ce6787c · inbound

A Techno-Economic Framework for Cost Modeling and Revenue Opportunities in Open and Programmable AI-RAN cites this paper.

A Techno-Economic Framework for Cost Modeling and Revenue Opportunities in Open and Programmable AI-RAN BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:43:32.076705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T00:41:18.795730Z digest=sha256:647535cd971b0c625d80e29d1d89ec1dec6e0435c1b97de597d3d71cbef5144c

Observation 3fd1657e-0628-474d-bcbe-2a55a2842876 · inbound

A Techno-Economic Framework for Cost Modeling and Revenue Opportunities in Open and Programmable AI-RAN cites this paper.

A Techno-Economic Framework for Cost Modeling and Revenue Opportunities in Open and Programmable AI-RAN BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:57:40.185779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T16:54:58.406395Z digest=sha256:49bb9a8ee524f84fa32955612878376784cb745bbb78b175600923c2e2092cb2

Observation 3a3fd922-297f-47cc-b1d0-44dfbbe1ce38 · inbound

Dual-Pool Token-Budget Routing for Cost-Efficient and Reliable LLM Serving cites this paper.

Dual-Pool Token-Budget Routing for Cost-Efficient and Reliable LLM Serving BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:01:02.497582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T17:49:45.604648Z digest=sha256:7573d6b26bb04b47821d84f1c44fa954d0e9cd326122e34ca70fcdc3a7561835

Observation ee4eb12b-bb62-4539-b8ff-3e76ec3fefab · inbound

When Is the Same Model Not the Same Service? A Measurement Study of Hosted Open-Weight LLM APIs cites this paper.

When Is the Same Model Not the Same Service? A Measurement Study of Hosted Open-Weight LLM APIs BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:11:14.189996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T01:27:54.164991Z digest=sha256:6df07597dbcf79f0c6ea32d4257ffa10e3ec61a3e5381f5535ce753721ada6c0

Observation dcf62596-9c8b-42cf-a116-9b9510ef970e · inbound

When Is the Same Model Not the Same Service? A Measurement Study of Hosted Open-Weight LLM APIs cites this paper.

When Is the Same Model Not the Same Service? A Measurement Study of Hosted Open-Weight LLM APIs BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:56:06.140516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T16:57:26.334739Z digest=sha256:a03dd51969be43b382a3f21f4df9708d717b7cf6b5c77aa6cec460c86ef0f261

Observation 1adb8f03-79d1-4a81-89f2-4af738e84889 · inbound

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL cites this paper.

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:31:16.212355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T05:14:14.168753Z digest=sha256:a7b23b5fe1e82294086c9a30e1711aef2922c0ff513370bf0df0e3d77e32085e

Observation 34eb5215-4a35-4947-8f3d-1b84e4ed133e · inbound

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL cites this paper.

ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:39:53.241704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T08:39:31.911497Z digest=sha256:96e1a5ce610dd29f7131837ad2ebbe44a8c38ac137b80f2186aeb4576359f4e6

Observation 872c7ed5-4293-4234-8d13-1e34e45fb66b · inbound

ADAPT: A Self-Calibrating Proactive Autoscaler for Container Orchestration cites this paper.

ADAPT: A Self-Calibrating Proactive Autoscaler for Container Orchestration BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T19:27:43.466513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T19:26:15.876258Z digest=sha256:fd8a1813c909bd09928cc188c0731d0c033f5c6a4f0fb3736c21d03ef27e32be

Observation e13b4565-a6f8-4007-b95e-e7e971669804 · inbound

GEM: GPU-Variability-Aware Expert to GPU Mapping for MoE Systems cites this paper.

GEM: GPU-Variability-Aware Expert to GPU Mapping for MoE Systems BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-20T04:33:03.850522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T04:31:11.271729Z digest=sha256:d424b2ab299295d91659751fc5a9936f72d5c9e1eac09cfa49f080caea5e5239

Observation 3a858c52-9166-43aa-adb4-c786d7515948 · inbound

Beyond Per-Token Pricing: A Concurrency-Aware Methodology for LLM Infrastructure Cost Estimation cites this paper.

Beyond Per-Token Pricing: A Concurrency-Aware Methodology for LLM Infrastructure Cost Estimation BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-03T12:48:11.790687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T08:45:42.781160Z digest=sha256:28397a91c94bf9586c19ca32ce27f40f5341442b1133a34d41ff7cb864f762e6

Observation bf706967-bcaa-4511-b468-cec43df32685 · inbound

CoCoScale: Leveraging Layer-wise Scaling to Unlock the Potential of Online LLM Serving cites this paper.

CoCoScale: Leveraging Layer-wise Scaling to Unlock the Potential of Online LLM Serving BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-11T21:08:24.159706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:08:24.159706Z digest=sha256:82fe2fc9cc9b02c400d9ef9b048e51990e8fa88db6b0c31391f6961fd5f33b67

Observation bed8a48e-ad5b-4099-8241-2d55c718d579 · inbound

Adaptive Inference Batching using Policy Gradients cites this paper.

Adaptive Inference Batching using Policy Gradients BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-07-07T20:34:09.762493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-07-07T20:29:21.563562Z digest=sha256:5dbdf7222266ad1e1d34508cc20fd2ca62cfeed3488787b17124bf08bed5e0b4

Observation bcb72876-5e1c-4716-bd4e-1a5261286d4e · inbound

Trusted Floors Under Untrusted Learners: A Runtime Assured-SLO Guard for ML Serving cites this paper.

Trusted Floors Under Untrusted Learners: A Runtime Assured-SLO Guard for ML Serving BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T07:31:31.140358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:31:31.140358Z digest=sha256:4b9b9ea1aede802a40d820c20139e481ce6b80c55547473452ddbc4f7493f8e0

Observation f12090fb-f6f4-4b0c-aa62-55d47fcd2087 · inbound

TokTier: Exact Stateful CPU+GPU Tokenization for Agentic LLM Serving cites this paper.

TokTier: Exact Stateful CPU+GPU Tokenization for Agentic LLM Serving BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T02:07:01.398171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:07:01.398171Z digest=sha256:80a6aca6894c567e477860ea32218eedc2ffbc2490a2a4ff353182be2fa245ef