Pith. sign in

Paper Citation Record · LEDGER

DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 35 inbound Pith citation observations for arXiv:2401.08671.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.08671 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:55:54.042509Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T11:59:50.582581Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9ea946e8-56ae-4cfe-9859-67f8ea75d258 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 279

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:39:33.490879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:06289e013be1522368ee3ed055e28f58cbd1a47581424eced6aa009a86ba0fea

Observation 852d3521-7c90-462c-832a-bcc96f884e8b · inbound

HybridFlow: A Flexible and Efficient RLHF Framework cites this paper.

HybridFlow: A Flexible and Efficient RLHF Framework DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:53:38.983721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T07:53:38.715353Z digest=sha256:9fb6cc812660d465d6eac4518d643f582f02185949dfd02b5f1fb53d85ab3e3a

Observation a129fb99-037a-4081-a384-96d4579e7a79 · inbound

BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching cites this paper.

BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-23T16:58:11.937207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T16:57:46.645061Z digest=sha256:2e21c7265b1be9899d47f0bc96998c1ed63931f0a697bffe25212ef182fef224

Observation 0c8721d8-0d08-4ff2-ad35-51587c21c02f · inbound

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design cites this paper.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:57:40.262382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:5d37672b1ea950337b936574030ff3a153a09ad3a2f14f03cdd5e62beb3b0eab

Observation cca8c0a0-3680-428a-9ec0-d540affbe1ac · inbound

Klotski: Efficient Mixture-of-Expert Inference via Expert-Aware Multi-Batch Pipeline cites this paper.

Klotski: Efficient Mixture-of-Expert Inference via Expert-Aware Multi-Batch Pipeline DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T17:55:54.042509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:55:54.042509Z digest=sha256:3a4da66639ae700236338353c37f22c305333f19b0d5a98e48c0b7ff4ea1078e

Observation 6bedb766-c3d8-4009-a6f4-8cabf5d05c7c · inbound

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees cites this paper.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.872509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.872509Z digest=sha256:1fae67ba280544561033f8dfb36d96d97305d74a5c088dc71ba588c72b2d2267

Observation ebbd282c-edd5-43e1-ae96-9cf2bf0587d4 · inbound

MIST: A Co-Design Framework for Heterogeneous, Multi-Stage LLM Inference cites this paper.

MIST: A Co-Design Framework for Heterogeneous, Multi-Stage LLM Inference DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:17:08.081299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T21:16:31.655330Z digest=sha256:b63546eb19228f5c610c8893724eb01bd4ccd0b73be742455238eb8ea681f9ad

Observation b06f25c7-381a-4421-a4e4-1855b33a2e6f · inbound

ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production cites this paper.

ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-22T15:44:58.084998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T15:42:05.266854Z digest=sha256:ff93c696430c4cabd1b08b1fc11cfd4deed241c4fe6525dcd535e5b5249e39a9

Observation 313443c8-0e08-4e59-b929-60f45ebddbc7 · inbound

AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity cites this paper.

AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:52.961383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:52.961383Z digest=sha256:7ea40e11cdbd9ba06c4f4dd4090f5ae9449c6d805294ef3c316c4ac43553e085

Observation 17c4b2d3-7d80-4f7e-a321-478c29739989 · inbound

Rectified Sparse Attention cites this paper.

Rectified Sparse Attention DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.938385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.938385Z digest=sha256:b8bdafa6450187f61779b5480531061c2153c719ef283a72dc2e63f3c68d7da7

Observation 3f75106c-2cda-4921-a977-c6699722549e · inbound

Chat-Ghosting: A Comparative Study of Methods for Auto-Completion in Dialog Systems cites this paper.

Chat-Ghosting: A Comparative Study of Methods for Auto-Completion in Dialog Systems DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:19:10.069938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:19:10.069938Z digest=sha256:33b8a319e0d78184809bc4f66c641980dd2545c2594bc874f6cb25fccbcfd255

Observation f846ec65-36d7-4250-8207-3cbfadd1c795 · inbound

Rethinking Caching for LLM Serving Systems: Beyond Traditional Heuristics cites this paper.

Rethinking Caching for LLM Serving Systems: Beyond Traditional Heuristics DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T16:21:09.103693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:21:09.103693Z digest=sha256:0636156520646e57b6572c80c6f92326362af28f27975ab49e1b5f99d48f9eba

Observation 804a2c0c-1dc3-47dc-a9f9-82497e7d3e21 · inbound

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference cites this paper.

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T15:53:29.009532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:53:29.009532Z digest=sha256:50bde247e7d272c7f0491eddd02c5d87a087997ba5aa6fb33327d448cd85c2b0

Observation c81406e8-a4fc-4464-adde-fb4e1b76914f · inbound

WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving cites this paper.

WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:24:51.224216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T12:22:56.972437Z digest=sha256:6539eb4b6b0f52a305b590b7d286a9d3e0bba21a321036777863a169b664beab

Observation 268c5e27-d697-4056-a7a9-becc9c17b726 · inbound

PrefixWall: Mitigating Prefix Caching Side Channels in Shared LLM Systems cites this paper.

PrefixWall: Mitigating Prefix Caching Side Channels in Shared LLM Systems DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-21T12:15:06.799217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T12:14:09.509302Z digest=sha256:2ae7967c87c0fda5679d2a908dd8931cbedc7b035f45cf84b18322a27c7c5004

Observation ffa085a0-9695-41a7-8065-f41d1ffbd4c5 · inbound

FASTER: Rethinking Real-Time Flow VLAs cites this paper.

FASTER: Rethinking Real-Time Flow VLAs DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T08:05:15.319681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T08:02:13.188363Z digest=sha256:f52fc8038eaaa350df417b71c8d444146e8a04987c34722769b48effbea232d4

Observation aaf05164-07a0-4f3b-a171-ee9197516154 · inbound

FASTER: Rethinking Real-Time Flow VLAs cites this paper.

FASTER: Rethinking Real-Time Flow VLAs DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T10:50:01.062812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T10:48:56.280105Z digest=sha256:c36805404056efd852050d7f2f78f9fb0d8d5fe61138802a714ca1c7e65630fe

Observation 9c00d7a0-282f-44cf-acd1-f1192872aa99 · inbound

Towards Efficient Large Vision-Language Models: A Comprehensive Survey on Inference Strategies cites this paper.

Towards Efficient Large Vision-Language Models: A Comprehensive Survey on Inference Strategies DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:43:00.670264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T21:40:55.343169Z digest=sha256:279752c1f88d73e5b9432cbd56ff01885dd6f6f4b7be4c99debcfe6fd5129292

Observation 876815a3-5aa7-486c-8378-6ce3c7e3a195 · inbound

Characterizing Performance-Energy Trade-offs of Large Language Models in Multi-Request Workflows cites this paper.

Characterizing Performance-Energy Trade-offs of Large Language Models in Multi-Request Workflows DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:30:00.274373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T12:28:38.809950Z digest=sha256:7a70e895c0a553643be4fd080339b1519893359bf8d5cd9b6f92539ca7cb2c82

Observation d57681e3-f3fb-42fc-8982-ad26d7846e04 · inbound

Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines cites this paper.

Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:59:03.464710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T09:52:40.057014Z digest=sha256:f89eb911f151eb99198c3ae0e790501f565646183c2ce1d279ffc6c74ed93dda

Observation 61b831f4-fcd0-47ff-adc3-18b849d53a70 · inbound

Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities cites this paper.

Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:16:10.210742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T09:45:57.201837Z digest=sha256:20a7fd50690e29901c726a47d0bba5450b6a166a46c11a8f453e7e19dcacabab

Observation c0ab0b09-ffc3-4798-9256-856c3064c131 · inbound

MARS: Efficient, Adaptive Co-Scheduling for Heterogeneous Agentic Systems cites this paper.

MARS: Efficient, Adaptive Co-Scheduling for Heterogeneous Agentic Systems DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:35:35.686543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T14:30:56.899306Z digest=sha256:6f8a7061f81905eb79e436a8b33c9e135b24241039daa2ed27c4a195a114618c

Observation bdb1d749-6c65-4651-bbda-4dec91c8d58d · inbound

Federation of Experts: Communication Efficient Distributed Inference for Large Language Models cites this paper.

Federation of Experts: Communication Efficient Distributed Inference for Large Language Models DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:51:07.551006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T13:40:50.411198Z digest=sha256:681de4106adfa90d7c9b475e849e75e38003833a19bb6a80d89aedff89364f25

Observation 4baef794-3b5c-4841-b783-3006804077fb · inbound

Focused Forcing: Content-Aware Per-Frame KV Selection for Efficient Autoregressive Video Diffusion cites this paper.

Focused Forcing: Content-Aware Per-Frame KV Selection for Efficient Autoregressive Video Diffusion DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:23:14.150643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T11:19:21.721494Z digest=sha256:ed9163ff4801cd90c12f0cb3da85e4d91453db2d093e6173261006969f8d37b6

Observation 13fd3d8f-a36d-4899-832c-359a99beb400 · inbound

AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System cites this paper.

AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:15:17.294037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T03:12:49.028342Z digest=sha256:519aada1af9e664f1bae078c0ebe53f7ebc9d8e5fbcd5e9d2292b0d815ff99b8

Observation 83d9298f-7b54-40a6-9b32-f483749567fe · inbound

Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving cites this paper.

Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:56:25.651067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T12:56:16.768455Z digest=sha256:a87e17f0967989b758c9a2e67d8a55516b4669e6337406b897660608aff71ee0

Observation 6e215635-2261-408f-9102-c420675d9305 · inbound

Multi-Segment Attention: Enabling Efficient KV-Cache Management for Faster Large Language Model Serving cites this paper.

Multi-Segment Attention: Enabling Efficient KV-Cache Management for Faster Large Language Model Serving DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:36:26.302619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T11:38:05.435505Z digest=sha256:c6280cc80e1a6222fb6e6187aa9413ad7a949ecf9717fe02b64bd7b44a25f8ba

Observation f690c26d-f209-43f4-b662-660346d06641 · inbound

Beyond Greedy Chunking: SLO-Aware Sliding-Window Scheduling for LLM Inference cites this paper.

Beyond Greedy Chunking: SLO-Aware Sliding-Window Scheduling for LLM Inference DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:27:06.039657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T23:49:28.318260Z digest=sha256:2060c6eefdec221d0712f83587820f025e41e85e0feb1cbb0ae97992b060a0b8

Observation e407d765-5959-43a4-b272-097e3e6cc9fa · inbound

Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents cites this paper.

Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:36:59.463922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T01:07:14.691347Z digest=sha256:78e453c99465e4fdec2af1959c0386bbbd8a8db269bad60e4a1880b07cfd62c7

Observation 338310bb-7333-4cb7-9032-c6507b74679f · inbound

ReMP: Low-Downtime Runtime Model-Parallelism Reconfiguration for LLM Serving cites this paper.

ReMP: Low-Downtime Runtime Model-Parallelism Reconfiguration for LLM Serving DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:39:24.932913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T19:32:31.882977Z digest=sha256:4bffece2d5274f74307f15ba51d3172d586370d6eaf20c7ee86ce1efe9f078fe

Observation a6e62147-93fa-408e-88c6-1da34df551c8 · inbound

LiveServe: Interaction-Aware Serving for Real-Time Omni-Modal LLMs cites this paper.

LiveServe: Interaction-Aware Serving for Real-Time Omni-Modal LLMs DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:59:50.585382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T07:26:07.356352Z digest=sha256:ac885675061264d1fd743f2cfa9cdb2c1114393dd609fdff54c9f4e9745b2097

Observation 286d72a8-abb5-474e-8411-d1aa48d0e656 · inbound

SmoothAgent: Efficient Long-Horizon LLM-Based Agent Serving with Lookahead Context Engineering cites this paper.

SmoothAgent: Efficient Long-Horizon LLM-Based Agent Serving with Lookahead Context Engineering DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:14.671196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T17:19:53.560039Z digest=sha256:8445db43f6e8be998f916d6814db710ca3a98817dc8708da30b743fbf8f79b86

Observation 208c00e4-9b51-4ca6-a064-f638bf559b5f · inbound

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving cites this paper.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:ec65faa90827c3383285e0f0b22120d8e66869f9ee3810fac6abc8e01d169649

Observation b03ec011-2f9d-4d74-9362-1893d5dec339 · inbound

BlockServe: Block-Grained Continuous Batching for High-Throughput Diffusion LLM Serving cites this paper.

BlockServe: Block-Grained Continuous Batching for High-Throughput Diffusion LLM Serving DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T05:43:12.690359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:43:12.690359Z digest=sha256:3f70f97c940b064df5972085192ba98c98d92671c225ae0377d9e9fb82fe6446

Observation a0422bac-16ce-46b9-823a-0e3a1899cb82 · inbound

Beyond Storage: State as a Runtime Control Problem in Parallel and Distributed Systems cites this paper.

Beyond Storage: State as a Runtime Control Problem in Parallel and Distributed Systems DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-01T19:51:06.098713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:51:06.098713Z digest=sha256:e775024081c5e1ab54140cf6f5ef6b04db87ecb7b5b5b3e4fc9415fcf551a2be