Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 34 inbound Pith citation observations for arXiv:2401.08671.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T10:10:22.872509Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T11:59:50.582581Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 9ea946e8-56ae-4cfe-9859-67f8ea75d258 · inbound
A Survey on Efficient Inference for Large Language Models DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 279
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 852d3521-7c90-462c-832a-bcc96f884e8b · inbound
HybridFlow: A Flexible and Efficient RLHF Framework DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a129fb99-037a-4081-a384-96d4579e7a79 · inbound
BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0c8721d8-0d08-4ff2-ad35-51587c21c02f · inbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6bedb766-c3d8-4009-a6f4-8cabf5d05c7c · inbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebbd282c-edd5-43e1-ae96-9cf2bf0587d4 · inbound
MIST: A Co-Design Framework for Heterogeneous, Multi-Stage LLM Inference DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b06f25c7-381a-4421-a4e4-1855b33a2e6f · inbound
ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 313443c8-0e08-4e59-b929-60f45ebddbc7 · inbound
AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17c4b2d3-7d80-4f7e-a321-478c29739989 · inbound
Rectified Sparse Attention DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f75106c-2cda-4921-a977-c6699722549e · inbound
Chat-Ghosting: A Comparative Study of Methods for Auto-Completion in Dialog Systems DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f846ec65-36d7-4250-8207-3cbfadd1c795 · inbound
Rethinking Caching for LLM Serving Systems: Beyond Traditional Heuristics DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 804a2c0c-1dc3-47dc-a9f9-82497e7d3e21 · inbound
HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c81406e8-a4fc-4464-adde-fb4e1b76914f · inbound
WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 268c5e27-d697-4056-a7a9-becc9c17b726 · inbound
PrefixWall: Mitigating Prefix Caching Side Channels in Shared LLM Systems DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ffa085a0-9695-41a7-8065-f41d1ffbd4c5 · inbound
FASTER: Rethinking Real-Time Flow VLAs DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation aaf05164-07a0-4f3b-a171-ee9197516154 · inbound
FASTER: Rethinking Real-Time Flow VLAs DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9c00d7a0-282f-44cf-acd1-f1192872aa99 · inbound
Towards Efficient Large Vision-Language Models: A Comprehensive Survey on Inference Strategies DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 876815a3-5aa7-486c-8378-6ce3c7e3a195 · inbound
Characterizing Performance-Energy Trade-offs of Large Language Models in Multi-Request Workflows DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d57681e3-f3fb-42fc-8982-ad26d7846e04 · inbound
Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 61b831f4-fcd0-47ff-adc3-18b849d53a70 · inbound
Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c0ab0b09-ffc3-4798-9256-856c3064c131 · inbound
MARS: Efficient, Adaptive Co-Scheduling for Heterogeneous Agentic Systems DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bdb1d749-6c65-4651-bbda-4dec91c8d58d · inbound
Federation of Experts: Communication Efficient Distributed Inference for Large Language Models DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4baef794-3b5c-4841-b783-3006804077fb · inbound
Focused Forcing: Content-Aware Per-Frame KV Selection for Efficient Autoregressive Video Diffusion DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 13fd3d8f-a36d-4899-832c-359a99beb400 · inbound
AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 83d9298f-7b54-40a6-9b32-f483749567fe · inbound
Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6e215635-2261-408f-9102-c420675d9305 · inbound
Multi-Segment Attention: Enabling Efficient KV-Cache Management for Faster Large Language Model Serving DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f690c26d-f209-43f4-b662-660346d06641 · inbound
Beyond Greedy Chunking: SLO-Aware Sliding-Window Scheduling for LLM Inference DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e407d765-5959-43a4-b272-097e3e6cc9fa · inbound
Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 338310bb-7333-4cb7-9032-c6507b74679f · inbound
ReMP: Low-Downtime Runtime Model-Parallelism Reconfiguration for LLM Serving DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a6e62147-93fa-408e-88c6-1da34df551c8 · inbound
LiveServe: Interaction-Aware Serving for Real-Time Omni-Modal LLMs DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 286d72a8-abb5-474e-8411-d1aa48d0e656 · inbound
SmoothAgent: Efficient Long-Horizon LLM-Based Agent Serving with Lookahead Context Engineering DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 208c00e4-9b51-4ca6-a064-f638bf559b5f · inbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b03ec011-2f9d-4d74-9362-1893d5dec339 · inbound
BlockServe: Block-Grained Continuous Batching for High-Throughput Diffusion LLM Serving DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0422bac-16ce-46b9-823a-0e3a1899cb82 · inbound
Beyond Storage: State as a Runtime Control Problem in Parallel and Distributed Systems DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.