Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 45 inbound Pith citation observations for arXiv:2401.02669.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:39:53.787340Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T09:39:46.796077Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 912586e1-1021-474c-b5a2-23ca6d36c563 · inbound
A Survey on Efficient Inference for Large Language Models Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 276
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8452f46f-dc76-404c-b737-f368d1906b84 · inbound
Pie: Pooling CPU Memory for LLM Inference Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b35285d-3aec-4217-85d3-054aaf04163c · inbound
LLMSteer: Improving Long-Context LLM Inference by Steering Attention on Reused Contexts Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bc970ca-837e-4ab2-907a-dc13a6579332 · inbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ff7a8d3-7e33-4302-a0d9-1ceac9c3347d · inbound
Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f0dbb30-0e8d-4c50-8c02-f067624bfa82 · inbound
A System for Microserving of LLMs Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51337c04-f5fd-47fc-96e1-641a79f20b87 · inbound
LLMQuoter: Enhancing RAG Capabilities Through Efficient Quote Extraction From Large Contexts Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bdd7da9-b375-422f-ae2e-922c8fea5844 · inbound
PICE: A Semantic-Driven Progressive Inference System for LLM Serving in Cloud-Edge Networks Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa98683e-5b9a-43f0-a904-6183b0d6c34c · inbound
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7f5fa46-7422-4d21-b0e6-525f8bb50e48 · inbound
Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2edf95f3-8250-47f2-b3ef-aa5241ad9e9d · inbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bb54491-cfd1-4fae-a01e-7b66d48df701 · inbound
MegaScale-Data: Scaling Dataloader for Multisource Large Foundation Model Training Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 732bb2dd-06db-4c48-af6f-edd9c20bb846 · inbound
SLO-Aware Scheduling for Large Language Model Inferences Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c40a79b-c316-4493-b732-015a2ccfaa86 · inbound
EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0a7575e-9c4d-4e48-afa1-13cff7c187d9 · inbound
Taming the Titans: A Survey of Efficient LLM Inference Serving Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d83a165-af28-4402-96b4-3953b225a919 · inbound
SpecRouter: Adaptive Routing for Multi-Level Speculative Decoding in Large Language Models Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6e991f7-3307-4c1a-9155-e0b38881e6c1 · inbound
CE-LSLM: Efficient Large-Small Language Model Inference and Communication via Cloud-Edge Collaboration Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d76d718f-dd07-4cee-b791-352a709d5d03 · inbound
From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ffa105b3-9b42-401d-9e69-c943be07b538 · inbound
Efficient Serving of LLM Applications with Probabilistic Demand Modeling Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ef143eb-2b8a-420b-8be2-f182a9b87772 · inbound
Breaking the Boundaries of Long-Context LLM Inference: Adaptive KV Management on a Single Commodity GPU Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2acfbd89-7815-46b2-8aaa-613e3183c044 · inbound
Lizard: An Efficient Linearization Framework for Large Language Models Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ed2d8fec-9b06-4bf3-9a06-6b3beb397a16 · inbound
A Distributed Learned Hash Table Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bb1337b-8b54-4029-a2da-6816ad942ff8 · inbound
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad07febf-f548-439d-87dc-15467a593779 · inbound
MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e70b8795-08b9-425e-b467-d9aee185cbda · inbound
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f8f9b874-4417-4e32-b77f-ac415e417089 · inbound
ODMA: On-Demand Memory Allocation Strategy for LLM Serving on LPDDR-Class Accelerators Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 16bb11fb-fa06-4431-8ca8-766fd9eaa14a · inbound
WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation da0b93bd-fe80-43ac-b4df-b32a3ef2febc · inbound
Efficient Remote KV Cache Reuse with GPU-native Video Codec Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2579936f-e5c7-486b-8930-e7604ba86662 · inbound
Towards Efficient Large Vision-Language Models: A Comprehensive Survey on Inference Strategies Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 75e5528a-0793-4a57-b499-daede83601df · inbound
AMMA: A Multi-Chiplet Memory-Centric Architecture for Low-Latency 1M Context Attention Serving Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c04477b0-eb27-4ef1-9abb-3cff65fd2872 · inbound
Unifying Sparse Attention with Hierarchical Memory for Scalable Long-Context LLM Serving Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6bad24d2-0ff9-4487-916c-a8d753e716db · inbound
Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation faf9b1b0-c28a-4933-83fb-b86bdfbc2e54 · inbound
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ff5cb41f-8c9e-4ce2-9dc4-d41dc71083a0 · inbound
NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 20d0172d-b5cc-4067-9b60-c9f16ce82159 · inbound
Move the Query, Not the Cache: Characterizing Cross-Instance Latent Attention Redistribution Across GPU Fabrics Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b602a9ce-8789-45ed-ad4b-bb55206c23d4 · inbound
Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8eac129a-9c81-4126-9857-e480cfe8c2a0 · inbound
Agent-Assisted Side-Channel Attacks on Non-Prefix KV Cache in RAG Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f99d961d-bf9f-4f77-a8b2-35589c7ddf1a · inbound
StickyInvoc: Rethinking Task Models for High-throughput Workflows in the LLM Era Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5178e071-885e-44f4-8fed-b50993fa7318 · inbound
ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0c43c7c5-d02a-4706-991e-1671fce3e692 · inbound
KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 47df7295-932e-4ec0-850d-898b84ec05d2 · inbound
Omni-Flow: A Unified Workflow Orchestration and Distributed KV Cache Sharing Framework for Multimodal Inference Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1beb5909-efba-4ad6-bee9-51a4df9b23de · inbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a4c2f96-f166-4bb8-a67b-e77b4b5cb1a8 · inbound
Auto-Scaling Heterogeneous Neural Processing Units for Energy and Cost-Efficient LLM Serving Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e28df95a-227a-46d1-95a5-62f39a064cc7 · inbound
C$^2$KV: Compressed and Composable KV Cache Reuse for Efficient LLM Inference Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91a2a6e7-f1be-498c-b55a-0a1f766c954f · inbound
Topology-Aware Data Movement for Disaggregated GPU Inference Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.