Pith. sign in

Paper Citation Record · LEDGER

Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 45 inbound Pith citation observations for arXiv:2401.02669.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.02669 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 45 of 45 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:39:53.787340Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:39:46.796077Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 912586e1-1021-474c-b5a2-23ca6d36c563 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 276

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:39:33.479689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:695d031f2ca6eb706775cc47fb0987d0b01a38d442b27fcb96419cf6d473cf0d

Observation 8452f46f-dc76-404c-b737-f368d1906b84 · inbound

Pie: Pooling CPU Memory for LLM Inference cites this paper.

Pie: Pooling CPU Memory for LLM Inference Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T20:53:26.655414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:53:26.655414Z digest=sha256:260ef2175024568406605212b93f1a08fb2ba1e3401a6177491efe3494f72875

Observation 1b35285d-3aec-4217-85d3-054aaf04163c · inbound

LLMSteer: Improving Long-Context LLM Inference by Steering Attention on Reused Contexts cites this paper.

LLMSteer: Improving Long-Context LLM Inference by Steering Attention on Reused Contexts Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-12T17:00:04.076364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:00:04.076364Z digest=sha256:60ff37fad929fa77272afe9d6fdecfc4f7d572095259d87579f0ab986152c976

Observation 4bc970ca-837e-4ab2-907a-dc13a6579332 · inbound

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference cites this paper.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T20:18:17.762149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:18:17.762149Z digest=sha256:44a1b95e817470d8cfb02a0b968b4692f9993bbd9c6742ece023981f26691ef3

Observation 3ff7a8d3-7e33-4302-a0d9-1ceac9c3347d · inbound

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries cites this paper.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.956930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.956930Z digest=sha256:bb4a4094b84b5da3abf3405082a694d472b7147b7a54262d9053dffd2908f5b4

Observation 1f0dbb30-0e8d-4c50-8c02-f067624bfa82 · inbound

A System for Microserving of LLMs cites this paper.

A System for Microserving of LLMs Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T14:05:01.827674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:05:01.827674Z digest=sha256:10f136a4e8109f3040ba4f6745f48f9a50967f3f8b468f65f1ff00dba9eac42f

Observation 51337c04-f5fd-47fc-96e1-641a79f20b87 · inbound

LLMQuoter: Enhancing RAG Capabilities Through Efficient Quote Extraction From Large Contexts cites this paper.

LLMQuoter: Enhancing RAG Capabilities Through Efficient Quote Extraction From Large Contexts Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T21:16:23.940987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:16:23.940987Z digest=sha256:f42d6ebfb9b47004a0bb0e8b14ed080fdecf25b7809617afb05f5bcb03f6aa3d

Observation 7bdd7da9-b375-422f-ae2e-922c8fea5844 · inbound

PICE: A Semantic-Driven Progressive Inference System for LLM Serving in Cloud-Edge Networks cites this paper.

PICE: A Semantic-Driven Progressive Inference System for LLM Serving in Cloud-Edge Networks Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T20:13:42.886796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:13:42.886796Z digest=sha256:12efc957878dac4db4ccd958db1303faf98bb8d93710ce0f3856b37483600740

Observation fa98683e-5b9a-43f0-a904-6183b0d6c34c · inbound

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference cites this paper.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.293818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.293818Z digest=sha256:aeaeb139ce6c9db89fcbebccc29a146cdb88037985847e886df4df05957e8db6

Observation e7f5fa46-7422-4d21-b0e6-525f8bb50e48 · inbound

Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents cites this paper.

Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T14:13:38.164388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:13:38.164388Z digest=sha256:e1de16204941e8e6687812b082452f62b17973c9c9263dfdc31d001eefde73f6

Observation 2edf95f3-8250-47f2-b3ef-aa5241ad9e9d · inbound

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees cites this paper.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.973785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.973785Z digest=sha256:49659fba367c8356ab167554281938cb7e4ce4035315e746b1ba0ce301cab5af

Observation 5bb54491-cfd1-4fae-a01e-7b66d48df701 · inbound

MegaScale-Data: Scaling Dataloader for Multisource Large Foundation Model Training cites this paper.

MegaScale-Data: Scaling Dataloader for Multisource Large Foundation Model Training Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:15:09.596423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T21:12:22.201810Z digest=sha256:402b9547d00eb968ec311afdb797d76529746c61337f9fec27d704301ef6dd0a

Observation 732bb2dd-06db-4c48-af6f-edd9c20bb846 · inbound

SLO-Aware Scheduling for Large Language Model Inferences cites this paper.

SLO-Aware Scheduling for Large Language Model Inferences Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T11:39:53.787340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:39:53.787340Z digest=sha256:50b9566fb87268ed0e4fcb2b0c3e68969da08de80bc2096a27770bb4241a1e02

Observation 0c40a79b-c316-4493-b732-015a2ccfaa86 · inbound

EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration cites this paper.

EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:52.452135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:52.452135Z digest=sha256:7c7dee746b57482d8f3d3aa9e647f616d309c2a05f317d39ffbf7fa9ff7c2269

Observation d0a7575e-9c4d-4e48-afa1-13cff7c187d9 · inbound

Taming the Titans: A Survey of Efficient LLM Inference Serving cites this paper.

Taming the Titans: A Survey of Efficient LLM Inference Serving Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:09.959910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:49:09.959910Z digest=sha256:e410b21c1b0f6982238770c79f0635e53090ed539d4efa84e9eac3e0d0d809df

Observation 3d83a165-af28-4402-96b4-3953b225a919 · inbound

SpecRouter: Adaptive Routing for Multi-Level Speculative Decoding in Large Language Models cites this paper.

SpecRouter: Adaptive Routing for Multi-Level Speculative Decoding in Large Language Models Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T22:15:13.711912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:15:13.711912Z digest=sha256:05540a6b84affa9ea7d5624043c3c934767825dc1b018b7b8bca6107bd31b0e9

Observation c6e991f7-3307-4c1a-9155-e0b38881e6c1 · inbound

CE-LSLM: Efficient Large-Small Language Model Inference and Communication via Cloud-Edge Collaboration cites this paper.

CE-LSLM: Efficient Large-Small Language Model Inference and Communication via Cloud-Edge Collaboration Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:21.786114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:21.786114Z digest=sha256:77a0055f054ee30a9122fb1d9ececab4005dd9ab4587c06d691685bee7e5966f

Observation d76d718f-dd07-4cee-b791-352a709d5d03 · inbound

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems cites this paper.

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:52:16.207614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T11:49:36.574471Z digest=sha256:5fbefe91e6044659421e78e05f220ae61ef866489ff1fc84f543d124fcc358f8

Observation ffa105b3-9b42-401d-9e69-c943be07b538 · inbound

Efficient Serving of LLM Applications with Probabilistic Demand Modeling cites this paper.

Efficient Serving of LLM Applications with Probabilistic Demand Modeling Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:46.533712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:46.533712Z digest=sha256:651994ced7292b2387b74267392e713e61e818d57fd2ef4ef049cf9450b3f2bc

Observation 3ef143eb-2b8a-420b-8be2-f182a9b87772 · inbound

Breaking the Boundaries of Long-Context LLM Inference: Adaptive KV Management on a Single Commodity GPU cites this paper.

Breaking the Boundaries of Long-Context LLM Inference: Adaptive KV Management on a Single Commodity GPU Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T23:00:11.123886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:00:11.123886Z digest=sha256:f3a522cf533d2360db67040b4f43549ac9044bbbbb4b934494d8d3c02a401d82

Observation 2acfbd89-7815-46b2-8aaa-613e3183c044 · inbound

Lizard: An Efficient Linearization Framework for Large Language Models cites this paper.

Lizard: An Efficient Linearization Framework for Large Language Models Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:42:04.568701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T04:37:55.034479Z digest=sha256:d5f90513e61a052d1f70454407aac504926a3953b2b95906672546b92cc14318

Observation ed2d8fec-9b06-4bf3-9a06-6b3beb397a16 · inbound

A Distributed Learned Hash Table cites this paper.

A Distributed Learned Hash Table Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T18:46:59.973133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:46:59.973133Z digest=sha256:6988c19920a91c68c5b103e1f6d1ce947bb180d73992967fc44729f8abe982b0

Observation 9bb1337b-8b54-4029-a2da-6816ad942ff8 · inbound

Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism cites this paper.

Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T20:53:32.388187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:53:32.388187Z digest=sha256:414ae5bad51a8b2c414d52349eeddd31d166dc5fa0a3e1a6586e80b000f00b46

Observation ad07febf-f548-439d-87dc-15467a593779 · inbound

MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness cites this paper.

MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T18:00:22.906958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:00:22.906958Z digest=sha256:d6e7ab6f5e4544afa1ad5d803f0130f4410fd2be886530252ba68d8eeb83f24b

Observation e70b8795-08b9-425e-b467-d9aee185cbda · inbound

Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services cites this paper.

Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:01:31.480503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T14:59:38.194894Z digest=sha256:91874035dce9e453015173501de12fcf3a423e8e8120561b848bf3ac4dbe5c6a

Observation f8f9b874-4417-4e32-b77f-ac415e417089 · inbound

ODMA: On-Demand Memory Allocation Strategy for LLM Serving on LPDDR-Class Accelerators cites this paper.

ODMA: On-Demand Memory Allocation Strategy for LLM Serving on LPDDR-Class Accelerators Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:58:43.002547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T23:54:08.052482Z digest=sha256:5eeece9e8238209827e6e97b193bb674b0d4aa9e2dc3db3b5593ce152bb67ea8

Observation 16bb11fb-fa06-4431-8ca8-766fd9eaa14a · inbound

WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving cites this paper.

WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:24:51.260073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T12:22:56.972437Z digest=sha256:09940322217675b82fdcded0d025745dcd246fc0a4b72161f86c6d481df15d67

Observation da0b93bd-fe80-43ac-b4df-b32a3ef2febc · inbound

Efficient Remote KV Cache Reuse with GPU-native Video Codec cites this paper.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:22:22.746402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:67c31da8cc393f6cc026e727d5151714956e48597ae1c375f9d1f3b354709042

Observation 2579936f-e5c7-486b-8930-e7604ba86662 · inbound

Towards Efficient Large Vision-Language Models: A Comprehensive Survey on Inference Strategies cites this paper.

Towards Efficient Large Vision-Language Models: A Comprehensive Survey on Inference Strategies Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:43:00.683591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-14T21:40:55.343169Z digest=sha256:4fc4d444d8a06d60be0947a1276432a4b5f5d7703f451c96abfc949b36660702

Observation 75e5528a-0793-4a57-b499-daede83601df · inbound

AMMA: A Multi-Chiplet Memory-Centric Architecture for Low-Latency 1M Context Attention Serving cites this paper.

AMMA: A Multi-Chiplet Memory-Centric Architecture for Low-Latency 1M Context Attention Serving Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:46:13.115621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-07T14:16:46.241126Z digest=sha256:e9ca06edd74c8664b6f09afeffaaf9d66e7a3e4bef955207b815b9cfc5adbb80

Observation c04477b0-eb27-4ef1-9abb-3cff65fd2872 · inbound

Unifying Sparse Attention with Hierarchical Memory for Scalable Long-Context LLM Serving cites this paper.

Unifying Sparse Attention with Hierarchical Memory for Scalable Long-Context LLM Serving Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:01:25.951780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-07T13:15:21.201950Z digest=sha256:dfc13b8e891a022b0bd25d50c6274394e542379b33721de9e952c8d0db48dd66

Observation 6bad24d2-0ff9-4487-916c-a8d753e716db · inbound

Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference cites this paper.

Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:14:10.664745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T05:01:02.726796Z digest=sha256:c9181d99b507713eae2e411df24a6ba055d01b053ccae87b3d8cfe36d22b4a0b

Observation faf9b1b0-c28a-4933-83fb-b86bdfbc2e54 · inbound

MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems cites this paper.

MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:46:21.116269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T04:29:57.548702Z digest=sha256:2b5b0b223908267e4df971d6fafb8dab5278c70b2f27a353dec67f6d345c6cb0

Observation ff5cb41f-8c9e-4ce2-9dc4-d41dc71083a0 · inbound

NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding cites this paper.

NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:53:54.623374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T01:51:38.671365Z digest=sha256:ceb416bd1b0e0cd2e1e9d8659e9133f881c720d2efa7c41c61e91e79cae4caec

Observation 20d0172d-b5cc-4067-9b60-c9f16ce82159 · inbound

Move the Query, Not the Cache: Characterizing Cross-Instance Latent Attention Redistribution Across GPU Fabrics cites this paper.

Move the Query, Not the Cache: Characterizing Cross-Instance Latent Attention Redistribution Across GPU Fabrics Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:56:15.745685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T16:02:53.294235Z digest=sha256:de28d43b9c837b8d333ea36a985908cffdad87b7d506af82ef12ba4cf8d09eda

Observation b602a9ce-8789-45ed-ad4b-bb55206c23d4 · inbound

Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents cites this paper.

Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:36:59.487421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T01:07:14.691347Z digest=sha256:178674aa7a3d584a44d2585d3554a63b8e4a7eea25c987b4ef32069083cd950b

Observation 8eac129a-9c81-4126-9857-e480cfe8c2a0 · inbound

Agent-Assisted Side-Channel Attacks on Non-Prefix KV Cache in RAG cites this paper.

Agent-Assisted Side-Channel Attacks on Non-Prefix KV Cache in RAG Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.406576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T12:19:30.666750Z digest=sha256:47684530fda1fd017e09ad3f237ff307ab8e51661fa8086a318e7e39bf26a03f

Observation f99d961d-bf9f-4f77-a8b2-35589c7ddf1a · inbound

StickyInvoc: Rethinking Task Models for High-throughput Workflows in the LLM Era cites this paper.

StickyInvoc: Rethinking Task Models for High-throughput Workflows in the LLM Era Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:39:42.046455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T11:16:53.967397Z digest=sha256:ab599e3d252789070c4de44ef06ff03f15b828f05f374b588132fb818077ec9a

Observation 5178e071-885e-44f4-8fed-b50993fa7318 · inbound

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill cites this paper.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:39:46.797595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:8a5082a0eec2b262805a3a623d02ae232ecfaad6183c1befb00f759cf182248a

Observation 0c43c7c5-d02a-4706-991e-1671fce3e692 · inbound

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding cites this paper.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-30T03:04:14.208841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:230928d3b00746ba088e5601d6a2acd032f01193e8d58a6627152a35a3935161

Observation 47df7295-932e-4ec0-850d-898b84ec05d2 · inbound

Omni-Flow: A Unified Workflow Orchestration and Distributed KV Cache Sharing Framework for Multimodal Inference cites this paper.

Omni-Flow: A Unified Workflow Orchestration and Distributed KV Cache Sharing Framework for Multimodal Inference Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:35:43.654306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-01T04:13:53.722933Z digest=sha256:b7704fd2ba322c6f1c60fb9d39aa94c6383da543e9edfddacc9b368288fa44e2

Observation 1beb5909-efba-4ad6-bee9-51a4df9b23de · inbound

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving cites this paper.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:9e6d9e635fd341d16ba40165c9857115252afc129b69dccd23c2a50730580543

Observation 6a4c2f96-f166-4bb8-a67b-e77b4b5cb1a8 · inbound

Auto-Scaling Heterogeneous Neural Processing Units for Energy and Cost-Efficient LLM Serving cites this paper.

Auto-Scaling Heterogeneous Neural Processing Units for Energy and Cost-Efficient LLM Serving Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T20:54:14.172249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T20:54:14.172249Z digest=sha256:1e67afa508cd96608da87f6c24e55718d633d72f12de0f5a49dfba1d94400e18

Observation e28df95a-227a-46d1-95a5-62f39a064cc7 · inbound

C$^2$KV: Compressed and Composable KV Cache Reuse for Efficient LLM Inference cites this paper.

C$^2$KV: Compressed and Composable KV Cache Reuse for Efficient LLM Inference Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T17:16:38.402742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:16:38.402742Z digest=sha256:c9c65765d9b1b0016287acd1b8b2df5e4905c60dcec1718889b04c87e74df20f

Observation 91a2a6e7-f1be-498c-b55a-0a1f766c954f · inbound

Topology-Aware Data Movement for Disaggregated GPU Inference cites this paper.

Topology-Aware Data Movement for Disaggregated GPU Inference Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:34.973260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:34.973260Z digest=sha256:55943523d48ea0fc5ff1d63d1cf74552305ec4b9ff70dc774447131dcf362ddc