Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T14:57:57.666022Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 99 of 99 outbound references and 0 inbound Pith citation observations for arXiv:2608.03555.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T14:57:57.666022Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
99 of 99 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8d7dd8e5-e608-48ad-a652-6c400b045177 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Nvidia h200 sxm 141 gb,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d23cb950-4071-4111-9db1-c23f2d3ddbc6 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention TensorRT-LLM,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d994fae3-7ee8-4ec8-b00e-d5c596aeb7bd · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Lmcache: Turboboosting vllm with 7x faster access to 100x more kv caches,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation daf5a3b2-9e11-4e41-8adf-64eae6a3906f · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50801cb7-55e5-45f7-8934-e0966f654ede · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Openai models,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05d2b9d8-8946-40c5-b0f0-fabe509eb385 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Taming throughput-latency tradeoff in llm inference with sarathi-serve,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97c43c0c-55b5-4cba-ac65-1b894d4faa99 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Claude Opus 5,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6def015d-e2ff-4748-916b-0361ec8adddd · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Introducing claude opus 4.6,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d5c634b-397f-411d-b485-69997b07cd7d · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Indexcache: Accelerating sparse attention via cross-layer index reuse,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d889abdc-da70-4da1-b7bb-a0f4fc511191 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention SWE-chat: Coding Agent Interactions From Real Users in the Wild
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e795f554-5ee1-49e2-980a-e83d6624a0de · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9785c5b9-9ee9-42fa-8391-63d5b9581fda · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Magicpig: LSH sampling for efficient LLM generation,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27524ffa-a2f9-4341-8c8b-ed4d34ffd408 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Llmservingsim: A hw/sw co-simulation infrastructure for llm inference serving at scale,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fb6dc67-2ed9-4e65-b40a-bb1bd3beef83 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention 37.3 a 2nm all-digital 14.4gb/s/pin lpddr6 phy with quarter-rate clocking architecture and multi-level fifo-based speculative dfe,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b00c4d4e-a972-48a4-8271-5d7730099d97 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0607bc4-f56d-4fcc-8a80-0e078adbc902 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Deepseek-v3.2: Pushing the frontier of open large language models,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41eab178-fbd1-4a7e-aa98-be1626eb6d5e · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention DeepSeek-V3 Technical Report
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4698871e-9424-4072-8744-99c442d1eaf6 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9c54e13-78da-4267-8386-4ae5bc89ecfe · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention com/EPIC-RPI/STARC, 2025, accessed: 2026-08-01
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4c4fff2-a8d7-44ad-81e9-fc3a4fe0344f · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Starc: Selective token access with remapping and clustering for efficient llm decoding on pim systems,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d98f96b-0faf-44ae-b8f7-66296dc882d2 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Mtia: First generation silicon targeting meta’s recommendation systems,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc31ffd5-61dd-4a79-9e0d-42d1c00e7906 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention kv-cache-tester: Inference server cache performance test- ing suite,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c3fe003-8250-4d43-8360-0e285464e383 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Rdma over ethernet for distributed training at meta scale,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3c33a8c-2718-48b3-b248-4acb9635012f · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Seerattention: Self-distilled attention gating for efficient long-context prefilling,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dccc7dd2-bd06-474d-97f5-bd388c7914fc · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Glm-5: from vibe coding to agentic engineering,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a398715-204c-44e4-bc15-693948fe7d10 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention The Llama 3 Herd of Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb81e072-00eb-482a-a58b-409c75c5b507 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Pim is all you need: A cxl-enabled gpu-free system for large language model inference,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4511efe-4c01-4c32-b566-97888a99794f · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Fastermoe: modeling and optimizing training of large-scale dynamic pre-trained models,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7554c8b6-9059-4353-9e6f-0eecee72c60e · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Shyla: 3d-stacked nvm-dram hybrid llm-inference ar- chitecture exploiting data and memory heterogeneity,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0606b9f1-2d6e-4b29-bccd-e19b59a675d8 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Transparent offloading and mapping (tom): enabling programmer-transparent near-data processing in gpu systems,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92794a87-85ac-4989-aae7-21eccd6726f9 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Hybridspec: Exploiting hybrid-bonding memory to accelerate llm serving through heterogeneous architecture and specu- lative decoding,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65ce5f7e-a2de-4057-a3af-d8051eeb72cb · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention codex_swebenchpro_traces,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ff04b0c3-e21d-481b-bce7-1aeceec7758c · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention A cost-effective near-storage processing solution for offline inference of long-context llms,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4357d5ee-87d6-4671-bd19-378565f75dbc · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention 2023, standard
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 292a8f94-8850-42e8-abe0-a5218dc819a0 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Scalable processing-near-memory for 1m-token llm inference: Cxl-enabled kv-cache management beyond gpu limits,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 27e6b1ca-3cef-41fc-802e-ed0b75920cb8 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Toward standardized near-data processing with unrestricted data placement for gpus,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdb24d7b-fb77-4ad6-bc34-c67397c21fca · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Samsung pim/pnm for transfmer based ai : Energy efficiency on pim/pnm cluster,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3455db69-509a-4875-bd3a-7c51965ffa3f · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention A silicon-proven unified low-latency CXL controller and port-based routing switch for memory-centric fabrics,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 465ec14d-4bfc-402e-8413-0ec92401e7ef · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Efficient memory management for large language model serving with pagedattention,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00f128ec-72ba-4720-aa6a-5f82733ceee4 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention MiniMax Sparse Attention
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6655629-ac27-42ea-94e6-43883d1bb4c8 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Hardware architecture and software stack for pim based on commercial dram technology,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a83c82e6-8200-4657-bce0-f1358b096d23 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Pond: Cxl-based memory pooling systems for cloud platforms,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f38912b-813f-451a-9ab3-85f17b37e38b · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Snapkv: Llm knows what you are looking for before generation,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7f6498db-72a4-45f6-90a3-f75c2a6c6dfb · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Meridian: In-memory acceleration for rag with document attention decomposition,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 71aad5d4-c91f-49c2-858b-41b4ce41751c · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Chime: A case for efficient long-context attention- fc disaggregated inference with dimm-pim,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 869f7343-2506-4a36-b35b-e6b316f3042e · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Nvidia’s ad102 officially revealed, how close were the previ- ous estimates?
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 829d515a-1328-484d-a2a7-e0dfb5104868 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention MoBA: Mixture of block attention for long- context LLMs,
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4b1472e6-8049-4b9b-a3fd-0d908548c702 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Sac: Disaggregated kv cache system for sparse attention llms with cxl,
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f4247f77-a36b-4cbb-bf91-6c1dde323abc · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Tpp: Transparent page placement for cxl-enabled tiered-memory,
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64f75a88-0540-4e21-9f37-6939e5b7260f · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention HBM3E product brief,
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7c1242cb-69e5-425b-9301-cc03127b2e7c · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention g ed., May 2025, production Data Sheet
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7e09f9a6-ee34-4308-aa07-d9b8b4f6cbd2 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Early silicon of raptor: The first 3d-dram accelerator for generative inference,
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9836d34e-ce16-4ebf-89ad-ae9c2ff349d2 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention NVIDIA DGX B200 datasheet,
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1e5d862b-f97e-4795-b713-89da268bb6e0 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention NVIDIA DGX H200 datasheet,
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bc8a10d5-e6e6-443c-b7bc-1192e7f07541 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention SLA-based planner,
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 40a02988-eb34-4758-ae43-b55623d62339 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Multi-process service: Appendix: Tools and in- terface reference,
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 91f4f66e-9c28-4902-9863-85885036bb75 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention (2026) Overall architecture — NVIDIA Dynamo documentation
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 28c3453f-11da-4719-85fc-e20312900b88 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Planner design,
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d2f30c89-eb67-4ec1-9181-69f50c4b9246 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Supported MIG profiles — NVIDIA multi- instance GPU user guide,
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0a47aa5c-f6f7-4039-821b-fa75e4d27b88 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Fine-grained dram: energy-efficient dram for extreme bandwidth systems,
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99c0a576-0a4e-4b7c-ba8f-b01174e431c9 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Exegpt: Constraint-aware resource scheduling for llm inference,
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e245943-ff76-4cb3-a140-40ef552e5bfe · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention ChatGPT (GPT-5),
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ace60453-3340-4802-81d5-92b0b982d7e2 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Attacc! unleashing the power of pim for batched transformer-based generative model inference,
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4be5218b-1826-48f2-908a-3323625201bd · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention A 192-gb 12-high 896-gb/s hbm3 dram with a tsv auto-calibration scheme and machine-learning-based layout opti- mization,
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3bfbb05d-b684-4ff6-8cb7-ac3d9e0cfbc4 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention An lpddr-based cxl-pnm platform for tco- efficient inference of transformer-based large language models,
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1341262d-1f59-4dc4-b3de-789a5a15811a · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Apple m2 die shot and architecture analysis – big cost increase and a15 based ip,
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 38d07b6e-9fa6-4c2c-a51e-0ade696dcac8 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Splitwise: Efficient generative llm inference using phase splitting,
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4eaec88c-7719-437a-afd7-fd967a84a59f · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Mooncake: A kvcache-centric disaggregated architecture for llm serving,
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be2c78d2-20ee-4db1-91b9-59affbe87787 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4717908-7610-4d28-b5b2-78b1b4a751c9 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Longsight: Compute-enabled memory to accelerate large-context llms via sparse attention,
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f44faa9-2424-42e0-b960-19ece375ed0d · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Drex: Accurate and scalable dense retrieval acceleration via algorithmic-hardware codesign,
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25ac2814-db2d-45f2-8542-7c5d76ec5873 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Sparq attention: bandwidth-efficient llm inference,
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ecf9d468-0bbf-4022-b109-7beef9a8a122 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention AttAcc simulator,
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4abada64-ac32-424b-9d45-198da8bbcdc5 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention LLMSimulator,
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9bdce448-1a96-4225-ae22-eda8614c7bd6 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Toolformer: Language models can teach themselves to use tools,
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 78165bad-bf47-4527-a14a-fc563f2e7431 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Ianus: Integrated accelerator based on npu-pim unified memory system,
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 218598e1-5126-43d3-8251-a7b889123e80 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Dynamollm: Designing llm inference clusters for performance and energy efficiency,
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da7ae6ab-79c6-4dbb-a06e-cce59b3194c7 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Quest: query-aware sparsity for efficient long-context llm inference,
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 667ab8f0-9d40-440b-b974-95062c1b1d3e · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Astra-sim2.0: Modeling hierarchical networks and disaggregated systems for large-model training at scale,
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 839c0458-9fd5-4f3c-a139-8460017ad2ea · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention MemExplorer: Navigating the Heterogeneous Memory Design Space for Agentic Inference NPUs
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e503753f-a17d-4fc5-8d8e-ab0748aa2ea3 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Efficient streaming language models with attention sinks,
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 548abcf4-5487-455f-8d76-53761d9825d2 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Hisparse: Turbocharging sparse attention with hierarchical memory,
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 06aa777f-2162-43e5-90a6-415a16325328 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Strata: Hierarchical Context Caching for Long Context Language Model Serving
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61352f64-28ee-4604-9d51-645e6cc9726d · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Beluga: A cxl-based memory architecture for scalable and efficient llm kvcache management,
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5a859f43-3852-4bb0-b0a9-ed6b4028d22d · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention ReAct: Synergizing reasoning and acting in language models,
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 850db07b-93a5-45a3-9604-7db34f945c25 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Tract: Disaggregated llm serving with cxl shared memory kv cache at rack-scale,
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a5ed225-56ef-45e6-8cbb-da43b1be8389 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Available: https://arxiv.org/abs/2601.06288
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1737c7af-5fd7-4fad-b374-3efcad806983 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Duplex: A device for large language models with mixture of experts, grouped query attention, and continuous batching,
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1732f4f-713b-41b3-a128-263a10bb69cd · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention H2o: heavy-hitter oracle for efficient generative inference of large language models,
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 71d4a0a6-e8ec-4b87-a3d9-5714e972d590 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Deepep: an efficient expert-parallel communication library,
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d5be8cc2-997c-4877-83ef-d897fb643fe6 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Patterns behind chaos: Forecasting data movement for efficient large-scale moe llm inference,
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 41ca0024-80fe-42d2-b55e-9edb60e1fb5b · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Sglang: efficient execution of structured language model programs,
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3907fcb6-17cb-48ca-8e76-7207edf8feee · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Distserve: disaggregating prefill and decoding for goodput-optimized large language model serving,
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e239bc84-b2d3-4b0d-ae2d-c14ff1bd4364 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Octopus: Enhancing CXL memory pods via sparse topology,
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0c47f604-1cf0-4652-8f5f-2db614f4fe2f · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention InfLLM-v2: Dense-sparse switchable attention for seamless short-to-long adaptation,
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8e9f74b2-3a44-4952-8547-279579b2989a · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Megascale-infer: Efficient mixture-of-experts model serving with disaggregated expert parallelism,
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f2a2bbea-f4d4-4576-85e9-cffcebd4f8ee · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Available: https://doi.org/10.1145/3579371.3589348
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 278d6a21-d664-440c-900d-9e99ab57a0c5 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 577d7638-0e16-430e-a748-bf2dae12ae21 · outbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention GLM-5: from Vibe Coding to Agentic Engineering
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.