Pith. sign in

Paper Citation Record · LEDGER

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention

As of 22 August 2026, this Paper Citation Record lists 99 of 99 outbound references and 0 inbound Pith citation observations for arXiv:2608.03555.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03555 v1

Coverage vector

measured 99 of 99 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:57:57.666022Z

measured 99 of 99 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

99 of 99 outbound references displayed

  • verified exact2
  • verified fuzzy40
  • unresolved56
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8d7dd8e5-e608-48ad-a652-6c400b045177 · outbound

This paper cites Nvidia h200 sxm 141 gb,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Nvidia h200 sxm 141 gb,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.149323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.149323Z digest=sha256:e51779fd132d78f349e38f09cd82668a5a39bbdddc4393ba2453274c3a946d94

Observation d23cb950-4071-4111-9db1-c23f2d3ddbc6 · outbound

This paper cites TensorRT-LLM,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention TensorRT-LLM,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.153945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.153945Z digest=sha256:46846578638e145dfd7d12440312ab610c9fda3982b44b3610fce03d4eebc9d1

Observation d994fae3-7ee8-4ec8-b00e-d5c596aeb7bd · outbound

This paper cites Lmcache: Turboboosting vllm with 7x faster access to 100x more kv caches,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Lmcache: Turboboosting vllm with 7x faster access to 100x more kv caches,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.160343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.160343Z digest=sha256:fd53f05e5eae521d36ecca9a1901e4b51e53985d5e54512ccba14f2608c68069

Observation daf5a3b2-9e11-4e41-8adf-64eae6a3906f · outbound

This paper cites Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.165570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.165570Z digest=sha256:f8ede5e75e9770ac6ad146ec703616444be6edde301b8cb363ffab8b3aa52f78

Observation 50801cb7-55e5-45f7-8934-e0966f654ede · outbound

This paper cites Openai models,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Openai models,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.170342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.170342Z digest=sha256:f8cc8e4356a850f9fc0042344150a4ad9419afc5655d2011d0aa278b9fe53e32

Observation 05d2b9d8-8946-40c5-b0f0-fabe509eb385 · outbound

This paper cites Taming throughput-latency tradeoff in llm inference with sarathi-serve,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Taming throughput-latency tradeoff in llm inference with sarathi-serve,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.175839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.175839Z digest=sha256:399130ce31531ac565ff3cebc3d4849262027fa57bfd81533650618397aff8ea

Observation 97c43c0c-55b5-4cba-ac65-1b894d4faa99 · outbound

This paper cites Claude Opus 5,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Claude Opus 5,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.180450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.180450Z digest=sha256:2a4480a4be71b18339e89266d56076f150b0873f29217c609e74630faf65f0ce

Observation 6def015d-e2ff-4748-916b-0361ec8adddd · outbound

This paper cites Introducing claude opus 4.6,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Introducing claude opus 4.6,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.185398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.185398Z digest=sha256:4d4c906bccb0aa64d356ee4650682d8abf9f9f1088392bbe7a86f16457de9147

Observation 4d5c634b-397f-411d-b485-69997b07cd7d · outbound

This paper cites Indexcache: Accelerating sparse attention via cross-layer index reuse,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Indexcache: Accelerating sparse attention via cross-layer index reuse,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.190238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.190238Z digest=sha256:0a32369e6c9014fcb96e598d6f7ac2fb82b9ba0367d969d9bf02c18869577e7f

Observation d889abdc-da70-4da1-b7bb-a0f4fc511191 · outbound

This paper cites SWE-chat: Coding Agent Interactions From Real Users in the Wild.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention SWE-chat: Coding Agent Interactions From Real Users in the Wild

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.194675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.194675Z digest=sha256:8ae01150339a1d483c6cde2316fdde02944ea1feec7b875421fa88084bcca809

Observation e795f554-5ee1-49e2-980a-e83d6624a0de · outbound

This paper cites RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.200352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.200352Z digest=sha256:554e41c2f4b96fed208df8fba484bf6ad04521a9c97e091dff5358dcf3976d3d

Observation 9785c5b9-9ee9-42fa-8391-63d5b9581fda · outbound

This paper cites Magicpig: LSH sampling for efficient LLM generation,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Magicpig: LSH sampling for efficient LLM generation,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.205506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.205506Z digest=sha256:52829aeea8c0570ed5917bd62b2d94e8bf34f15bf75ed9c389ebe813a7d65d50

Observation 27524ffa-a2f9-4341-8c8b-ed4d34ffd408 · outbound

This paper cites Llmservingsim: A hw/sw co-simulation infrastructure for llm inference serving at scale,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Llmservingsim: A hw/sw co-simulation infrastructure for llm inference serving at scale,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.210220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.210220Z digest=sha256:180d24ab96320afacb4cd76942e48c5dcfda51669aa6bb23c061499b47973a4e

Observation 7fb6dc67-2ed9-4e65-b40a-bb1bd3beef83 · outbound

This paper cites 37.3 a 2nm all-digital 14.4gb/s/pin lpddr6 phy with quarter-rate clocking architecture and multi-level fifo-based speculative dfe,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention 37.3 a 2nm all-digital 14.4gb/s/pin lpddr6 phy with quarter-rate clocking architecture and multi-level fifo-based speculative dfe,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.214380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.214380Z digest=sha256:e1255e4395b90021d674da6f159dc65c1207908433b43528375e714366bcb0ee

Observation b00c4d4e-a972-48a4-8271-5d7730099d97 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.218841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.218841Z digest=sha256:4db90e1b3891f72f3e93edcdb89686f83a686cc131936d40585d346e85c3f441

Observation c0607bc4-f56d-4fcc-8a80-0e078adbc902 · outbound

This paper cites Deepseek-v3.2: Pushing the frontier of open large language models,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Deepseek-v3.2: Pushing the frontier of open large language models,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.227132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.227132Z digest=sha256:33a8f86b87c35d76fb3a98310f4369b96f4adb7de26e605576bdcfe09c9dbfda

Observation 41eab178-fbd1-4a7e-aa98-be1626eb6d5e · outbound

This paper cites DeepSeek-V3 Technical Report.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention DeepSeek-V3 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.232105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.232105Z digest=sha256:402c6bf1c90968fe9baf56d0f58fce2e966778887c9c84f604f5313460b5331b

Observation 4698871e-9424-4072-8744-99c442d1eaf6 · outbound

This paper cites SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.236994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.236994Z digest=sha256:5cd133c237db77207502cdc44a505f5c265f454bca5db05822f212a642cb7091

Observation b9c54e13-78da-4267-8386-4ae5bc89ecfe · outbound

This paper cites com/EPIC-RPI/STARC, 2025, accessed: 2026-08-01.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention com/EPIC-RPI/STARC, 2025, accessed: 2026-08-01

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.242889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.242889Z digest=sha256:cc77d75cc76a70a6c190ada6c733234f14f25f330a8bd70eabfc3991d6c18891

Observation f4c4fff2-a8d7-44ad-81e9-fc3a4fe0344f · outbound

This paper cites Starc: Selective token access with remapping and clustering for efficient llm decoding on pim systems,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Starc: Selective token access with remapping and clustering for efficient llm decoding on pim systems,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.247307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.247307Z digest=sha256:19a7a3a6b5b04573719f2ddd0781f555623097c700144f1ac4cb4ad0d35fefd7

Observation 5d98f96b-0faf-44ae-b8f7-66296dc882d2 · outbound

This paper cites Mtia: First generation silicon targeting meta’s recommendation systems,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Mtia: First generation silicon targeting meta’s recommendation systems,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.251826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.251826Z digest=sha256:802cb8bb9cddfad01d397a20fae7e80a84887e4e91683eb79c59ad50bec8b3df

Observation bc31ffd5-61dd-4a79-9e0d-42d1c00e7906 · outbound

This paper cites kv-cache-tester: Inference server cache performance test- ing suite,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention kv-cache-tester: Inference server cache performance test- ing suite,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.261315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.261315Z digest=sha256:3bd09288070bcf4f8ede7624c16d0b27e50150bacbe81b6d4a9ef64977590b6c

Observation 0c3fe003-8250-4d43-8360-0e285464e383 · outbound

This paper cites Rdma over ethernet for distributed training at meta scale,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Rdma over ethernet for distributed training at meta scale,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.266049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.266049Z digest=sha256:4ff8bf6472519e50bbd367fa97589da142b7616df630b5e931c5788684892416

Observation b3c33a8c-2718-48b3-b248-4acb9635012f · outbound

This paper cites Seerattention: Self-distilled attention gating for efficient long-context prefilling,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Seerattention: Self-distilled attention gating for efficient long-context prefilling,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.272118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.272118Z digest=sha256:93707d325ec0c2e4ee907b8939042b9064c669b843d203bcfbd3394be90e2db3

Observation dccc7dd2-bd06-474d-97f5-bd388c7914fc · outbound

This paper cites Glm-5: from vibe coding to agentic engineering,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Glm-5: from vibe coding to agentic engineering,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.277907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.277907Z digest=sha256:10480933ccb8933407f77209de1b95f95e59f118d9816a3aec9f3fbf9ca0fec9

Observation 5a398715-204c-44e4-bc15-693948fe7d10 · outbound

This paper cites The Llama 3 Herd of Models.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention The Llama 3 Herd of Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.290966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.290966Z digest=sha256:8c48eb3d6b70813fa88267f913a4c2a0a4b5b51f5f59375e4252739ec0780839

Observation eb81e072-00eb-482a-a58b-409c75c5b507 · outbound

This paper cites Pim is all you need: A cxl-enabled gpu-free system for large language model inference,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Pim is all you need: A cxl-enabled gpu-free system for large language model inference,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.297214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.297214Z digest=sha256:40ad7f3a5a648ec31e1a00196c1fe178173ad21a1f6411ea8510e22f1bcd2eaf

Observation e4511efe-4c01-4c32-b566-97888a99794f · outbound

This paper cites Fastermoe: modeling and optimizing training of large-scale dynamic pre-trained models,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Fastermoe: modeling and optimizing training of large-scale dynamic pre-trained models,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.307638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.307638Z digest=sha256:5ee1a85c71671364b31991d9cfbe95763c00fd5406afba5f5dc1f9060936182e

Observation 7554c8b6-9059-4353-9e6f-0eecee72c60e · outbound

This paper cites Shyla: 3d-stacked nvm-dram hybrid llm-inference ar- chitecture exploiting data and memory heterogeneity,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Shyla: 3d-stacked nvm-dram hybrid llm-inference ar- chitecture exploiting data and memory heterogeneity,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.312140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.312140Z digest=sha256:bd577d6f139ab024271d1e0b123d28a7162b995adabf401b1df76442ae5d2b2f

Observation 0606b9f1-2d6e-4b29-bccd-e19b59a675d8 · outbound

This paper cites Transparent offloading and mapping (tom): enabling programmer-transparent near-data processing in gpu systems,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Transparent offloading and mapping (tom): enabling programmer-transparent near-data processing in gpu systems,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.322232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.322232Z digest=sha256:bb98f37a3adc6c05e59209dd42bb2c1612b458674395c6ffb8565e4e4765c8fe

Observation 92794a87-85ac-4989-aae7-21eccd6726f9 · outbound

This paper cites Hybridspec: Exploiting hybrid-bonding memory to accelerate llm serving through heterogeneous architecture and specu- lative decoding,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Hybridspec: Exploiting hybrid-bonding memory to accelerate llm serving through heterogeneous architecture and specu- lative decoding,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.327773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.327773Z digest=sha256:b64b6759a99f3cb4a8c908539b871f0edea2bbeddf8d37f128b44b8fbfe842aa

Observation 65ce5f7e-a2de-4057-a3af-d8051eeb72cb · outbound

This paper cites codex_swebenchpro_traces,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention codex_swebenchpro_traces,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:00.405442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.332368Z digest=sha256:ed4b53dd3954fd81f6ee2e28e9102243a53e4e28fa231d74110c39bf7a75a0cd

Observation ff04b0c3-e21d-481b-bce7-1aeceec7758c · outbound

This paper cites A cost-effective near-storage processing solution for offline inference of long-context llms,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention A cost-effective near-storage processing solution for offline inference of long-context llms,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.337036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.337036Z digest=sha256:b094f3a04dacddd4667af7c7b934742aa397d4cff1a8997dec5e4066693f9b3b

Observation 4357d5ee-87d6-4671-bd19-378565f75dbc · outbound

This paper cites 2023, standard.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention 2023, standard

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:00.388832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.342985Z digest=sha256:a2c9678d87a486078f60858d705e5b7bba2bed5615d00d27b30f5999723351a2

Observation 292a8f94-8850-42e8-abe0-a5218dc819a0 · outbound

This paper cites Scalable processing-near-memory for 1m-token llm inference: Cxl-enabled kv-cache management beyond gpu limits,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Scalable processing-near-memory for 1m-token llm inference: Cxl-enabled kv-cache management beyond gpu limits,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:00.373609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.347564Z digest=sha256:c2299edefaee7aac09652be9ac8dc41010d4135839002614625c3763ed8cd317

Observation 27e6b1ca-3cef-41fc-802e-ed0b75920cb8 · outbound

This paper cites Toward standardized near-data processing with unrestricted data placement for gpus,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Toward standardized near-data processing with unrestricted data placement for gpus,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.351827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.351827Z digest=sha256:f38df7d072312434783cf64fba571b353303f0543c0aa55a93da1dbcd77a6539

Observation bdb24d7b-fb77-4ad6-bc34-c67397c21fca · outbound

This paper cites Samsung pim/pnm for transfmer based ai : Energy efficiency on pim/pnm cluster,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Samsung pim/pnm for transfmer based ai : Energy efficiency on pim/pnm cluster,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:00.357756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.361114Z digest=sha256:57fcf12d34bfdc7905618f5edbeeb5ab029b14c16c748bed24f606c00e718baf

Observation 3455db69-509a-4875-bd3a-7c51965ffa3f · outbound

This paper cites A silicon-proven unified low-latency CXL controller and port-based routing switch for memory-centric fabrics,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention A silicon-proven unified low-latency CXL controller and port-based routing switch for memory-centric fabrics,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:00.342205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.365715Z digest=sha256:233f57c703f559efde13c8596f2b4aac239ccf4ef52d980be86b9004c60978a8

Observation 465ec14d-4bfc-402e-8413-0ec92401e7ef · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Efficient memory management for large language model serving with pagedattention,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.370496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.370496Z digest=sha256:5115c41d15edb09ce50a065545a0c913d999f6f775656571f1560b5f14806bb5

Observation 00f128ec-72ba-4720-aa6a-5f82733ceee4 · outbound

This paper cites MiniMax Sparse Attention.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention MiniMax Sparse Attention

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.375091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.375091Z digest=sha256:f465f5a20465880eda958820a7ca0cc6aae26f3acf69c0b5f43150434960c754

Observation b6655629-ac27-42ea-94e6-43883d1bb4c8 · outbound

This paper cites Hardware architecture and software stack for pim based on commercial dram technology,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Hardware architecture and software stack for pim based on commercial dram technology,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.380038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.380038Z digest=sha256:e6f7f2da584211eb948d239b5043258f0b6b5f69dae19094f2818e124e83f73f

Observation a83c82e6-8200-4657-bce0-f1358b096d23 · outbound

This paper cites Pond: Cxl-based memory pooling systems for cloud platforms,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Pond: Cxl-based memory pooling systems for cloud platforms,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.384812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.384812Z digest=sha256:72f617545bbe325a63180bd380957c7af6f278058c8f2e02e66e7816d377b985

Observation 7f38912b-813f-451a-9ab3-85f17b37e38b · outbound

This paper cites Snapkv: Llm knows what you are looking for before generation,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Snapkv: Llm knows what you are looking for before generation,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:00.327089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.389331Z digest=sha256:1caba695990a48fd278753aecdbbb9492e2438b76bf4476f82d7013e4a6440d5

Observation 7f6498db-72a4-45f6-90a3-f75c2a6c6dfb · outbound

This paper cites Meridian: In-memory acceleration for rag with document attention decomposition,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Meridian: In-memory acceleration for rag with document attention decomposition,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:00.311907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.393819Z digest=sha256:407298362942522f4dc50c01c94a5fb2e98928db4ce555b71ef99be734a3e931

Observation 71aad5d4-c91f-49c2-858b-41b4ce41751c · outbound

This paper cites Chime: A case for efficient long-context attention- fc disaggregated inference with dimm-pim,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Chime: A case for efficient long-context attention- fc disaggregated inference with dimm-pim,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:00.296688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.398419Z digest=sha256:5fd174cd7ddfd6bd300fa92154719a5b398f2267db3b2a4353501e198f028ad5

Observation 869f7343-2506-4a36-b35b-e6b316f3042e · outbound

This paper cites Nvidia’s ad102 officially revealed, how close were the previ- ous estimates?.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Nvidia’s ad102 officially revealed, how close were the previ- ous estimates?

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:00.281663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.403138Z digest=sha256:f090ae4c3f741464394e0ab4fa3a7ba33aae17e44714418f5b3d3887f7a53197

Observation 829d515a-1328-484d-a2a7-e0dfb5104868 · outbound

This paper cites MoBA: Mixture of block attention for long- context LLMs,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention MoBA: Mixture of block attention for long- context LLMs,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:00.266183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.407621Z digest=sha256:4e8654b71e76e877a850dcb77f204205b8658b81c125890cfc838b3254da067e

Observation 4b1472e6-8049-4b9b-a3fd-0d908548c702 · outbound

This paper cites Sac: Disaggregated kv cache system for sparse attention llms with cxl,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Sac: Disaggregated kv cache system for sparse attention llms with cxl,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:00.249064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.412221Z digest=sha256:f2148748fa772c25228c70fec75adbe84820d3f8eb90757c32342a30514d0460

Observation f4247f77-a36b-4cbb-bf91-6c1dde323abc · outbound

This paper cites Tpp: Transparent page placement for cxl-enabled tiered-memory,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Tpp: Transparent page placement for cxl-enabled tiered-memory,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.416938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.416938Z digest=sha256:a956f52ff05e6902bfee58b6ab77f8d86481f8c0ab95ba7e72e63f8badb1e929

Observation 64f75a88-0540-4e21-9f37-6939e5b7260f · outbound

This paper cites HBM3E product brief,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention HBM3E product brief,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:00.230476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.421441Z digest=sha256:8df5f9a13d0a3b0f17d8f9d4575a653124cccd9fa52908307447f08c4822a31c

Observation 7c1242cb-69e5-425b-9301-cc03127b2e7c · outbound

This paper cites g ed., May 2025, production Data Sheet.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention g ed., May 2025, production Data Sheet

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:00.214140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.426206Z digest=sha256:8544d9a43787f1710360e58329be546a19a10f80eb3027b598a6c428bb6eb374

Observation 7e09f9a6-ee34-4308-aa07-d9b8b4f6cbd2 · outbound

This paper cites Early silicon of raptor: The first 3d-dram accelerator for generative inference,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Early silicon of raptor: The first 3d-dram accelerator for generative inference,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:00.196710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.435190Z digest=sha256:18c31ef413ca73066a8bc033cb4e7a341bb51b990dd9e0496ca5ac9556568547

Observation 9836d34e-ce16-4ebf-89ad-ae9c2ff349d2 · outbound

This paper cites NVIDIA DGX B200 datasheet,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention NVIDIA DGX B200 datasheet,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:00.177834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.440118Z digest=sha256:67c43f979ed11d3c2520c971d240ebe72dd0c50d4aa4066beda80e3523ef5d99

Observation 1e5d862b-f97e-4795-b713-89da268bb6e0 · outbound

This paper cites NVIDIA DGX H200 datasheet,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention NVIDIA DGX H200 datasheet,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:00.155632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.444590Z digest=sha256:821a35ec7f2c08a3bbd5af0fd24130e78ecca457bfca34dff95a3d3d22b97e29

Observation bc8a10d5-e6e6-443c-b7bc-1192e7f07541 · outbound

This paper cites SLA-based planner,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention SLA-based planner,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:00.133856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.449134Z digest=sha256:e65657b67bb784c4dc970b27ad4770e8740496bbb6f36cd6aed4e410d78c5654

Observation 40a02988-eb34-4758-ae43-b55623d62339 · outbound

This paper cites Multi-process service: Appendix: Tools and in- terface reference,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Multi-process service: Appendix: Tools and in- terface reference,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:00.114628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.453507Z digest=sha256:14a894e3ad82a2f3c3deb0b158e5a1f1b2e137304f29274c76a0911f5278a312

Observation 91f4f66e-9c28-4902-9863-85885036bb75 · outbound

This paper cites (2026) Overall architecture — NVIDIA Dynamo documentation.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention (2026) Overall architecture — NVIDIA Dynamo documentation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:00.096891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.457817Z digest=sha256:5b58d0e62e7c375705afac357c288ea2c73dc9abd272ad2e0c92dcd651946324

Observation 28c3453f-11da-4719-85fc-e20312900b88 · outbound

This paper cites Planner design,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Planner design,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:00.078680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.462454Z digest=sha256:b5d37c54e9293c3c02393ecf7cea6ad2360fb428e070d732f89b2bce9e9627b5

Observation d2f30c89-eb67-4ec1-9181-69f50c4b9246 · outbound

This paper cites Supported MIG profiles — NVIDIA multi- instance GPU user guide,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Supported MIG profiles — NVIDIA multi- instance GPU user guide,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:00.057883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.466901Z digest=sha256:8cb386425fbad4fc7381f6e854007e7c1ee1ecc816c18dc047e60a97ed306248

Observation 0a47aa5c-f6f7-4039-821b-fa75e4d27b88 · outbound

This paper cites Fine-grained dram: energy-efficient dram for extreme bandwidth systems,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Fine-grained dram: energy-efficient dram for extreme bandwidth systems,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.471604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.471604Z digest=sha256:376a23173a60c63a540eb29957fa39f3e853c51904e8d2491fbd1ebf582a48e6

Observation 99c0a576-0a4e-4b7c-ba8f-b01174e431c9 · outbound

This paper cites Exegpt: Constraint-aware resource scheduling for llm inference,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Exegpt: Constraint-aware resource scheduling for llm inference,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.476730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.476730Z digest=sha256:b1ea401da58d6e6aff0c33f84495248e57ba62a716a8ec5eac554afb599ce8f3

Observation 1e245943-ff76-4cb3-a140-40ef552e5bfe · outbound

This paper cites ChatGPT (GPT-5),.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention ChatGPT (GPT-5),

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:58:00.013479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.481464Z digest=sha256:b7f97312933e89d670d762eff63ead068d6a7890c50c8277b1fc8693a828b75f

Observation ace60453-3340-4802-81d5-92b0b982d7e2 · outbound

This paper cites Attacc! unleashing the power of pim for batched transformer-based generative model inference,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Attacc! unleashing the power of pim for batched transformer-based generative model inference,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.490666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.490666Z digest=sha256:f63baef296e6ec8dca828664d2a29fd22234b94d73a88b6a74f8fbd4f8ffc304

Observation 4be5218b-1826-48f2-908a-3323625201bd · outbound

This paper cites A 192-gb 12-high 896-gb/s hbm3 dram with a tsv auto-calibration scheme and machine-learning-based layout opti- mization,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention A 192-gb 12-high 896-gb/s hbm3 dram with a tsv auto-calibration scheme and machine-learning-based layout opti- mization,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:57:59.987872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.495167Z digest=sha256:874ce1520c3936d0a40d523e8cf3eaa50069913dc4aba4f956a992e7f51d17fc

Observation 3bfbb05d-b684-4ff6-8cb7-ac3d9e0cfbc4 · outbound

This paper cites An lpddr-based cxl-pnm platform for tco- efficient inference of transformer-based large language models,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention An lpddr-based cxl-pnm platform for tco- efficient inference of transformer-based large language models,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.500163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.500163Z digest=sha256:8a37f32111d4d646ea72813d93a9430f66817146f65361c5b6ec41c03977aec4

Observation 1341262d-1f59-4dc4-b3de-789a5a15811a · outbound

This paper cites Apple m2 die shot and architecture analysis – big cost increase and a15 based ip,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Apple m2 die shot and architecture analysis – big cost increase and a15 based ip,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:57:59.969080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.504641Z digest=sha256:9091adec6d71bb707fb09cc26293c113c75d9d86781f2ec07ead8bc474e644ff

Observation 38d07b6e-9fa6-4c2c-a51e-0ade696dcac8 · outbound

This paper cites Splitwise: Efficient generative llm inference using phase splitting,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Splitwise: Efficient generative llm inference using phase splitting,

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.509113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.509113Z digest=sha256:45b74b87d4a12e1678528dd9b63b7f80c515704195ed19c644500566d6539c3a

Observation 4eaec88c-7719-437a-afd7-fd967a84a59f · outbound

This paper cites Mooncake: A kvcache-centric disaggregated architecture for llm serving,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Mooncake: A kvcache-centric disaggregated architecture for llm serving,

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.513839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.513839Z digest=sha256:b9e3778f1b91321e9d692f718a251ff9cf2a201183443729d802dd749a2e3f57

Observation be2c78d2-20ee-4db1-91b9-59affbe87787 · outbound

This paper cites Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.518164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.518164Z digest=sha256:34c3c2cb6b817f93ce5f17d659c210fe3d6fc12756c06cf934b73caba9d22199

Observation d4717908-7610-4d28-b5b2-78b1b4a751c9 · outbound

This paper cites Longsight: Compute-enabled memory to accelerate large-context llms via sparse attention,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Longsight: Compute-enabled memory to accelerate large-context llms via sparse attention,

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.527863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.527863Z digest=sha256:c913bec20e0b93394c610f5b87f64fd7eaaf06dd24ca1b9350701dcf406a91ad

Observation 0f44faa9-2424-42e0-b960-19ece375ed0d · outbound

This paper cites Drex: Accurate and scalable dense retrieval acceleration via algorithmic-hardware codesign,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Drex: Accurate and scalable dense retrieval acceleration via algorithmic-hardware codesign,

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.532714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.532714Z digest=sha256:4909d39f42a55411fd4f6856ac327c6ea42e120be7b85959a548556aeb58a027

Observation 25ac2814-db2d-45f2-8542-7c5d76ec5873 · outbound

This paper cites Sparq attention: bandwidth-efficient llm inference,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Sparq attention: bandwidth-efficient llm inference,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:57:59.951661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.537246Z digest=sha256:facf3719455188ad946a8d1386e648177684d318a058cfa20fe91cbdb107045a

Observation ecf9d468-0bbf-4022-b109-7beef9a8a122 · outbound

This paper cites AttAcc simulator,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention AttAcc simulator,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:57:59.933904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.541691Z digest=sha256:7582306260cc587bcadfa5345bf8d3f6c6d72c2145c803952339d336f08e7564

Observation 4abada64-ac32-424b-9d45-198da8bbcdc5 · outbound

This paper cites LLMSimulator,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention LLMSimulator,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:57:59.915842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.547009Z digest=sha256:4c869653f097d4e0372be42196d6afe72ec5913cbc6e361b0444fcd1436efabb

Observation 9bdce448-1a96-4225-ae22-eda8614c7bd6 · outbound

This paper cites Toolformer: Language models can teach themselves to use tools,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Toolformer: Language models can teach themselves to use tools,

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:57:59.896406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.553027Z digest=sha256:f7ef6e35d8207fe3733a14c5013ca449df448f5aa155d6e554635e6eb9228bca

Observation 78165bad-bf47-4527-a14a-fc563f2e7431 · outbound

This paper cites Ianus: Integrated accelerator based on npu-pim unified memory system,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Ianus: Integrated accelerator based on npu-pim unified memory system,

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.561610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.561610Z digest=sha256:b89f46c729ea36e8b09d25e46df78dd7209bb97c99e08f4b344973529ed2ab9d

Observation 218598e1-5126-43d3-8251-a7b889123e80 · outbound

This paper cites Dynamollm: Designing llm inference clusters for performance and energy efficiency,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Dynamollm: Designing llm inference clusters for performance and energy efficiency,

Reference 82

Resolution
malformed identifier
no resolver link, observed 2026-08-15T14:57:57.567291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.567291Z digest=sha256:9a06d6c3d8a7180b620ff172f81de08e13feec5161a65beb0ac992f9511e4d23

Observation da7ae6ab-79c6-4dbb-a06e-cce59b3194c7 · outbound

This paper cites Quest: query-aware sparsity for efficient long-context llm inference,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Quest: query-aware sparsity for efficient long-context llm inference,

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:57:59.877514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.572723Z digest=sha256:7b7f4001553a6e4405315c166235a8563adc9643928b7d896250833839c92b46

Observation 667ab8f0-9d40-440b-b974-95062c1b1d3e · outbound

This paper cites Astra-sim2.0: Modeling hierarchical networks and disaggregated systems for large-model training at scale,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Astra-sim2.0: Modeling hierarchical networks and disaggregated systems for large-model training at scale,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:57:59.859074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.578688Z digest=sha256:75c4b83091e92d9f98e2fcba6598cffb82201c04c3e2e11b5b82c845533aaad3

Observation 839c0458-9fd5-4f3c-a139-8460017ad2ea · outbound

This paper cites MemExplorer: Navigating the Heterogeneous Memory Design Space for Agentic Inference NPUs.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention MemExplorer: Navigating the Heterogeneous Memory Design Space for Agentic Inference NPUs

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.583167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.583167Z digest=sha256:df6330dcb6f5e06c4ea4964ac77a30a4412b29b81dc2804800014e299aad77ec

Observation e503753f-a17d-4fc5-8d8e-ab0748aa2ea3 · outbound

This paper cites Efficient streaming language models with attention sinks,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Efficient streaming language models with attention sinks,

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:57:59.840387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.588354Z digest=sha256:e5ccd9bf4b0b9a42f58b6901095e4d4f5ab264634bc683d4c21c885719ea4332

Observation 548abcf4-5487-455f-8d76-53761d9825d2 · outbound

This paper cites Hisparse: Turbocharging sparse attention with hierarchical memory,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Hisparse: Turbocharging sparse attention with hierarchical memory,

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:57:59.817534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.593155Z digest=sha256:32611b083e9044cc746347a7caa6c0369f04d41f03e23f55592fa063fca2fb80

Observation 06aa777f-2162-43e5-90a6-415a16325328 · outbound

This paper cites Strata: Hierarchical Context Caching for Long Context Language Model Serving.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Strata: Hierarchical Context Caching for Long Context Language Model Serving

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.598073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.598073Z digest=sha256:f0a909bcd192388caf388451e643a70d89a4effcb1aaf001b10069319fa0000d

Observation 61352f64-28ee-4604-9d51-645e6cc9726d · outbound

This paper cites Beluga: A cxl-based memory architecture for scalable and efficient llm kvcache management,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Beluga: A cxl-based memory architecture for scalable and efficient llm kvcache management,

Reference 90

Resolution
verified exact
doi, observed 2026-08-15T14:57:57.723542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.613579Z digest=sha256:8b9228a446a000e23572ba0c0414f6a2c81c8fc9704465cdbfba04495ec35f59

Observation 5a859f43-3852-4bb0-b0a9-ed6b4028d22d · outbound

This paper cites ReAct: Synergizing reasoning and acting in language models,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention ReAct: Synergizing reasoning and acting in language models,

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:57:59.799077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.618785Z digest=sha256:148c5d90e5363589ca596afe9b0a8409bd94d8ddfbba25eef8bd76d3d89433df

Observation 850db07b-93a5-45a3-9604-7db34f945c25 · outbound

This paper cites Tract: Disaggregated llm serving with cxl shared memory kv cache at rack-scale,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Tract: Disaggregated llm serving with cxl shared memory kv cache at rack-scale,

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.623910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.623910Z digest=sha256:d649d13069597f6bf563b017fc5fd035fefc632c92d251cb45fc007703b71ec1

Observation 2a5ed225-56ef-45e6-8cbb-da43b1be8389 · outbound

This paper cites Available: https://arxiv.org/abs/2601.06288.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Available: https://arxiv.org/abs/2601.06288

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.608756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.608756Z digest=sha256:e227da22bb7632e15d818892d6df572dd8ba1f61430d2de2ea09c74f93a528bf

Observation 1737c7af-5fd7-4fad-b374-3efcad806983 · outbound

This paper cites Duplex: A device for large language models with mixture of experts, grouped query attention, and continuous batching,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Duplex: A device for large language models with mixture of experts, grouped query attention, and continuous batching,

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.634586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.634586Z digest=sha256:a1a75acd6f4f8cfe2f52c5c9f25e2efa216232f9f4c998fc543d7aaddb366a09

Observation e1732f4f-713b-41b3-a128-263a10bb69cd · outbound

This paper cites H2o: heavy-hitter oracle for efficient generative inference of large language models,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention H2o: heavy-hitter oracle for efficient generative inference of large language models,

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:57:59.766301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.639094Z digest=sha256:4f5617c55a610aa5a8c7327563010f0201fa860a360fc26a5ee993b91793ac44

Observation 71d4a0a6-e8ec-4b87-a3d9-5714e972d590 · outbound

This paper cites Deepep: an efficient expert-parallel communication library,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Deepep: an efficient expert-parallel communication library,

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:57:59.749108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.644038Z digest=sha256:879cb80537bfe18b0ba1371579d6ea30097e2eacd537ef54a1b41d5292aabafc

Observation d5be8cc2-997c-4877-83ef-d897fb643fe6 · outbound

This paper cites Patterns behind chaos: Forecasting data movement for efficient large-scale moe llm inference,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Patterns behind chaos: Forecasting data movement for efficient large-scale moe llm inference,

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:57:59.782718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.628881Z digest=sha256:ee979e0522ebb3c1170b27042c0cff27193e83bb18d54158d44e2750c42dca36

Observation 41ca0024-80fe-42d2-b55e-9edb60e1fb5b · outbound

This paper cites Sglang: efficient execution of structured language model programs,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Sglang: efficient execution of structured language model programs,

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:57:59.715187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.653026Z digest=sha256:24556d4e6855d7668a91898aaf839eebcfbb49464345f7952cd1afd262e29b87

Observation 3907fcb6-17cb-48ca-8e76-7207edf8feee · outbound

This paper cites Distserve: disaggregating prefill and decoding for goodput-optimized large language model serving,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Distserve: disaggregating prefill and decoding for goodput-optimized large language model serving,

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:57:59.695869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.657267Z digest=sha256:555a73faed721df53d99777f1cd97aa3dc0e16a2240900338894b1eb711b458d

Observation e239bc84-b2d3-4b0d-ae2d-c14ff1bd4364 · outbound

This paper cites Octopus: Enhancing CXL memory pods via sparse topology,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Octopus: Enhancing CXL memory pods via sparse topology,

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:57:59.678165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.661606Z digest=sha256:332c380a048e674d1409f960916a6b2490b4c040b63bf21d514232df069479d7

Observation 0c47f604-1cf0-4652-8f5f-2db614f4fe2f · outbound

This paper cites InfLLM-v2: Dense-sparse switchable attention for seamless short-to-long adaptation,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention InfLLM-v2: Dense-sparse switchable attention for seamless short-to-long adaptation,

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:57:59.732854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.648545Z digest=sha256:3d96f8aa8b2154d83c5c61fb1d47086e38fcd8f24415cbf7c5ce96d368757725

Observation 8e9f74b2-3a44-4952-8547-279579b2989a · outbound

This paper cites Megascale-infer: Efficient mixture-of-experts model serving with disaggregated expert parallelism,.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Megascale-infer: Efficient mixture-of-experts model serving with disaggregated expert parallelism,

Reference 105

Resolution
verified exact
doi, observed 2026-08-15T14:57:57.707092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T14:57:57.666022Z digest=sha256:5d4ab228040925af18783ca9623e727d421a9cd54cda2934a35e6b8bb4a6f207

Observation f2a2bbea-f4d4-4576-85e9-cffcebd4f8ee · outbound

This paper cites Available: https://doi.org/10.1145/3579371.3589348.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Available: https://doi.org/10.1145/3579371.3589348

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.256627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.256627Z digest=sha256:98dec504d281f4a752c795232a5f18996453ed3b984afb1bbc576576e4293380

Observation 278d6a21-d664-440c-900d-9e99ab57a0c5 · outbound

This paper cites Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.523025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.523025Z digest=sha256:d2ba4d92f89560d2929bb26a941199e3536f64e2c3c3f271a111af2d2f80e9c6

Observation 577d7638-0e16-430e-a748-bf2dae12ae21 · outbound

This paper cites GLM-5: from Vibe Coding to Agentic Engineering.

Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention GLM-5: from Vibe Coding to Agentic Engineering

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:57.283172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:57.283172Z digest=sha256:735ee9275aca432d03cfdf77715ebd8c1b36faf7d9100eed9abb39cd4c33cd8f

Pith citing papers

No inbound Pith citation observations are available.