Pith. sign in

Paper Citation Record · LEDGER

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration

As of 9 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2506.11104.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11104 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:03:19.387482Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:13:27.644720Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T05:13:28.150379Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact3
  • verified fuzzy1
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 39b75541-84d0-436c-a687-1d832d3d8a51 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:03:19.875957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.265408Z digest=sha256:dd5c4fb7afd07b5f27be8d4b276a413741cb7d78f92f3e4103c87043d7357f8e

Observation fd441fa5-0ea3-4132-a475-96c32bb8aacd · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:03:19.866737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.269172Z digest=sha256:648f69dac46b9733d801e5559a2ae882bd0748391349d9792e2cc94746ba14b8

Observation 3260fa7d-8bec-493d-a95d-2eb25927ded6 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:03:19.857783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.272325Z digest=sha256:07b68c36b472f63d69089c2f0c621ed0bfa2fe1384ba1d63290b9b3b32d7cee4

Observation 0aa85b62-32dd-445b-8034-ac4fda8bcde0 · outbound

This paper cites Longformer: The Long-Document Transformer.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Longformer: The Long-Document Transformer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.278649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.278649Z digest=sha256:5df291450f9f3d07ea0dfc2097d3f3bed0c3851dec05507af92738cf26c3a027

Observation 2a75ea7a-6630-4698-aa96-2d9059574388 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.281563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.281563Z digest=sha256:146f349fbc7714bb7f848a2e35549e3d0f6429fc19de254a14ab1c4011554899

Observation 109f8c1b-18d4-40e9-9e3d-309081fc6931 · outbound

This paper cites NACL: A General and Effective KV Cache Eviction Framework for LLMs at Inference Time.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration NACL: A General and Effective KV Cache Eviction Framework for LLMs at Inference Time

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.284561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.284561Z digest=sha256:23188989e5a6a6b6cd4dd43f4c4ab31d63def19d03742576b9027575ec08c5ca

Observation b5e93402-0bcf-4b35-aa7d-d1b0f66e2241 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Generating Long Sequences with Sparse Transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.287621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.287621Z digest=sha256:0b7081960d4285239a99a421e749b562b22ffd72ee5d31a053702a0390e2ff6b

Observation ace05112-41fc-445a-8fff-53ea8b05380c · outbound

This paper cites Masked Language Modeling for Proteins via Linearly Scalable Long-Context Transformers.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Masked Language Modeling for Proteins via Linearly Scalable Long-Context Transformers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.290731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.290731Z digest=sha256:566f5a7bb03b1bdb7f5bac0c1c807aee5be5fe2a9a538f476b54eadd45d95d0b

Observation 2ed08cd9-9e48-4ab4-a16c-a33e2b3f3383 · outbound

This paper cites Adaptively Sparse Transformers.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Adaptively Sparse Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.293695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.293695Z digest=sha256:e6739f9ea233df17eac7292b1f929744b5f190d619bf6384adb2158487a33539

Observation 72b51fed-4cae-4cf4-a006-764cc2c0b555 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.296945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.296945Z digest=sha256:e2b54d45bbdc2726a49155c727c1dea5326e65c6fdf5b3b7435edd166e738418

Observation f462fe05-3d87-4621-b1b1-f213bc3d6e00 · outbound

This paper cites Multi-News: a Large-Scale Multi-Document Summarization Dataset and Abstractive Hierarchical Model.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Multi-News: a Large-Scale Multi-Document Summarization Dataset and Abstractive Hierarchical Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.299889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.299889Z digest=sha256:5c3c4b05deaa875dacebeb595d0dec6e1ddd05a702c49f06c37a41dee8010988

Observation a38fa083-ed9b-4ed8-9083-7a3cbef28524 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.302732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.302732Z digest=sha256:d211da54056a966172172dbed538e4b7c258c457fc9419ddbb0ffaa94e9ab1d6

Observation b593f040-8eea-4bca-9588-07bbd603a0e1 · outbound

This paper cites Semsa: Semantic sparse attention is hidden in large language models.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Semsa: Semantic sparse attention is hidden in large language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:19.849396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.305392Z digest=sha256:dc1a7c3ce1132ebb1a4fca1d407fbea4d96cabba5a3fb53e33cf75fe4d213db7

Observation 38126815-345c-449c-97d8-caf3610bff33 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:03:19.840241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.308260Z digest=sha256:af06dcfb525c1e7d40e7734ee04110a10844e8efb0e11b923fb566f5d2b12d47

Observation 11123903-98f1-4399-8696-71a83d51e448 · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.311000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.311000Z digest=sha256:a95988b1b368762cd02ec1b61456091998f7539e43f649d06a2af8e61195d5ff

Observation 7dfb9399-0a44-4a50-9657-3da18bb4e273 · outbound

This paper cites A Dynamic Head Importance Computation Mechanism for Neural Machine Translation.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration A Dynamic Head Importance Computation Mechanism for Neural Machine Translation

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:03:19.616366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.313839Z digest=sha256:10d85fef2a2b20988565a8743ddc794330f099b2d94b11895914b402e2928cc5

Observation a9404664-de59-48ed-99fe-06cc174f009e · outbound

This paper cites Axial Attention in Multidimensional Transformers.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Axial Attention in Multidimensional Transformers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.316871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.316871Z digest=sha256:0eef89e92d65f7a5b18eba9d468fb383b8ea1d8d2b62de0fede1e8def9ac2020

Observation 964a02ee-e3b0-467c-b77d-b3cd16ae2297 · outbound

This paper cites MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.319925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.319925Z digest=sha256:0750246203eefe5a9205b939da423cf9ec365aa19553109f994bed4bc2cbdf11

Observation 60a43518-9018-4fbb-a493-0f4020fed5a6 · outbound

This paper cites Reformer: The Efficient Transformer.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Reformer: The Efficient Transformer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.322704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.322704Z digest=sha256:07381c199bb6bbe6111bbe39004bffc34553fa1a3614e43ee5b780abaf01a13b

Observation 77f0c218-0e94-4bb1-80cc-3958813350b3 · outbound

This paper cites LongEval: Guidelines for Human Evaluation of Faithfulness in Long-form Summarization.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration LongEval: Guidelines for Human Evaluation of Faithfulness in Long-form Summarization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.325755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.325755Z digest=sha256:c883d2df579a97450e5561a1c21eecb04d55d6c000a3795b89ea9ca1797df257

Observation dff2d4cb-47cf-4d2b-82dc-4e4ec51871a2 · outbound

This paper cites SnapKV: LLM Knows What You are Looking for Before Generation.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration SnapKV: LLM Knows What You are Looking for Before Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.328603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.328603Z digest=sha256:3a3f20c287f8414b35df7e223e042d761e65ad6d688d5672e4cf15a4e7f9e91e

Observation c5eacb4a-1b9d-4430-91be-00d18496f4a5 · outbound

This paper cites Global Attention Mechanism: Retain Information to Enhance Channel-Spatial Interactions.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Global Attention Mechanism: Retain Information to Enhance Channel-Spatial Interactions

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.331633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.331633Z digest=sha256:7012e1f3654ba3bde98d12e4b7ece6e86eb952be1a823e72573f55fa11069730

Observation e8386c03-f072-4bf1-82c1-b59efe1bbb23 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:03:19.831550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.334602Z digest=sha256:96f633879ff854d98d8f4489aa07cfc24c21f01d827f857bb83f3a2729772991

Observation 47463b5e-831e-45b2-a3ba-4524dcb270d4 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:03:19.822960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.337352Z digest=sha256:e4595f6ec4b790c008ff3010bb12b63738f7922c4e43e7011f111cab014f9bf5

Observation b849773a-ef26-49c0-b491-92f618feca64 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.340101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.340101Z digest=sha256:ddbc00a15838b856c1192fbe5a4036570b95d34baaa64f2dc5e03243c25c1a70

Observation 723e7c91-852a-4b51-baeb-20f365ac3591 · outbound

This paper cites Lightweight and Efficient Neural Natural Language Processing with Quaternion Networks.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Lightweight and Efficient Neural Natural Language Processing with Quaternion Networks

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:03:19.557693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.342944Z digest=sha256:cde930752a777a7e5e3d2ee4a66cd9a43675851205c3380be3fc55ed84e2c7de

Observation aee541e8-2104-418f-88f2-f80226d18deb · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.345983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.345983Z digest=sha256:0ffcbc6c0508ee965abac1fba5f0164eec0d314524bc670dffe1a11ba5911086

Observation 53206f46-2ae0-49ae-9df8-68285976e6a3 · outbound

This paper cites Multi-Head Self-Attention with Role-Guided Masks.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Multi-Head Self-Attention with Role-Guided Masks

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:03:19.545146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.348741Z digest=sha256:1e0c99e3c2538c79b98bf2c36253c3465dd07891cab6fd41317fae089821298d

Observation 78a2e262-6928-4375-ae3e-d37f98c80da9 · outbound

This paper cites Improving Transformers with Dynamically Composable Multi-Head Attention.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Improving Transformers with Dynamically Composable Multi-Head Attention

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.353231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.353231Z digest=sha256:e62f8e0ceb92f82a310c3582afdaaf05cdcd24d9f160eb3ffd1790a32c18f62c

Observation 4ad6a6f0-2888-4cc0-bc53-f5b9ee758c26 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Efficient Streaming Language Models with Attention Sinks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.357102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.357102Z digest=sha256:f257432b02b98b65b77bf8becde1b8a9495003009c6bb01afecef1e4e66aa345

Observation f770c5da-57e3-4456-a076-a60800a339bf · outbound

This paper cites LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.359687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.359687Z digest=sha256:519faf2178e2c36090482cda30b36eb09312cf3119f4a81fd8a211172150f1dd

Observation d96aaaba-19ec-46e6-b609-794c52746810 · outbound

This paper cites ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.362463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.362463Z digest=sha256:d973358f6d332c2c7999a65114b6cf7718de69da234a1d7c9739c97b211428f5

Observation 582542cf-e636-4fcf-a262-cdc2a48eb119 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.365339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.365339Z digest=sha256:87007d52320bc1d3c6a81d2d961de79365c87676db069d5af2b82a87ffc36e75

Observation 92cf65c7-5490-460f-8c4f-d8e12ac71434 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:03:19.803606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.368297Z digest=sha256:a1b5e475ac7d1abe8adeff612e6fc707d29afb05d6d60f7028826e9da083db83

Observation 6af78f03-c580-41bb-a8e2-244859840541 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.370887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.370887Z digest=sha256:c3a4d20d374f157e16199f93553b36d286c7e026e6dc8745c1205362346e80ff

Observation 5814662f-8955-4c90-9230-7afdb02297a1 · outbound

This paper cites DiffKV: Differentiated Memory Management for Large Language Models with Parallel KV Compaction.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration DiffKV: Differentiated Memory Management for Large Language Models with Parallel KV Compaction

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.373594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.373594Z digest=sha256:533dda1f7a4e2b34e9da8e8da57cfafc88fa24ef92158252f30f01e19b785c59

Observation e15d8fcf-d490-405b-bfa2-f6f4f5218f1c · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.376481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.376481Z digest=sha256:8c8d4bf637bbb2f47169209a7b440e8f036fecef13a7724d5fb2a33d310c543b

Observation ac3977f9-50d3-4bac-a657-5d6f08e176a4 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:03:19.784410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.379033Z digest=sha256:1492f3af7ea29385bf79c21784fcd4be2a09fe39ff4d00c6e33981c5702c401f

Observation 0fa2375a-4a90-4325-82ad-523cb12c7c76 · outbound

This paper cites BUZZ: Beehive-structured Sparse KV Cache with Segmented Heavy Hitters for Efficient LLM Inference.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration BUZZ: Beehive-structured Sparse KV Cache with Segmented Heavy Hitters for Efficient LLM Inference

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.381771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.381771Z digest=sha256:bf57928f658563e99845a12f6021253c53605445804054fb2998b70c595a5204

Observation 8ad58d27-eea1-4510-8dc9-59d415f56a26 · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration SGLang: Efficient Execution of Structured Language Model Programs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.384593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.384593Z digest=sha256:6956b915ec992f6c0ee56773bc76677922bcfeeec4705e385d3e7b099a171759

Observation a6411a1e-082c-4e41-a447-5238a7d44b9c · outbound

This paper cites BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.387482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.387482Z digest=sha256:ab5a105c1753d23fb762c630f4bc658289a0461c144cfe02ba491be3f4a9cb73

Pith citing papers

Observation 3d0b4a51-b3ee-462f-9e79-f2da65fab369 · inbound

An Overview of Algorithms for Contactless Cardiac Feature Extraction from Radar Signals: Advances and Challenges cites this paper.

An Overview of Algorithms for Contactless Cardiac Feature Extraction from Radar Signals: Advances and Challenges DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-06T05:13:28.210636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T05:13:27.644720Z digest=sha256:47463f7f3ec534a5ea35deeeaecf349347c8589203f9c66622babf931b583b65