Pith. sign in

Paper Citation Record · LEDGER

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention

As of 19 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 1 inbound Pith citation observation for arXiv:2605.18753.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.18753 v1

Coverage vector

measured 78 of 78 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T10:50:12.926232Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T05:40:56.557646Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

78 of 78 outbound references displayed

  • verified exact28
  • verified fuzzy42
  • unresolved4
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6d5ba521-272d-448a-8673-71cacd619d3a · outbound

This paper cites Is it really long context if all you need is retrieval? towards genuinely difficult long context NLP.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Is it really long context if all you need is retrieval? towards genuinely difficult long context NLP

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.146562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:cb1e8eec21cc8485586937d9083233f74c72a29a038baf292e39fca7d4be1a61

Observation 929d89d6-30cd-40ee-8a68-1b6381af3f11 · outbound

This paper cites an unresolved cited work.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-05-20T10:53:26.135459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:ef3d76fee49eabd7d334566e329de7b0942ba6e9184066f79a2b95952816392f

Observation 60e5711f-49c1-4cba-a6b4-9e445e33c176 · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Attention is all you need.Advances in neural information processing systems, 30

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.133698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:1ace97ebd7df1ba1b00137044ea24bebc898414cb0162f2ca3045d4d5c687d90

Observation 7cde45ae-baeb-4fda-b8f7-8d05985f8c39 · outbound

This paper cites Softmax is not enough (for sharp size generalisation).

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Softmax is not enough (for sharp size generalisation)

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.144002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:52edee8fa2607d91ef89ca7190d93113be506ec9de92a0fae9e069e82fc18d74

Observation c3e2f403-7f6e-4f7d-8a65-294a897559f5 · outbound

This paper cites Native sparse attention: Hardware-aligned and natively trainable sparse attention.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Native sparse attention: Hardware-aligned and natively trainable sparse attention

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.129008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:d9bf0c6f8b48183e705f7debfa6f0b5ef28bcb644654b6a0f3511c0ce15bb5d7

Observation 03f36460-b66b-4402-8a46-86c456e2fb4f · outbound

This paper cites InfLLM-v2: Dense-sparse switchable attention for seamless short-to-long adaptation.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention InfLLM-v2: Dense-sparse switchable attention for seamless short-to-long adaptation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.137533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:d02a95a5a5e1b4c90ee691f188b3f9a399982b9ec6b5374fc87143306c8460a9

Observation 0068619b-f1ab-4ccf-b610-bbd56a088321 · outbound

This paper cites an unresolved cited work.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-20T10:53:26.141925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:63c15bdb6a51d9560cb59947ddad5d7bcf8de843d6171877a1eb8811ab79f853

Observation 13a84dfa-7a41-42fd-85bd-749ddcd466e9 · outbound

This paper cites Long-context generalization with sparse attention.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Long-context generalization with sparse attention

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.133509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:d4b7f3b86b1450a428090e657632d7421bc8daf7a1bc08b55091fb36c941254d

Observation f54f691a-91c0-4bc2-8ee2-a6fdd31a56a4 · outbound

This paper cites MoBA: Mixture of block attention for long-context LLMs.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention MoBA: Mixture of block attention for long-context LLMs

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.102675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:a28e800bf32b887f7598096f3ae883348214408c7827bb8f5d2305b3d684fcb2

Observation 96437640-d12c-414b-a20d-477cfe944867 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with IO-awareness.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Flashattention: Fast and memory-efficient exact attention with IO-awareness

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.094757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:6cf6e8238885158bfea65f9dce9589c1f11d77057e66f163619a7d98111c4a3e

Observation c1e71adb-e692-4559-91b5-c1e49de57ea7 · outbound

This paper cites From softmax to sparsemax: A sparse model of attention and multi-label classification.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention From softmax to sparsemax: A sparse model of attention and multi-label classification

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.090382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:0041a05721fcb3b1bf6f9a9353ff8fadf8d9cc82b913e2f605df897cfb300c4a

Observation 6d29233f-6a4b-4841-bc28-040776463dbf · outbound

This paper cites Adasplash: Adaptive sparse flash attention.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Adasplash: Adaptive sparse flash attention

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.122761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:0ebacd77993150c2e7d3c4968afb8ee447b0754d7a5d8371ad69357af384c78e

Observation 28d35cc7-d532-45bc-bb60-7d3e98144cba · outbound

This paper cites Adasplash-2: Faster differentiable sparse attention.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Adasplash-2: Faster differentiable sparse attention

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.092660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:fbe291d10be50030071ebc4440bed7164685bc7c3385947213f889eceb3785e4

Observation 411b1c93-e3a7-4f79-8df0-f8e783e4eea6 · outbound

This paper cites Gqa: Training generalized multi-query transformer models from multi-head checkpoints.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Gqa: Training generalized multi-query transformer models from multi-head checkpoints

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.109441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:14f1e968cb69b5da922f4ba156056c4230c6f6f239dd89105223aa59780528ce

Observation 6c096a2f-2461-42cd-a942-6d7e043880d6 · outbound

This paper cites Learning classifiers with fenchel-young losses: Generalized entropies, margins, and algorithms.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Learning classifiers with fenchel-young losses: Generalized entropies, margins, and algorithms

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.135662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:c481ea312382617142aa7a9083d1fa53b98e1c248356f0093f1180c9adf34146

Observation 2127a77a-bee5-4144-8add-8b4dc98c9bba · outbound

This paper cites Flashattention-2: Faster attention with better parallelism and work partitioning.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Flashattention-2: Faster attention with better parallelism and work partitioning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.107670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:1a388625a02476b9ee96d01b2d4450c73f5d60407e5f645746780081708ae369

Observation 745df542-c48f-46b2-90d7-bb65983a9acf · outbound

This paper cites an unresolved cited work.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-05-20T10:53:26.107447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:54389d97bdcb0b1f646312dc753495c9dceba36fd3c25aefa1938b241bc6f9b8

Observation 382caca0-f26e-4ce1-a1d1-42e43378e158 · outbound

This paper cites Infllm-v2-data-5b dataset.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Infllm-v2-data-5b dataset

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.111009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:a24bb026740213a816b35f64dbe11615d365f8923842c6a1fce32a41f6232ffe

Observation a219ed25-0cbd-4564-a9bb-26d52baf25a1 · outbound

This paper cites MiniCPM4: Ultra-Efficient LLMs on End Devices.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention MiniCPM4: Ultra-Efficient LLMs on End Devices

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:53:13.541414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:d4c345d7bd53c6bcaab55cac3c200912efd28a840b9d8b4f2b5137ae4b6b8b50

Observation 5c9972ef-8e6a-4392-8a57-6ab1c3edf280 · outbound

This paper cites RULER: What’s the real context size of your long-context language models? InFirst Conference on Language Modeling.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention RULER: What’s the real context size of your long-context language models? InFirst Conference on Language Modeling

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.112768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:cce78147369b59d8f2ae7d3adf8bfe8647cd1990ec50e9da40b885d08040aaa4

Observation c454c62d-091b-4d39-a19f-f9b27fb03b8f · outbound

This paper cites HELMET: How to evaluate long-context models effectively and thoroughly.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention HELMET: How to evaluate long-context models effectively and thoroughly

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.109072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:2322b29b57c0c7190a3d512a388d3ffbdf7f08ada123d9a42728b99cd1d574ce

Observation 1e11a913-1c19-45e9-b9e3-b527eda3301f · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.Advances in Neural Information Processing Systems, 37:95266–95290.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.Advances in Neural Information Processing Systems, 37:95266–95290

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.142415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:2b548f5bfdf4d1a8e45413fde6dddd9cf2ad1132457340fc3864fadc57a9e4ba

Observation 066a030b-ac45-4354-b6a3-e1b65e8d6471 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Measuring Massive Multitask Language Understanding

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:53:13.574632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:d0d6177e4308896373269a0e152e978be2f47d391e74eb2edf9d7b51eedceacb

Observation 2ea2c98e-808c-43b1-8933-551ad4bf317a · outbound

This paper cites Commonsenseqa: A question answering challenge targeting commonsense knowledge.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Commonsenseqa: A question answering challenge targeting commonsense knowledge

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.124815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:dbed307e8cc8472dea975fcc68df154c49b320acff8cdc7d9a89a16edcd9ad6a

Observation 4c26b643-b332-4ed3-92a0-779477d99aff · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Instruction-Following Evaluation for Large Language Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:53:13.559200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:637e911a5a3d25d705a4e0b3e2c27753e53b1aec01473dadd396e2d04997a3f4

Observation 721e1ca2-1129-40ab-a4fa-c36d7c65a36c · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? InProceedings of the 57th annual meeting of the association for computational linguistics, pages 4791–4800.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Hellaswag: Can a machine really finish your sentence? InProceedings of the 57th annual meeting of the association for computational linguistics, pages 4791–4800

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.116548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:8bff65d30c5e88e9699854f310a6038786a2ad0b1fedb7a12e8cfd0cfed5da2f

Observation 51b7de01-03e3-4df9-820f-149602f4de4c · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Training Verifiers to Solve Math Word Problems

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:53:13.502184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:9ac576291879b3758d36d1fca8f12691b9f3324ad4c3b386f071437e78252e55

Observation 4fce58cb-1e38-4b55-9c22-2bccc3a88622 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Measuring Mathematical Problem Solving With the MATH Dataset

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:53:13.577323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:a920378da6da7b944b623e930d28e288078146bdd14c67d84a635af474aead31

Observation 6af8292c-5eeb-4c28-a865-35a21d11dbfe · outbound

This paper cites Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.126760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:a3e63bfdc4c8bde5636828d6be3989760695696a5310854a189e8e1f23a4e143

Observation 4431f30c-6e91-48fc-8b48-4d39681f5306 · outbound

This paper cites Program Synthesis with Large Language Models.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Program Synthesis with Large Language Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:53:13.556324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:42010741a54149de98b6ed97661c177d9317f63ccb0a7f07204a6c79861257d2

Observation b9bd9f9c-2486-4072-8da8-92df5d05d670 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Evaluating Large Language Models Trained on Code

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:53:13.519507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:5f0a885c0417d7db5793f4121c4c15a2dfc74a614f8b1932a723b4bb4b37312d

Observation d210fddf-4bd7-43fc-9f7c-54e5f20d2873 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Efficient memory management for large language model serving with pagedattention

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.118486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:ecf0ce98c43466e36f66c02eebc3035f956f72416ee2b3a9545849f4e9a577d4

Observation 308a8b91-229c-4f72-a63b-f3d9af120b49 · outbound

This paper cites Sglang: Efficient execution of structured language model programs.Advances in neural information processing systems, 37:62557–62583.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Sglang: Efficient execution of structured language model programs.Advances in neural information processing systems, 37:62557–62583

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.144440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:4e228d2f9d5ad383a0c6a4ca28d181056614703b3588ea97f6b16218d464de6d

Observation 61c50e82-deb7-4d3b-83e3-0f96536ea180 · outbound

This paper cites Flashattention-3: Fast and accurate attention with asynchrony and low-precision.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Flashattention-3: Fast and accurate attention with asynchrony and low-precision

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.119438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:e22b7dcb87c1561c3f20a67d1fb001259400e4480f39d57ced3a13451bd8be41

Observation 5991786b-911d-483f-b8fe-2a03898e7ed1 · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:53:13.596117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:6fa57461d2867bbf1af3dcde0ac17a08c074c5f4485a6f1753934f841af7b86f

Observation 7fd93b9f-ad9a-46b0-a97e-c4efac714fa1 · outbound

This paper cites Pyramidinfer: Pyramid kv cache compression for high-throughput llm inference.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Pyramidinfer: Pyramid kv cache compression for high-throughput llm inference

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.115897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:1cb08263bdb30d0f3402b89d3b6ae5c463055fe9499e769634182cfe784f74df

Observation 89fc6086-267b-4d24-960d-16c253156d00 · outbound

This paper cites Efficient streaming language models with attention sinks.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Efficient streaming language models with attention sinks

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.075215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:1d6adc920b1ccc99c75396a62b67ccb8ae240326d3e2b7bcb2fc104e6cc6b6aa

Observation 856ad4ff-07df-434d-8a66-40f646d9527c · outbound

This paper cites Longformer: The Long-Document Transformer.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Longformer: The Long-Document Transformer

Reference 38

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T10:53:13.553401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:c4cb23d993032e393a125d1b7e661b1d2b2d55ef38d62659c36280920a4ad5a9

Observation fcd1287e-80be-43f2-bb8a-2b4996452c3e · outbound

This paper cites Big bird: Transformers for longer sequences.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Big bird: Transformers for longer sequences

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.073490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:2efec268b35836536523910426a238943912dcba0eb843b670079affeaa38fee

Observation 4af70864-ffa5-4506-b5c9-9afcf20246bf · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.114210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:63e0fa8e65c493e81e7cac523edaddc0d7b1eac8573782aa8a598588bca44af7

Observation 44cbbdd6-45e9-46b0-93d7-d27c7a5d104e · outbound

This paper cites MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T10:53:13.547437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:39dc02189f931ebfbaf63c22df42430365b9705fc3ba992bb0512f056b5c86b5

Observation f444d31c-07f9-4262-a33b-523a679281c1 · outbound

This paper cites Reformer: The efficient transformer.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Reformer: The efficient transformer

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.117767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:964644ecadf3d42db020e84211bf1f5c37de01d003ed1c42199674d199a387e1

Observation 61ed1033-3856-4c6c-932f-6bc89d8bd961 · outbound

This paper cites Spargeattn: Accurate sparse attention accelerating any model inference.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Spargeattn: Accurate sparse attention accelerating any model inference

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:53:13.565342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:23d62337894a0a2885757b897c56d62b036e0ff0e745c3cde2a8f66d917a986a

Observation 2de1d86a-f818-4c0c-924d-118e52d8a62a · outbound

This paper cites Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:53:13.550501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:6a9a6e0fdab91c8fc42ac6272e4154f0a268cf7204d2048ccdcf2c4d37f970e3

Observation f76f756c-f240-4be3-b87e-f23f48ce80e8 · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:53:13.571507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:3ffba1222a874c7e5fbbe838fd08e7b0a71ddec4ec38a1b8f672b938ac4f64c6

Observation a4ec7517-e142-4d4d-be6b-264149bff7ed · outbound

This paper cites DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:53:13.580322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:a967ff972c7a96e897d6ffe71975118fbbbe51fe84b38088af17241a663043d7

Observation 5bea1bae-d609-4e79-8070-001742a9a948 · outbound

This paper cites Lycheedecode: Accelerating long-context llm inference via hybrid-head sparse decoding.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Lycheedecode: Accelerating long-context llm inference via hybrid-head sparse decoding

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:53:13.568362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:a155772a849a929cac69d088a4ef2d93d3f7387baf524b2f3e119ee4aef50ff8

Observation 0a4c63d7-30fd-4472-a06f-48ab893eb12c · outbound

This paper cites Snapkv: Llm knows what you are looking for before generation.Advances in Neural Information Processing Systems, 37:22947–22970.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Snapkv: Llm knows what you are looking for before generation.Advances in Neural Information Processing Systems, 37:22947–22970

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.105830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:842a4ea84b92976496db45f3717608d0fe5cb3e4e66c2ffd01e6ff7569ae82c8

Observation 242bae82-e641-4f8c-994f-c90ba31b61d1 · outbound

This paper cites R-KV: Redundancy-aware KV cache compression for reasoning models.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention R-KV: Redundancy-aware KV cache compression for reasoning models

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:53:13.583802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:847d52d3c0034af1a1b04a96426b7221a8aa67c802ed5e736ce25c17cb50a1f6

Observation 84c5d80e-72f7-430f-b5f6-4a638ee239d4 · outbound

This paper cites Indexcache: Accelerating sparse attention via cross-layer index reuse.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Indexcache: Accelerating sparse attention via cross-layer index reuse

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:53:13.529441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:f4966f3a5c2cc182e8e12e005b47d9cbdf8b34cf855d49a24278ad37f29135f4

Observation 5aa1b51d-5b99-4871-8de2-8d4d7dc237ee · outbound

This paper cites Infllm: Training-free long-context extrapolation for llms with an efficient context memory.Advances in neural information processing systems, 37:119638–119661.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Infllm: Training-free long-context extrapolation for llms with an efficient context memory.Advances in neural information processing systems, 37:119638–119661

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.121386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:a105c45c4bacce05a611010f71432c77079d8a98ee0991ebe4afdfdd441b322a

Observation 2de173d6-f8da-40a1-b937-7a2692db3acb · outbound

This paper cites ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:53:13.593270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:99e8aea1e693e623b516b6c58389c94db00a57c4f253349315a515f9069fbe7d

Observation 0a017d96-554a-4abc-852f-ab2b4b2a37fe · outbound

This paper cites Nosa: Native and offloadable sparse attention.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Nosa: Native and offloadable sparse attention

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:53:13.589945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:9d2190d83303e75c50b90079187d9db6b5b89f2f7394b1ddf53359ead60f8fdc

Observation 178987df-f6f3-4739-a48d-eabf26415659 · outbound

This paper cites an unresolved cited work.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-05-20T10:53:26.112291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:13a7b749659e5f6b09da45b7d7ff6c6d0990ff709f682b1f63b2421a09904bc4

Observation d5245d6c-a565-4287-8f3e-f21f15e45ead · outbound

This paper cites Inference-time hyper-scaling with KV cache compression.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Inference-time hyper-scaling with KV cache compression

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.148533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:027e9ef5e27a5a8230e2cc5e2dcc634a7a29706da6a88a545c4314611e596153

Observation 6be08aa3-d628-42e7-8402-535a22e41e32 · outbound

This paper cites Kvquant: Towards 10 million context length llm inference with kv cache quantization.Advances in Neural Information Processing Systems, 37:1270–1303.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Kvquant: Towards 10 million context length llm inference with kv cache quantization.Advances in Neural Information Processing Systems, 37:1270–1303

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.140118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:8ec01d36fcab29a4d39d9f2d764e1479a9b3a5c43623db711042a302a0f61696

Observation 4b0f79cf-da3a-4097-840d-827433bdaddb · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:53:13.598927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:2569a5f1d92d3d0305dc1f78fe5a04f1e4ec2b2b4accb589889d1b294eaa8061

Observation d54f0181-1c18-4a4f-b48f-72eeb58fc5ae · outbound

This paper cites Pqcache: Product quantization-based kvcache for long context llm inference.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Pqcache: Product quantization-based kvcache for long context llm inference

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.114446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:73cf3a633adc13f8696aba3bea101f2d8a650a314a1d205173b92fd3af1b8e36

Observation a8990c60-f268-4918-920b-394a6500e38c · outbound

This paper cites MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:53:13.513638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:a96d2488f98593d100fe9e564b8dca6c482022345823d0470ccd49dc75de4425

Observation 015213c6-84eb-4fbb-bb4f-1b04d4a97998 · outbound

This paper cites SeerAttention-R: Sparse Attention Adaptation for Long Reasoning.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention SeerAttention-R: Sparse Attention Adaptation for Long Reasoning

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:53:13.586873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:a903b4023c61904c476fe0e13740007c6445a1151d2d4f5e66fc1fbcd07acb7e

Observation 22512576-7a2a-4be3-8518-2108e49e5c76 · outbound

This paper cites SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:53:13.535272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:70ac4197ac4fb3f4185c368465520c55625cda7c22dba769e5d8dc4193a4347b

Observation 8c463151-981c-4626-9282-0d893fa8f0bc · outbound

This paper cites Flash sparse attention: An alternative efficient implementation of native sparse attention kernel.arXiv e-prints, pages arXiv–2508.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Flash sparse attention: An alternative efficient implementation of native sparse attention kernel.arXiv e-prints, pages arXiv–2508

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.099196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:f25c093a4a0c6ffaad6af33c722eb0be5745657bcc5987271d55009cb19b1009

Observation 523dda38-9d8c-4fb4-bf91-921b201947fe · outbound

This paper cites Hsa: Head-wise sparse attention for efficient and accurate long-context inference.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Hsa: Head-wise sparse attention for efficient and accurate long-context inference

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.065707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:064da0e0fa573272d93b9cdd427043a23428337003719b09dcd87e0588528549

Observation 8bc6fa7f-847b-4318-b89a-b5337e39fa49 · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:53:13.538082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:4c0e0dd2748b738b816fc7f2bf6444190e156a90bcc74d766ab9431a3f4cc17a

Observation 28691fa5-f1df-4e3a-a74a-f9d6191c0ed2 · outbound

This paper cites Deepseek-v4: Towards highly efficient million-token context in- telligence.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Deepseek-v4: Towards highly efficient million-token context in- telligence

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.063949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:da73fb0d0f6f202a47fca2722c7ec84ff288635d00eb56307114f435914df229

Observation a06cb1cf-cf5e-4d2e-ac87-937f1b96d7e7 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:53:13.562151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:ed3b098ee0931b150bbfaf2e2b5a1a738c4d75cef51bcc73d4a4dd99c471c98d

Observation 6840e049-0ae7-4fbc-92b4-1063b144f783 · outbound

This paper cites Spargeattention2: Trainable sparse attention via hybrid top-k+ top-p masking and distillation fine-tuning.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Spargeattention2: Trainable sparse attention via hybrid top-k+ top-p masking and distillation fine-tuning

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:53:13.510469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:547bf574c63341efb99305cb3e1debf386667f5395beec9d865cdb4d2196ca40

Observation b802fea6-bcc7-4764-a155-64c75f823508 · outbound

This paper cites Double-p: Hierarchical top-p sparse attention for long-context llms.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Double-p: Hierarchical top-p sparse attention for long-context llms

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:53:13.526500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:3fb12303936027ab33fafb5d24f129a4e7cfae5594dd92a6114ff597e29c117d

Observation b387206a-3673-41b1-9161-664b978492c4 · outbound

This paper cites Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:53:13.523245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:710789425fec362753a58ea758cca415409b48baacc0fd9417d4a138b093c073

Observation 5e5eac3a-8f1c-4e36-af15-f69d8cda4fb0 · outbound

This paper cites Possible generalization of boltzmann-gibbs statistics.Journal of statistical physics, 52(1):479–487.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Possible generalization of boltzmann-gibbs statistics.Journal of statistical physics, 52(1):479–487

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.102971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:8ff727c339328f72318dbee7324ff7430f4beade5051e2754de76ef58db8d4e8

Observation 52831df3-351c-45eb-bef1-1cf126c01724 · outbound

This paper cites MiniCPM: Unveiling the potential of small language models with scalable training strategies.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention MiniCPM: Unveiling the potential of small language models with scalable training strategies

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.077431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:34dd4cb6c5ef8e9f98d88165f3783dd8de2a73100b261538de8bdc10301307b1

Observation 8909391e-6eb6-4e71-bdd6-9f125bc24e1b · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:53:13.516609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:d6dad5ea541a330eea02d38655f220236c6a75379904e4cd69e38abceb2f8976

Observation 15a95ccd-30b8-4cc5-a967-a66f05443e93 · outbound

This paper cites Olmes: A standard for language model evaluations.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Olmes: A standard for language model evaluations

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.110754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:0d085d2144a70d87106929920961d81a2ef4fc911844e0b44b60d50851749676

Observation 0ae6d7d9-f2eb-44a5-81e2-f604dafe161f · outbound

This paper cites The language model evaluation harness, 07 2024.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention The language model evaluation harness, 07 2024

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.101138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:cc06a5ad027f4f0735eea0d97531a679121b53bddc18e0fc2b47eb621d1223e1

Observation 1a315cbe-b00b-4ae8-90bf-c80d01e79514 · outbound

This paper cites Hardware-aligned hierarchical sparse attention for efficient long-term memory access.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Hardware-aligned hierarchical sparse attention for efficient long-term memory access

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:53:13.544336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:13d96c4d1155322d9339a7d644770e241f144f48afa00abf2ef4a6709116251e

Observation 63a076b0-92e8-4a4b-ae1c-c3535cfb464d · outbound

This paper cites Every to- ken counts: Generalizing 16m ultra-long context in large language models.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Every to- ken counts: Generalizing 16m ultra-long context in large language models

Reference 76

Resolution
malformed identifier
arxiv_id, observed 2026-05-20T10:53:13.532290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:77077713e9018c2752500f4fd649d1d2f9c9e7c5c578fef44bee820b47c45309

Observation a1a735b0-6b28-4105-9879-e05470164930 · outbound

This paper cites lim n→∞ H aggrsoftmax z(1),z (2),· · ·,z (H);θ logn = 1.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention lim n→∞ H aggrsoftmax z(1),z (2),· · ·,z (H);θ logn = 1

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.060183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:15e59f0a9fb9653d3a4d3dc75f6cea7ee7804180f724db117bb10451979d991e

Observation 92d0096f-f18d-4542-a5a9-224759a06778 · outbound

This paper cites Proof.We first prove that softmax head aggregation is dispersive.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Proof.We first prove that softmax head aggregation is dispersive

Reference 78

Resolution
malformed identifier
raw_fallback, observed 2026-05-20T10:53:26.066170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:742a6e8323e6a554290f62dac30183c97d8ee13c2d9c9ea25bc935b9de3ad421

Pith citing papers

Observation db6737fa-c185-44b1-a272-8815764a89af · inbound

Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling cites this paper.

Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-12T05:40:56.557646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:40:56.557646Z digest=sha256:1d4480cf9d3f82cb5271b898286fcada6b148745896977ef97bf9dd79bdbf36d