Pith. sign in

Paper Citation Record · LEDGER

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning

As of 15 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 12 inbound Pith citation observations for arXiv:2506.08889.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08889 v1

Coverage vector

measured 78 of 78 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:06:32.755775Z

measured 90 of 90 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:11:48.900029Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T13:36:59.523914Z

Reference resolution

78 of 78 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved64
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 603767e6-da74-41e3-9625-00bbc598aeab · outbound

This paper cites URLhttps://github.com/tile-ai/tilelang.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning URLhttps://github.com/tile-ai/tilelang

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:06:33.702936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:06:32.546420Z digest=sha256:9a00c9d1cfc42e60c7e54c70dbcc133f302144a53afe5a42940dd1ee5ed2ae8a

Observation 3fca4ecb-e378-4f01-908d-ef2b18764eae · outbound

This paper cites Keyformer: Kv cache reduction through key tokens selection for efficient generative inference.Proceedings of Machine Learning and Systems, 6:114–127, 2024.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Keyformer: Kv cache reduction through key tokens selection for efficient generative inference.Proceedings of Machine Learning and Systems, 6:114–127, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:06:33.644684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:06:32.550117Z digest=sha256:df8f013ecd786f4459fd6da0cf9daf0d44d47283134cc226bc9a004c49c17281

Observation 7f5078a8-1d1e-4198-979f-7049585bdb4c · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.553125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.553125Z digest=sha256:131466731bd9132c739b56ef47ec76a1d54d59ea83fb804293e7fbbd26da379c

Observation 283fb409-0322-4c71-9453-dd78c3e8b1cb · outbound

This paper cites xLSTM: Extended Long Short-Term Memory.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning xLSTM: Extended Long Short-Term Memory

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.556320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.556320Z digest=sha256:d94c78b701785b597ddd2bea32a62280408a79be1d4342c3a5e2df3f85bae475

Observation 09743327-ca1d-4c10-b8f3-6af12b746a81 · outbound

This paper cites RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.559683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.559683Z digest=sha256:691235d426e73f3ae14cfc0d5f9c921aed4cf5a0d051f54da947850abf5b6da7

Observation 6000f946-9a71-4274-9b21-340fc650a4f6 · outbound

This paper cites Longformer: The Long-Document Transformer.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Longformer: The Long-Document Transformer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.563190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.563190Z digest=sha256:7388d6f3c8ab35c1382f9f8c000a1f839b37083235546cae874a159ebc1b3d09

Observation 0181f93e-8425-4ed1-a1f6-06a70775a97f · outbound

This paper cites Reducing transformer key-value cache size with cross-layer attention.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Reducing transformer key-value cache size with cross-layer attention

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:06:33.609507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:06:32.566599Z digest=sha256:409ae3abcb4e75b28d655a93f4a6dd2e6c1dfbc3940bd5ecf735db6d2e063b20

Observation 4815cf23-38f5-4b5a-ba4e-31639b59f255 · outbound

This paper cites R-kv: Redundancy-aware kv cache compression for training-free reasoning models acceleration.arXiv preprint arXiv:2505.24133, 2025.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning R-kv: Redundancy-aware kv cache compression for training-free reasoning models acceleration.arXiv preprint arXiv:2505.24133, 2025

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.569593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.569593Z digest=sha256:498665a352948fdfcd0f431dde2601bf454189a90c45428a9b5b882781ff04e8

Observation 8f496848-d855-4471-8142-2c892c297b6b · outbound

This paper cites SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.572234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.572234Z digest=sha256:7d0d4a76550d428eaa999db9f7b838766156ea326e0290b7338fd5db0b39bf1a

Observation 44a057aa-bae1-493a-bc57-03ec9dab8d18 · outbound

This paper cites RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.575162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.575162Z digest=sha256:e830d7a5dab9c5b7b1edc13bdac257d53dea46963d134536eeb1418a971da6d1

Observation bfb725fa-9e3e-45ef-80a5-30ddc1b6c7da · outbound

This paper cites MagicPIG: LSH Sampling for Efficient LLM Generation.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning MagicPIG: LSH Sampling for Efficient LLM Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.577962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.577962Z digest=sha256:224c3f972dd884eaca76cde9e572d2114a028cb2e233877f7eba47f785d80dfd

Observation 2bf6d40a-4219-422e-b3fc-64ba19b1d892 · outbound

This paper cites PipeThreader: Software-defined pipelining for efficient dnn execution.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning PipeThreader: Software-defined pipelining for efficient dnn execution

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:06:33.570272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:06:32.580743Z digest=sha256:eebd640ec77dddc1de2fdd9a7312126db3449bede939aaf5ba7f4b71763e3806

Observation 3abc9a08-dfa0-4a4e-b361-d1f14374e254 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Generating Long Sequences with Sparse Transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.583149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.583149Z digest=sha256:5bd608295d82d7c4c8008bec19acbe608dafbef14eeb081c63b4890e4fe3e434

Observation 8673676c-8c53-4f7c-a57a-84f2dca653bf · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.585847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.585847Z digest=sha256:01fb422801d83f073f5830c6f0235f67bc726396d6796cfdad6b38b24b3f4af6

Observation 97b6ae02-56b1-4b19-9cf1-daf8bc52f13f · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.588405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.588405Z digest=sha256:a86d1fe9accacd190e86379d114c4ae3b713b7dac148fecff7869fec23ae049c

Observation 5d536d28-56ee-48fd-98f7-a858af3525ea · outbound

This paper cites Hymba: A Hybrid-head Architecture for Small Language Models.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Hymba: A Hybrid-head Architecture for Small Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.591124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.591124Z digest=sha256:d3d049286d92efa4c85918ea4d7a694448e7d67873a2a2431df6be0fd323a508

Observation e4e97183-6acd-47af-9892-7ea5a68e7b7b · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.594106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.594106Z digest=sha256:42011ecc97e27c8c8317ea7fd5dd2f85e734d889e3a2fede144b7b180523f6e7

Observation ec087002-3424-487c-802c-c6f20a8d8573 · outbound

This paper cites Moa: Mixture of sparse attention for automatic large language model compression.arXiv preprint arXiv:2406.14909, 2024.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Moa: Mixture of sparse attention for automatic large language model compression.arXiv preprint arXiv:2406.14909, 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.596786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.596786Z digest=sha256:296cee2b6aeb0c73b6bef257ccc686e66acb045dde134b7f481441307f877c98

Observation 59d96b03-d758-480d-8dea-c50af189e0f7 · outbound

This paper cites SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.599244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.599244Z digest=sha256:0c131743919b40e6cc77b01aa70a4a466d1c5fb210f8e549922ea53cf55ee4ac

Observation 6574c471-194b-4c87-9736-87f3e24d4f67 · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.602023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.602023Z digest=sha256:ece5d187abd5d0445f2fd400ce4c2a9048fb642f980353df1b0880ad4eb97d7f

Observation f3700472-5d13-4c40-a724-ad7110675c01 · outbound

This paper cites Better & Faster Large Language Models via Multi-token Prediction.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Better & Faster Large Language Models via Multi-token Prediction

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.604608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.604608Z digest=sha256:5be883e36357c529f810362bbb8c816d40fa6ed8316f81a49e59c9e0a674e589

Observation 3d1573ba-6df3-4109-8a22-a7c0cb30fbfb · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.607376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.607376Z digest=sha256:74b1331edbeb64ebb6db780af37dac3ea34973cb1f3c11f9a4bf8a1fc72d96c0

Observation 9fafa75e-6103-4783-8a4f-1d254ddb0115 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.610150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.610150Z digest=sha256:eeb2c7f2d67890298164f5967da54467a8f71dd3b7a12480d5b98abce5d2ae56

Observation f224880f-aba3-45d4-b743-5c6e89313233 · outbound

This paper cites Omnikv: Dynamic context selection for efficient long-context llms.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Omnikv: Dynamic context selection for efficient long-context llms

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:06:33.538895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:06:32.612606Z digest=sha256:c073d7ff1eff5e69e26f70254ba1b5a70b6a975debb999a78fef466a69aa3a9a

Observation 26d5b8c7-b6ca-4bdb-b8bb-248d491e1837 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Measuring Massive Multitask Language Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.615350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.615350Z digest=sha256:a568e78aa9d9555dddd9c7f01fbcf859feca5758b34aa08869d49768ceb92f01

Observation a125b793-69f1-4931-b3c5-c781c13f3bc4 · outbound

This paper cites Squeezed attention: Accelerating long context length llm inference.arXiv preprint arXiv:2411.09688, 2024.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Squeezed attention: Accelerating long context length llm inference.arXiv preprint arXiv:2411.09688, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.618100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.618100Z digest=sha256:0b4b4754c15c8abb6ad2044859316c735c6dcc8f763972f29cf2b7d633f504de

Observation e30bed75-7a91-4ce3-afde-a4304d5cc93a · outbound

This paper cites RaaS: Reasoning-Aware Attention Sparsity for Efficient LLM Reasoning.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning RaaS: Reasoning-Aware Attention Sparsity for Efficient LLM Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.620481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.620481Z digest=sha256:9cf24622604f827b72758c412bbce8e8800554da01bd52fff3f2729f684ba8cd

Observation f630c6ae-6e46-4b5c-bb03-cf706b18b457 · outbound

This paper cites OpenAI o1 System Card.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning OpenAI o1 System Card

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.622907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.622907Z digest=sha256:5e5e66192c8de44ba7ea0e3a8197190b5561e0adf76f23ff29e3136a0f454b72

Observation 4999cafd-4f82-4a86-8e13-90bd51e734a5 · outbound

This paper cites MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.625484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.625484Z digest=sha256:bac0d7dce2427e0ed222531454985bb567a9b97f4c1d55c58c9014c9a2127d57

Observation e755f89d-3c32-44d1-89a3-42874f1b507c · outbound

This paper cites Kullback-leibler divergence.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Kullback-leibler divergence

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:06:33.518944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:06:32.628140Z digest=sha256:c6c3c50d5fa130a0f759afe0d6961633916cd4883ee38803b8192fb6f3eb1e04

Observation 6ebb34e9-7d68-4134-a99f-d5fbb211f785 · outbound

This paper cites Transformers are rnns: Fast autoregressive transformers with linear attention.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Transformers are rnns: Fast autoregressive transformers with linear attention

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.630511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.630511Z digest=sha256:8e680e6299ede4006903b83bc1bf039f6fb92494d65799dba4f28b8a6df745a4

Observation bcbff325-2586-45a4-86c8-cc1488914b75 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Gonzalez, Hao Zhang, and Ion Stoica

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.632901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.632901Z digest=sha256:7c80e69befb70d123b1952828ae5e013e9cc8071859c5df9da6d3429ed5e9d20

Observation 94ff5f04-9d57-4434-b600-a547ef219e7c · outbound

This paper cites FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.635500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.635500Z digest=sha256:bb9a803365b0ffc16389b25be4f5a7d8f22e1fd8ea2e07e7706b8ce0a9278c38

Observation 7879778f-56b8-4abc-885c-8df83eb59889 · outbound

This paper cites Fast inference from transformers via speculative decoding.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Fast inference from transformers via speculative decoding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.637982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.637982Z digest=sha256:c061aada85ec7651dd6e088d497a173f30414d55e4d38b1b31a11e2a21c89cff

Observation 67473948-fa99-448e-bea7-ef575a77dd23 · outbound

This paper cites MiniMax-01: Scaling Foundation Models with Lightning Attention.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning MiniMax-01: Scaling Foundation Models with Lightning Attention

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.640407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.640407Z digest=sha256:1137bf5eabfcf488064d5ab323e807a7fdeff041894208b57ef079482b550fb2

Observation 1d7ae517-d409-4357-b5c4-6d3b18cdd479 · outbound

This paper cites SCBench: A KV Cache-Centric Analysis of Long-Context Methods.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning SCBench: A KV Cache-Centric Analysis of Long-Context Methods

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.642988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.642988Z digest=sha256:11d709254fb712f1ba9c5793886606ffe66f9917b81a0edbaa853d8faa3fd5fa

Observation 03f08edf-0805-4bd4-9d62-5ae7f5e8cf61 · outbound

This paper cites Snapkv: Llm knows what you are looking for before generation.Advances in Neural Information Processing Systems, 37:22947–22970, 2024.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Snapkv: Llm knows what you are looking for before generation.Advances in Neural Information Processing Systems, 37:22947–22970, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.645647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.645647Z digest=sha256:45d258634a74146edf7bee0ee67c7c8effaa45b555915c8515bb85fd326b4aee

Observation f3838f42-b8e1-41e0-8dfb-6e1323801559 · outbound

This paper cites Twilight: Adaptive attention sparsity with hierarchical top- p pruning.arXiv preprint arXiv:2502.02770, 2025.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Twilight: Adaptive attention sparsity with hierarchical top- p pruning.arXiv preprint arXiv:2502.02770, 2025

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.648079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.648079Z digest=sha256:f8b71bf231e4e3208c23129121a5691fe79280c1b71a040e1f4013e0f6805e19

Observation babc598e-3810-4f5a-bff9-881b61898fe8 · outbound

This paper cites Adaptive Computation Pruning for the Forgetting Transformer.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Adaptive Computation Pruning for the Forgetting Transformer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.650507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.650507Z digest=sha256:58a10549527bf10fb10e29eac83627337fc7366556376420e25d1490900edd54

Observation 0ff6959f-9958-4507-a9ec-1f287dd68ea6 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.652878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.652878Z digest=sha256:cc2c68143ff30ae7f43118e7fbe5cf516bf5fd72faa1f23b545e4ce30992a67f

Observation c219a013-8686-4648-8283-db95ac9584b7 · outbound

This paper cites DeepSeek-V3 Technical Report.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning DeepSeek-V3 Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.655660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.655660Z digest=sha256:223d2c25985c5bdebd310a8ee29d6673fa161f53e33ec36dfee716d6a6d455de

Observation 2608c643-fffb-422b-877e-691614308d6c · outbound

This paper cites RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.658257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.658257Z digest=sha256:d16bfbfd0707f0140071fbdc53244eccd9ef02cbcad09453ead82af0e6a928ce

Observation 3530e94f-8a37-4537-b556-73ba454253da · outbound

This paper cites Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.660824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.660824Z digest=sha256:69d0084655c99ea1143034fdcf1fbb6df14d9611938d9eaf616534e132f39f2a

Observation 5e68eb24-e125-495b-beec-fbcb87c254a1 · outbound

This paper cites an unresolved cited work.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.663343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.663343Z digest=sha256:9026c110c99afbc672af2fd902f757fda99a9d78c1f6f42cc260f44dab46e71d

Observation 293d79b4-1d15-40dc-a4f6-305781561c89 · outbound

This paper cites MoBA: Mixture of Block Attention for Long-Context LLMs.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.665648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.665648Z digest=sha256:e44628d2bde3219a48ef539e58ecb4552434aad1533512267569d194cd9c418d

Observation 7a362d9b-f617-41b7-bbf1-6fb73bd032fe · outbound

This paper cites Inference-time sparse attention with asymmetric indexing.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Inference-time sparse attention with asymmetric indexing

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.671170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.671170Z digest=sha256:33016f76ba9111f4bb78eaf754caeb61c7a9aa4e03877dbc465a5ca039000b74

Observation d667d25b-3933-414d-a5b6-143357e4e96c · outbound

This paper cites Aime problems and solutions.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Aime problems and solutions

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:06:33.472215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:06:32.673524Z digest=sha256:80845429f0d86f587fed9f0267e74cf4512142a3cea050b7c65824a5e80bfffb

Observation c01b0e2c-6de0-4f45-ad94-953af6858da4 · outbound

This paper cites RWKV: Reinventing RNNs for the Transformer Era.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning RWKV: Reinventing RNNs for the Transformer Era

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.676223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.676223Z digest=sha256:56cddc247a94c3e6d612f5c04b9cfc8260c697607de9d06b9d9ee77c4b2c517c

Observation 0cf432c8-7479-46db-b6a4-144757bb9505 · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Gpqa: A graduate-level google-proof q&a benchmark

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.678846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.678846Z digest=sha256:6e95b93e0eb791875aa8ce7c80e52a7bc1984902de2a9d4607a20366aaec20d2

Observation d49d1105-3f23-496a-83b3-036244b4a626 · outbound

This paper cites Flashattention-3: Fast and accurate attention with asynchrony and low-precision.Advances in Neural Information Processing Systems, 37:68658–68685, 2024.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Flashattention-3: Fast and accurate attention with asynchrony and low-precision.Advances in Neural Information Processing Systems, 37:68658–68685, 2024

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.681295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.681295Z digest=sha256:422640b27964a41906c0958120b9362c629aad5dbf396eeeb200a84d56138e79

Observation 8b4bec04-2885-409e-a38b-b7d3854429c1 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Fast Transformer Decoding: One Write-Head is All You Need

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.683732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.683732Z digest=sha256:8747e965e2f79696c8d94226e6f6fe1f6304ab5777a7aad584fa5eb1c5305b40

Observation d984e696-1486-4856-95d9-d4a857901aa0 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.686218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.686218Z digest=sha256:3c02c505f409efd79df8fc800e0f8e8955b62ff7b7b0359a74cec653f3156200

Observation e4e05d7f-6813-44c6-813c-7b8878ce31b7 · outbound

This paper cites Retentive Network: A Successor to Transformer for Large Language Models.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Retentive Network: A Successor to Transformer for Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.688555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.688555Z digest=sha256:d6fa279733fb21cdc2fe2da44c703e5f81cf37fbbb5b0b436ee7d1983e265129

Observation fd83167e-1222-400d-8695-d27c9f5de1de · outbound

This paper cites You only cache once: Decoder-decoder architectures for language models.Advances in Neural Information Processing Systems, 37:7339–7361, 2024.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning You only cache once: Decoder-decoder architectures for language models.Advances in Neural Information Processing Systems, 37:7339–7361, 2024

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.691080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.691080Z digest=sha256:af773676a9fbfe6146b66552916aea588627d999de92f4157e7ec19042a7a545

Observation 973f0b67-f390-49b2-b41c-433cfb715649 · outbound

This paper cites Rectified Sparse Attention.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Rectified Sparse Attention

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.693476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.693476Z digest=sha256:d594daacb81956a12df723c2e8f02d20cf25d67e77b7bf51f7e8c90d53a5620b

Observation e7e5477f-7825-4fd7-9640-a0efa6e1f46c · outbound

This paper cites Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.696425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.696425Z digest=sha256:e26251ee66b5268dd9205722383035cbf68829e35317a1650cc3cd9d849229c1

Observation a67b2b37-dab9-4ff1-a390-54ec09637381 · outbound

This paper cites Minicpm4: Ultra-efficient llms on end devices.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Minicpm4: Ultra-efficient llms on end devices

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:06:33.432664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:06:32.699380Z digest=sha256:239175fc07ac28f4314aeb2e5992ce278fcf374dd719b65f109c21159acf24f9

Observation cd3ca030-3cf0-40aa-a4f0-706877f84057 · outbound

This paper cites Triton: an intermediate language and compiler for tiled neural network computations.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Triton: an intermediate language and compiler for tiled neural network computations

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.701827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.701827Z digest=sha256:5d2abfc4d8de8d5bb59ff90600bb342c573912644c8a798670aff2d78689a4a8

Observation 6ac5c415-4962-4918-ba04-6d5b0b62b380 · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:06:33.411160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:06:32.704664Z digest=sha256:0c6a8f0255d6afac4f1c1a1b95d02fd3ee98087b3cbdffdb232d1dc865ce5dea

Observation a4ba9d5f-4361-4aed-88d1-e2f654aeaaf8 · outbound

This paper cites Ladder: Enabling efficient low-precision deep learning computing through hardware-aware tensor transformation.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Ladder: Enabling efficient low-precision deep learning computing through hardware-aware tensor transformation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:06:33.401791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:06:32.707584Z digest=sha256:77cc22f88989c830358e3c3d093ebb3b0cc0d79e1b93e6009984c25e0a913c3f

Observation e51bd4e9-f1ae-497b-b21f-cb827e7bdd73 · outbound

This paper cites InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.710225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.710225Z digest=sha256:8bb515479e8015fabda5e27c7ea372e8e0192d92be4631c0e02b7c67efa36bab

Observation 0da04754-dc55-459a-abb4-c029a17bb797 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Efficient Streaming Language Models with Attention Sinks

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.712682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.712682Z digest=sha256:6f64219951bb7d6b4e3c231fd98fd001e6c564496c916cb0b98d74c80f247c56

Observation af72d851-2eef-4c38-93dc-e2ca446dc95f · outbound

This paper cites DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.715316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.715316Z digest=sha256:5188fd992b34d82c2e30338496db3465c240cd9147513acf216a8e0021855da2

Observation c1f77f0b-62cc-436f-a5a8-9fb99f5859c8 · outbound

This paper cites XAttention: Block Sparse Attention with Antidiagonal Scoring.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning XAttention: Block Sparse Attention with Antidiagonal Scoring

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.718359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.718359Z digest=sha256:a17262f2417aa4374be4cca6c89c077da416556b833ee13938d81e4f582c257d

Observation c2b5e48b-48d0-4ef7-8fc6-80bd77976df3 · outbound

This paper cites Qwen3 Technical Report.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Qwen3 Technical Report

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.721183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.721183Z digest=sha256:35ff35171363f79bd47fe0a18e8d463c2f7e576eeb7f93605feca24e0a9df1cf

Observation e6a73f34-695e-4872-ae7d-4196c0ddbddb · outbound

This paper cites LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.724124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.724124Z digest=sha256:0b966ad7126f1a6285c272e7664b5bb4e83b2f799aaa359897b47c154f410563

Observation bab96539-ebe4-47de-8355-181e7446dd8d · outbound

This paper cites Post-Training Sparse Attention with Double Sparsity.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Post-Training Sparse Attention with Double Sparsity

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.726866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.726866Z digest=sha256:e6e5c72430b940247adb2bda84c4a623f6aae2be0d0d6386e0d1b2e5ab8fddc8

Observation 95dfe03b-fd92-4e07-8d3f-2a2c169094e6 · outbound

This paper cites Gated Linear Attention Transformers with Hardware-Efficient Training.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Gated Linear Attention Transformers with Hardware-Efficient Training

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.729535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.729535Z digest=sha256:f34be3dbc06f7dfab8c4836e1cb1f100f90e3ac96065ce20c093b5f8f3c7181f

Observation ba368b4c-83f6-4f2c-9b86-76d1891b91af · outbound

This paper cites Parallelizing Linear Transformers with the Delta Rule over Sequence Length.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.732274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.732274Z digest=sha256:1ad5602b34d2bbf513579cf9e4da9ef3a881d5f45541b1811226fcadf18df9dd

Observation 92fa3dfc-1f8d-4cf3-90db-dd119862ed8b · outbound

This paper cites Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.735020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.735020Z digest=sha256:b9d272b579a557530f12cac499a1ef3178af9f059c504312c0a680edd1bfa3b4

Observation 292cce17-8c9a-4178-887e-7497c8c511e0 · outbound

This paper cites Hardware-Efficient Attention for Fast Decoding.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Hardware-Efficient Attention for Fast Decoding

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.737715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.737715Z digest=sha256:1ab33d79c93a86dd16dcb4d5ea038f3dffc7c21983177689a2b9f64c7073d7c7

Observation 563df4f2-cca6-45f9-a9b7-96a0005e3186 · outbound

This paper cites Big bird: Transformers for longer sequences.Advances in neural information processing systems, 33: 17283–17297, 2020.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Big bird: Transformers for longer sequences.Advances in neural information processing systems, 33: 17283–17297, 2020

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.740450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.740450Z digest=sha256:0d068f65617ddab49139604f633e00d363bc8b559f28358f4bab57409da7774a

Observation ce549a7a-91ea-40a8-b8bd-ec2b2a44d192 · outbound

This paper cites In-context KV-Cache Eviction for LLMs via Attention-Gate.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning In-context KV-Cache Eviction for LLMs via Attention-Gate

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.742963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.742963Z digest=sha256:0d5e1e03738558b38bbbfb778569664d3fdcfa75c7b295731fcacce3e350381a

Observation 506d2d3c-2bcb-4aea-9efb-b6071374512a · outbound

This paper cites PQCache: Product Quantization-based KVCache for Long Context LLM Inference.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning PQCache: Product Quantization-based KVCache for Long Context LLM Inference

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.745585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.745585Z digest=sha256:d2a975abc03b7c901cfd07bf42b89577d0ed995b3eadd364be5a3d3eea049de6

Observation f109a2ac-2869-4b5e-9f1b-2099b952ac7d · outbound

This paper cites Spargeattn: Accurate sparse attention accelerating any model inference.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Spargeattn: Accurate sparse attention accelerating any model inference

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:06:33.387280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:06:32.748548Z digest=sha256:c4c5d9206edf71df6f701d38cc0c7d4e5a0993e1a5b7e6053d4c1e3a1177019a

Observation ab92dd5a-4c2b-42ca-936f-3deb329e8c07 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning H2o: Heavy-hitter oracle for efficient generative inference of large language models

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:06:33.379076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:06:32.751054Z digest=sha256:734c942970c34a041bd5a16a92d1204e48f67ea0606692d7c57ef5b3f9b4f054

Observation 0e4c6dc1-84ad-4eee-897e-4fe0150e1e66 · outbound

This paper cites Sglang: Efficient execution of structured language model programs.Advances in Neural Information Processing Systems, 37:62557–62583, 2024.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Sglang: Efficient execution of structured language model programs.Advances in Neural Information Processing Systems, 37:62557–62583, 2024

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:06:33.371151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:06:32.753405Z digest=sha256:f61b5f2be62768ea7c35b8cff1911a94e08e457dec56cfe74ab8af7ead4efc8b

Observation 53b8b144-c7aa-40e1-9c49-15a0541be264 · outbound

This paper cites ROLLER: Fast and efficient tensor compilation for deep learning.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning ROLLER: Fast and efficient tensor compilation for deep learning

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:06:33.362723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:06:32.755775Z digest=sha256:2b62658298d05b61219db52ee40689e240d0c45762332e9ceb3efcc77a6464fe

Pith citing papers

Observation 9efce4c6-ac26-4d67-b107-bec22e6a9b2e · inbound

DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning cites this paper.

DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning SeerAttention-R: Sparse Attention Adaptation for Long Reasoning

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T07:26:02.860786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T07:25:56.953876Z digest=sha256:66b97910dc51edfcd2bb33bb25802d414ed1244e80f39881046d26a393e8e90d

Observation b4a83254-06b0-4481-9dd6-7cc3af0c7330 · inbound

BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding cites this paper.

BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding SeerAttention-R: Sparse Attention Adaptation for Long Reasoning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:21:18.608942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T22:20:53.856657Z digest=sha256:c8fb053a241f79af1a7071d1d8208283cf11b3e927cbffca42974849f914e8a2

Observation 1e34d13a-9f28-459a-b6b1-af2ff9ee8997 · inbound

Understand and Accelerate Memory Processing Pipeline for Large Language Model Inference cites this paper.

Understand and Accelerate Memory Processing Pipeline for Large Language Model Inference SeerAttention-R: Sparse Attention Adaptation for Long Reasoning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:28:29.860253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-14T00:25:49.807277Z digest=sha256:7bbe38004a7d41454f1ef9ad70bc3460fe6819c3a609b3ceb6224658fe57383a

Observation 3772f3b3-2b63-4da8-a594-fef652386ee2 · inbound

LongAct: Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning cites this paper.

LongAct: Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning SeerAttention-R: Sparse Attention Adaptation for Long Reasoning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:20:10.553722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-10T11:17:43.769244Z digest=sha256:2bf6a282177d4e6fd78ddd914dcc2ad2065a411827ec9afac897753f63f63185

Observation 7e737445-b0fd-423c-8c2c-b47f79990846 · inbound

Unifying Sparse Attention with Hierarchical Memory for Scalable Long-Context LLM Serving cites this paper.

Unifying Sparse Attention with Hierarchical Memory for Scalable Long-Context LLM Serving SeerAttention-R: Sparse Attention Adaptation for Long Reasoning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:01:25.896526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-07T13:15:21.201950Z digest=sha256:b3d85fcbbd5e496b878ddedc039c328ad2b289091120d45fe1f925fe938f8e8b

Observation ea5e97a0-7941-4c73-84d0-4223c643f029 · inbound

An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference cites this paper.

An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference SeerAttention-R: Sparse Attention Adaptation for Long Reasoning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:05:54.140296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T02:56:28.828593Z digest=sha256:41108164aff67b399dca95086fdc919a684d135e20243ca37dfded6fadf82564

Observation 4e375f8d-a619-4f26-8b6a-d8c8893f666e · inbound

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction cites this paper.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction SeerAttention-R: Sparse Attention Adaptation for Long Reasoning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:41:26.703884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:7b6e0843adae3fbc4af6be34650b4249342d00af4838d6172ce81038e8259766

Observation 015213c6-84eb-4fbb-bb4f-1b04d4a97998 · inbound

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention cites this paper.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention SeerAttention-R: Sparse Attention Adaptation for Long Reasoning

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:53:13.586873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:8100a54ac0401f2daa608370f723736d56b08deb945e5e8d699d19874c67692f

Observation 25b2bac7-6670-42dc-a6cd-e008de0a2d35 · inbound

You Only Index Once: Cross-Layer Sparse Attention with Shared Routing cites this paper.

You Only Index Once: Cross-Layer Sparse Attention with Shared Routing SeerAttention-R: Sparse Attention Adaptation for Long Reasoning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:36:59.525347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T01:06:04.896501Z digest=sha256:630bf501a626c9521d81a0e2d6d41968e867b3f6e0ec1eafe73a32d15a35d80c

Observation 5ee18d25-c704-4b04-9543-64c33b112150 · inbound

PIVOT: Efficient Query-Group Indexing for Token-Level Sparse Attention cites this paper.

PIVOT: Efficient Query-Group Indexing for Token-Level Sparse Attention SeerAttention-R: Sparse Attention Adaptation for Long Reasoning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-31T11:04:04.584216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T11:04:04.584216Z digest=sha256:4fa906cbc2ba71f64917caccf6bf45561cad7204ae8615abc20981db93a3b394

Observation dac1432e-871b-4332-b81a-15b0cca4a735 · inbound

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs cites this paper.

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs SeerAttention-R: Sparse Attention Adaptation for Long Reasoning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:53.318889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:43:53.318889Z digest=sha256:d630d5eb672b5b8eaa72940da5de29f6c783d6b2c61c13f7c7fba2a8bd752814

Observation 5ba11f5f-1fb7-44d2-be6d-de0d04b20416 · inbound

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs cites this paper.

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs SeerAttention-R: Sparse Attention Adaptation for Long Reasoning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T00:11:48.900029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:11:48.900029Z digest=sha256:acb1f199ebc690040409627ed2f4cec71e0138e0c378b828031274d13abecdef