Pith. sign in

Paper Citation Record · LEDGER

FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 35 inbound Pith citation observations for arXiv:2502.20766.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.20766 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:42:35.360123Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4e55e330-e88d-4e8e-bf2a-f6e8cf2c29fa · inbound

AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity cites this paper.

AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:53.833323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:53.833323Z digest=sha256:8cea8a1988a7f130adea2b720d75ad8caa383c6faf6258261693932cc0eba3e9

Observation 5bfbc421-98c1-467d-a47e-8c7bdcfe314a · inbound

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers cites this paper.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:15.864790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:15.864790Z digest=sha256:daf72da4715e90b946b55d2963a7f9fb4596a6d27686ffccfd32a99f52aabb64

Observation 94ff5f04-9d57-4434-b600-a547ef219e7c · inbound

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning cites this paper.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.635500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.635500Z digest=sha256:f69f6b75b0f2f1dbaf31805b1fa3a6713ea42d6ac0ab125bfcd6f5440b9d9c2a

Observation 17a278f5-ba27-41f6-b918-962f6bd530f4 · inbound

Lag-Relative Sparse Attention In Long Context Training cites this paper.

Lag-Relative Sparse Attention In Long Context Training FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:10:46.279283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:10:46.279283Z digest=sha256:d964fc5bbb6849b52c3194e35e04202c29bc82419b77786c1f7f3a08cf5e6f6d

Observation da99d513-7753-4c4a-9aec-2bf4bf27ee7d · inbound

Sparse Fine-Tuning of Transformers for Generative Tasks cites this paper.

Sparse Fine-Tuning of Transformers for Generative Tasks FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:30:14.855194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:30:14.855194Z digest=sha256:6461b831a01c8c6d01e602550015c9b69f3393f8955727f3f4f09fe5841406d5

Observation d4be6df4-9618-4fd4-9dac-81eac982d980 · inbound

DeltaLLM: A Training-Free Framework Exploiting Temporal Sparsity for Efficient Edge LLM Inference cites this paper.

DeltaLLM: A Training-Free Framework Exploiting Temporal Sparsity for Efficient Edge LLM Inference FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:17:44.573805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:17:44.573805Z digest=sha256:af81838940b2268c96b83c0b2ac59faccc39ce7161f83052a70fadcfa7889828

Observation fe522093-4b58-4108-9d8d-c7eacb0b20b3 · inbound

Accelerating Prefilling via Decoding-time Contribution Sparsity cites this paper.

Accelerating Prefilling via Decoding-time Contribution Sparsity FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:06:59.904965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T03:05:34.843274Z digest=sha256:c6580c23f62bd3cca8718c42dd23fa791fc1886ff51dad047406b7b8fea60b52

Observation b099f32b-71b1-4685-95ba-b605fdd08fd1 · inbound

ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference cites this paper.

ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:06:52.158913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T22:03:10.316005Z digest=sha256:41f7a540c579d2e1ee35b7e994b8b67bbb9d66a9b428470d2cd3be821379e343

Observation 041c5d06-2132-4dee-89d7-aaed139c5a31 · inbound

UltraImageGen: Efficient Ultra-High-Resolution Image Generation with Hierarchical Local Attention cites this paper.

UltraImageGen: Efficient Ultra-High-Resolution Image Generation with Hierarchical Local Attention FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T09:20:06.613749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:20:06.613749Z digest=sha256:d74b561d8ba454f5765130ac431617912e263c0fbab63e126ff940c11bb77c08

Observation 1efe3e05-b733-407c-9421-0bd925bc76f0 · inbound

BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding cites this paper.

BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:21:18.631790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T22:20:53.856657Z digest=sha256:3397600828dc1776319cbe389d6bd9bd1e03275b29ac3bca50e51d60fb8f085d

Observation a8ba8449-c17d-4dc4-b904-6eca1fe3294d · inbound

ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs cites this paper.

ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T03:37:50.122140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:37:50.122140Z digest=sha256:bff42501b9db51bb86ac2a64abf658141ca5dafd90cd6032273e09fcf696178a

Observation 2ab1a122-93ed-4304-8588-665627252c4a · inbound

Prism: Spectral-Aware Block-Sparse Attention cites this paper.

Prism: Spectral-Aware Block-Sparse Attention FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T03:22:54.033487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:22:54.033487Z digest=sha256:96e3dcca81cce7eb8b4a0664ddcd6bcd13ec3d8313296329b873551a54baf33c

Observation 4097eb8f-ba2f-4c77-b234-586b7e982609 · inbound

RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference cites this paper.

RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:00:17.890303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T20:59:33.902420Z digest=sha256:5845adb534facc252bc187696d18f12d67a190475e5d8b03d5a0014d16db217f

Observation 7d49031b-1479-480a-9a3c-491dd2765964 · inbound

RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference cites this paper.

RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-21T12:50:09.505344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T12:45:27.150368Z digest=sha256:fa8c97febc7b2e501f794bf8688f8377edfc45d53c15fe444d6e7a66b37202da

Observation a7b82960-56e5-44b9-8d0e-709c3268313c · inbound

RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference cites this paper.

RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T22:05:43.018334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:05:43.018334Z digest=sha256:06c7c31d5fbdcb98d713b2e988a0a5a96e80c8bff9c37852ad099040998aa8de

Observation 40daa6a6-231c-4877-b494-10baa294ce7a · inbound

Stem: Rethinking Causal Information Flow in Sparse Attention cites this paper.

Stem: Rethinking Causal Information Flow in Sparse Attention FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T02:39:28.798730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:39:28.798730Z digest=sha256:e1c04d965ab242e0b55b0ca009379b3442a4c8944fdc04953751bba54cc9e44c

Observation 70435810-0bee-42cc-8d50-8281bf98cce0 · inbound

Salca: A Sparsity-Aware Hardware Accelerator for Efficient Long-Context Attention Decoding cites this paper.

Salca: A Sparsity-Aware Hardware Accelerator for Efficient Long-Context Attention Decoding FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:11:18.913894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T17:56:39.124969Z digest=sha256:7dd70c23baf3efe50e581b6c1642599939113e405c0d616f6532c87055dafbf9

Observation 956182f1-e1bb-406d-b559-7b64a421254f · inbound

Sparse Attention as a Range Searching Problem: Towards an Inference-Efficient Index for KV Cache cites this paper.

Sparse Attention as a Range Searching Problem: Towards an Inference-Efficient Index for KV Cache FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:15:50.421216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T01:15:35.871863Z digest=sha256:f40dad90fcf9ce80268f77d9f36fbb779c967ef2ccd7011422e89fca02391cd7

Observation 96c16d51-7e17-4785-b372-5c58b88a422f · inbound

An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference cites this paper.

An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:05:54.249978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T02:56:28.828593Z digest=sha256:97d4ee39aa0981100f007463a1ff65f674fef2a1b57d10a0036abbbca2c0bf11

Observation 5f47c392-359e-47ba-872e-b937bd6e455c · inbound

ReST-KV: Robust KV Cache Eviction with Layer-wise Output Reconstruction and Spatial-Temporal Smoothing cites this paper.

ReST-KV: Robust KV Cache Eviction with Layer-wise Output Reconstruction and Spatial-Temporal Smoothing FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:26:30.101043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T02:52:33.076123Z digest=sha256:6553f3aa37cad668957d8e081f3528fa8d639abe3d0f2071862ba44f149a57da

Observation 691d2615-4b55-4092-9f23-e629ebced8cf · inbound

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection cites this paper.

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-19T21:22:47.999879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T21:19:31.263068Z digest=sha256:83733b3a109af368220f991304d843e89d6dfefcf97a537f3d6ef4b3de411aa7

Observation 9df60470-ba16-4bd9-9fa3-769b49d555f1 · inbound

KVDrive: A Holistic Multi-Tier KV Cache Management System for Long-Context LLM Inference cites this paper.

KVDrive: A Holistic Multi-Tier KV Cache Management System for Long-Context LLM Inference FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:18:13.988453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T11:13:46.098095Z digest=sha256:2e1afff761829854415079cca7f13d8beac2dde31e3ec915209bbccfce11db49

Observation 45c82e97-281a-475d-9fd9-670ff0b15d49 · inbound

SSV: Sparse Speculative Verification for Efficient LLM Inference cites this paper.

SSV: Sparse Speculative Verification for Efficient LLM Inference FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:14:45.989128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T07:14:04.899063Z digest=sha256:7ce282ee3a4bab300f1611797ef0171d46a5072b8830776e3a1fc7d432b70175

Observation d00311f8-5d3e-4a4d-9470-d7e43812f682 · inbound

PulseCol: Periodically Refreshed Column-Sparse Attention for Accelerating Diffusion Language Models cites this paper.

PulseCol: Periodically Refreshed Column-Sparse Attention for Accelerating Diffusion Language Models FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:09:38.711780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T05:05:55.767705Z digest=sha256:c54b97d3ae3a8ff0b8f45f5c668deff309bb6b0e988ae4b48081e25874865594

Observation c47c37df-1e68-4dfd-8707-18e5f507d54d · inbound

DFSAttn: Dynamic Fine-grained Sparse Attention for Efficient Video Generation cites this paper.

DFSAttn: Dynamic Fine-grained Sparse Attention for Efficient Video Generation FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:45:20.108521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T04:44:19.628926Z digest=sha256:e169e8b303ea693ac78092fd52b060d90632211565745d358edcc5e2b5a0f8f0

Observation 132123ff-4f1b-4575-9f2d-4375c6e787aa · inbound

SIFT: Selective-Index For Fast Compute of RAG Prefill by Exploiting Attention Invariance cites this paper.

SIFT: Selective-Index For Fast Compute of RAG Prefill by Exploiting Attention Invariance FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:37:31.238420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T16:24:31.109508Z digest=sha256:3ea447ecd52ec6bf8d993a9b35c3d103cadf34bd22cad2c36b8a9e1913255d33

Observation 52f26710-85f6-4606-9141-17a1b7facdf3 · inbound

Prefilling-dLLM: Predictive Prefilling for Long-Context Inference in Diffusion Language Models cites this paper.

Prefilling-dLLM: Predictive Prefilling for Long-Context Inference in Diffusion Language Models FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:57:38.696659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T13:30:19.620692Z digest=sha256:256c06cc8d28dca9840503223a6f1746ebcce932e0d4247b96c99fa9899dadf9

Observation 34acb13f-32eb-42b8-ac10-087a2646c71b · inbound

LEDGER: Scaling Agentic Document Editing with Dependency-aware Graph Retrieval cites this paper.

LEDGER: Scaling Agentic Document Editing with Dependency-aware Graph Retrieval FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T10:34:36.227719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T10:32:09.554608Z digest=sha256:f9dab4786bf0008bce2fe7275b68b561963651d87ebbf9ac0eddad4d2f5ca255

Observation d930d33b-b247-4e1f-8498-963cfe56059b · inbound

SAF3R: Dynamic Sparse Attention for Feed-Forward 3D Reconstruction Transformers cites this paper.

SAF3R: Dynamic Sparse Attention for Feed-Forward 3D Reconstruction Transformers FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-12T01:12:13.747675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:12:13.747675Z digest=sha256:4939b3b4080f6adb51ef10bd2fcea88de412109c7339286f17ce0f5f71009b1e

Observation 04cc3de6-669c-4141-a719-37c79c982594 · inbound

Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers cites this paper.

Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T13:25:54.951687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:25:54.951687Z digest=sha256:92686388da1a35873c089085860080ccf87e29d86c4b5fef2aa2395d0c316b72

Observation a2f80cf7-bab2-4401-ada4-f300e1b68339 · inbound

Dropping the Anchor: Statistical Context Summarization for Distributed Systems via Pulsar Attention cites this paper.

Dropping the Anchor: Statistical Context Summarization for Distributed Systems via Pulsar Attention FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T14:02:40.838142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:02:40.838142Z digest=sha256:536a69ab9f89440af54631e274d049f48dceaaf3a901de5c80d522ed34610268

Observation 615534be-7122-4418-8206-4f466bc122a5 · inbound

PIVOT: Efficient Query-Group Indexing for Token-Level Sparse Attention cites this paper.

PIVOT: Efficient Query-Group Indexing for Token-Level Sparse Attention FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-31T11:04:04.526969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T11:04:04.526969Z digest=sha256:ef2f3a6d02a3118b6c4be1537676d5dbd96a1d24323140097dc9f5be93ae436a

Observation ba7fe839-b193-4471-a8cc-6705ae71df6c · inbound

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention cites this paper.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:46.301880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:46.301880Z digest=sha256:117705d4c2a0f58ceaece5225f6196f685a5c5df8499e4ddbe49dc253e65587e

Observation 0380e261-edaa-46d6-8086-6632893ec8f6 · inbound

SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inference cites this paper.

SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inference FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T20:50:03.984247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:50:03.984247Z digest=sha256:cb816767590d7c1a37c7daefed19ec85e939a7e5dbc0a9080987a54b0b1bf04b

Observation 6083fe37-9c5f-40e6-97e3-d61f110a605a · inbound

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding cites this paper.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.360123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.360123Z digest=sha256:559530ed0f90152667031557955373d0f5ed1486c09327139196717be0289e2b