Pith. sign in

Paper Citation Record · LEDGER

Rectified Sparse Attention

As of 19 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 4 inbound Pith citation observations for arXiv:2506.04108.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04108 v2

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:52:48.987113Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:06:32.693476Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T13:46:58.816966Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 195f9a5b-f13b-4520-a5d3-b22ff14f473a · outbound

This paper cites SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills.

Rectified Sparse Attention SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.200672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.200672Z digest=sha256:0a68d12fbca91b58997d8229fca2b591d73c11b3000a927a9a5b833e080ec63f

Observation bbacd52a-ac00-421f-b26f-472ea2bbf55b · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Rectified Sparse Attention GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.285123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.285123Z digest=sha256:bcc083756b37a6f03c7f78594c2b3a0a54d145e0dd7362faf19491ca6f2dee83

Observation 16245e3d-3417-44ff-8c73-564e3e612f0b · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

Rectified Sparse Attention Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.349290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.349290Z digest=sha256:644eaa1306492d1d246e647d7d12aedfa86f2b47172ceafc886a54fee2245f66

Observation 6d71bcdd-d3e0-46c7-98e1-921e2eea93a6 · outbound

This paper cites MagicPIG: LSH Sampling for Efficient LLM Generation.

Rectified Sparse Attention MagicPIG: LSH Sampling for Efficient LLM Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.465707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.465707Z digest=sha256:9e398ceed20e90c026bfb8d0e0ac3783bdd0356f4bf6c0c53b1548cd110cc54d

Observation 6ca1098b-5ddc-474f-a14e-ce7f8b4b2dfa · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Rectified Sparse Attention Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.527304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.527304Z digest=sha256:a4dcbf1b3826ef232219b50392c65c3adab952274628c1ffb8cc71548e59b86a

Observation ed759aec-6814-40db-a56d-41151166bdf9 · outbound

This paper cites Flash-Decoding for long-context inference.https://crfm.stanford.edu/2023/10/12/flashdecoding.html, 2023.

Rectified Sparse Attention Flash-Decoding for long-context inference.https://crfm.stanford.edu/2023/10/12/flashdecoding.html, 2023

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:52:49.268981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:52:48.550761Z digest=sha256:1d6f9f7556fd7e3a59ff18f55b80633765f957bde46989f8be8089244dd4968f

Observation 22abf7b4-91e0-48e1-b5b4-8b0ae3a4152c · outbound

This paper cites MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models.

Rectified Sparse Attention MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.650576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.650576Z digest=sha256:e0b371611edb27c61f4243c366c7f7b1b50698a1b43064b75300fa672c6d1c3e

Observation e892edf0-045c-41f9-81e0-73dd74f82b7f · outbound

This paper cites SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs.

Rectified Sparse Attention SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.764869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.764869Z digest=sha256:e76e22bc5fc867720257acc559ce5f0f3d180abc85bdeed42a40583748005ef7

Observation 99217db7-9ac4-4fd3-91b0-9f5e57ab2e5c · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Rectified Sparse Attention DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.868706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.868706Z digest=sha256:626b93842dd6ff9221250f87852ef9df3961e41a7c532403c95c7f507173563b

Observation ee8639a1-dc19-4f1b-8861-83e1711016ae · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Rectified Sparse Attention OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.901428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.901428Z digest=sha256:de1ab00748c53f1edf67514e3c80eeeb098ff814a944743f28c942d24bce3087

Observation b8c5d9b4-69ab-498a-9051-24b27898a446 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

Rectified Sparse Attention Measuring mathematical problem solving with the math dataset

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.935265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.935265Z digest=sha256:b2f423ffe48f734814f5d3a098c4d47b65d67fd5cfe1465ca19af431c1966552

Observation 17c4b2d3-7d80-4f7e-a321-478c29739989 · outbound

This paper cites DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference.

Rectified Sparse Attention DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.938385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.938385Z digest=sha256:df160c1f77f64f9526365c703707b32777469633a7fa33a5e6e1e777d64fa7df

Observation 12ea0f1f-3189-47b6-864f-d1ea47ef7f87 · outbound

This paper cites OpenAI o1 System Card.

Rectified Sparse Attention OpenAI o1 System Card

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.941415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.941415Z digest=sha256:6cbd6652059c4d62025b7e5e70b99577587f8e90e3cba6301b58a7fcb954e536

Observation 35a5dac5-0e4d-46d9-b136-1aecd9c73860 · outbound

This paper cites Fast inference from transformers via speculative decoding.

Rectified Sparse Attention Fast inference from transformers via speculative decoding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.944325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.944325Z digest=sha256:dfcc58a853b7b8c33a80328c681709239d4e1ef8ef3ceb704209ec4200ee3cc5

Observation be024e36-8e0f-44cc-909d-dd3b550312c7 · outbound

This paper cites Solving quantitative reasoning problems with language models.Advances in Neural Information Processing Systems, 35:3843–3857, 2022.

Rectified Sparse Attention Solving quantitative reasoning problems with language models.Advances in Neural Information Processing Systems, 35:3843–3857, 2022

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.947196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.947196Z digest=sha256:8057caa155bed40139cbebfd4fbcdaa8bbcacb29039be8fb4442308fe723f256

Observation 11fa7331-4345-4ad5-ac3d-8fe859d827c4 · outbound

This paper cites EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty.

Rectified Sparse Attention EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.949963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.949963Z digest=sha256:6065137e99cf6f9c846d282ee898080c699d90ce7c9347dbf2e46ab52cbf01cb

Observation 6ec6e455-732c-44d0-b336-17cb5fc7c4e6 · outbound

This paper cites MARIO: MAth Reasoning with code Interpreter Output -- A Reproducible Pipeline.

Rectified Sparse Attention MARIO: MAth Reasoning with code Interpreter Output -- A Reproducible Pipeline

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.952928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.952928Z digest=sha256:b91f011a4c335075c098a5a26bd5af8fa329cd056d3739a8d23ff41cf32359b6

Observation 89b1cfc7-6b65-4532-a222-6d99636ed20f · outbound

This paper cites ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression.

Rectified Sparse Attention ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.956989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.956989Z digest=sha256:fea976981af9667fa1a65dd458ab95071e6da8ca097c5c7b2d39fc71cb2ae58d

Observation 131d3691-37ae-43d6-81b2-a7d6ccfa2d99 · outbound

This paper cites MoBA: Mixture of Block Attention for Long-Context LLMs.

Rectified Sparse Attention MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.960314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.960314Z digest=sha256:57df9bf4788d206c226ba21839b6b3cc5c85bdd01c66f1cfe683b3f2507f047c

Observation ff56fa26-426a-4638-807d-425e491602f4 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Rectified Sparse Attention Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.963403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.963403Z digest=sha256:327fcff0b26fb5d7d72d38fa034683b9304a448d3b74289a3ae25484abb1eb57

Observation e9161153-b987-4d15-aee1-1ea702c8f66c · outbound

This paper cites MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding.

Rectified Sparse Attention MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.966152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.966152Z digest=sha256:49c3a071739b217f4bae38d007f001d6462d93ddc90775aac56e5ba401b45891

Observation ee44efbc-ac4b-45df-b458-0d636cf0ee21 · outbound

This paper cites TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding.

Rectified Sparse Attention TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.969030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.969030Z digest=sha256:1614a68ab71da51df41ecb09e8ca3629adf8a5d3393ab8527b81c722b1a3c895

Observation 4e5d68fd-4e9c-4eae-8d71-39bf16189b3d · outbound

This paper cites Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference.

Rectified Sparse Attention Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.972268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.972268Z digest=sha256:965bf40563db96b7c671bcf05f47b0ffcdf91225f09a687ba5c889944f77ac1c

Observation 1a8181b3-5ed7-4a97-975b-447725659275 · outbound

This paper cites InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory.

Rectified Sparse Attention InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.975137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.975137Z digest=sha256:cb2de56b6d4ae7a736634dec38536b9e63ea6316f142b2434fb3a1739b3568cf

Observation 524b4230-df1a-4c66-a691-39c2f3aa839b · outbound

This paper cites Qwen2.5 Technical Report.

Rectified Sparse Attention Qwen2.5 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.978094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.978094Z digest=sha256:d2e2a3667aa4cb8bc507df87f97e13bff24131dcfdf0c66ae061b090a19f6ff0

Observation afb1e160-cf1b-40f4-8c89-1d923ac34cf3 · outbound

This paper cites Qwen2.5-1M Technical Report.

Rectified Sparse Attention Qwen2.5-1M Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.981099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.981099Z digest=sha256:effbc04e34163235a87eae23fc10182b6d5cae9c69bf2af24d094222613ae3a7

Observation e9a5af27-b368-4d42-817a-6675be5981b5 · outbound

This paper cites Orca: A distributed serving system for Transformer-based generative models.

Rectified Sparse Attention Orca: A distributed serving system for Transformer-based generative models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:52:49.239761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:52:48.984029Z digest=sha256:82c4d334bfdf7a0e783421af2df39c292b9e5d4a42ff9a23eea513f18d5e24a0

Observation cb718aeb-3183-4d09-8cca-503448e3a0ba · outbound

This paper cites Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.

Rectified Sparse Attention Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.987113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.987113Z digest=sha256:9e5c32ff7c6e48bfe4ff2501953ff104bd2ba3e1e46ed9a729b6254322b09e1e

Pith citing papers

Observation 973f0b67-f390-49b2-b41c-433cfb715649 · inbound

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning cites this paper.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Rectified Sparse Attention

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.693476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.693476Z digest=sha256:9837264e287a24c7132872bd44602f9d8a828e953d89211a53b9ad749ba19686

Observation 3fd29474-9a0f-4b7a-89b8-11bf887c5093 · inbound

Flashlight: PyTorch Compiler Extensions to Accelerate Attention Variants cites this paper.

Flashlight: PyTorch Compiler Extensions to Accelerate Attention Variants Rectified Sparse Attention

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:41:30.042725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-22T11:40:41.762364Z digest=sha256:5bb5b0208dc87c2701e7fc1e92bb21e0600e522a86df55ab114fac1441f229ef

Observation cf2a2230-520c-4f33-b5b1-5d835d2570f9 · inbound

BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding cites this paper.

BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding Rectified Sparse Attention

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:21:18.626267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T22:20:53.856657Z digest=sha256:c6666bbc8a0d5f6ef2c98154f12f363e7c482967515fbd7b7151f6d7bb9c9942

Observation 5361cd55-4d42-455c-8d9c-b0a2d1a68d7f · inbound

You Only Index Once: Cross-Layer Sparse Attention with Shared Routing cites this paper.

You Only Index Once: Cross-Layer Sparse Attention with Shared Routing Rectified Sparse Attention

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:46:58.818402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:06:04.896501Z digest=sha256:6dbcc8ef52aaf4814fa5692c09365bf258322f32c58109adc8e85c5ceb1dbe7d