Pith. sign in

Paper Citation Record · LEDGER

SparQ Attention: Bandwidth-Efficient LLM Inference

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2312.04985.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.04985 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T13:46:11.419712Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9892386f-9661-4b1d-80ec-97f60e102f1a · inbound

Top-Theta Attention: Sparsifying Transformers by Compensated Thresholding cites this paper.

Top-Theta Attention: Sparsifying Transformers by Compensated Thresholding SparQ Attention: Bandwidth-Efficient LLM Inference

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T05:29:16.733334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:29:16.733334Z digest=sha256:001a20ca86764ec01548d4f48b1ab67663d48af4985595f4b078a933384fe0a3

Observation 870ee56a-545b-4041-89b2-ff0acca5654d · inbound

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification cites this paper.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification SparQ Attention: Bandwidth-Efficient LLM Inference

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.419712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.419712Z digest=sha256:75680923282b2933903d5433ddaced8d41c0a9a38223c8087d67527fd0ea50ef

Observation 733b74c5-244a-427c-924e-d71070f9fb19 · inbound

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM cites this paper.

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM SparQ Attention: Bandwidth-Efficient LLM Inference

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-22T17:01:48.528269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T16:58:11.105511Z digest=sha256:9b9df3d57bfa1df3100d1688099328414554ffb32fa854ff1a24f90e23ae2305

Observation 8094898b-a09f-4271-ac6e-1e47a32bf967 · inbound

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling cites this paper.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling SparQ Attention: Bandwidth-Efficient LLM Inference

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:07.482431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:07.482431Z digest=sha256:befb35decaac9c407084e2b1617e701ddeba114dde6a9014e0af618dcc1866a3

Observation c9d52bb4-15be-4955-bc9d-3d23740d9950 · inbound

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference cites this paper.

HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference SparQ Attention: Bandwidth-Efficient LLM Inference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:00.637951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:00.637951Z digest=sha256:48a6ddae57d1725ef7e1a0298489cf26f40ed731ca4d72964b09b8a6513f45e5

Observation 40aab145-4630-4637-9fd4-00bad92266b0 · inbound

Cartridges: Lightweight and general-purpose long context representations via self-study cites this paper.

Cartridges: Lightweight and general-purpose long context representations via self-study SparQ Attention: Bandwidth-Efficient LLM Inference

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:33.981581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:33.981581Z digest=sha256:54941fbdd57f11044d4394d7fbcc2eb734a5ffdcbf872133f327daae0a9745a2

Observation 1e514596-25c8-486d-96ff-73b94b3258e6 · inbound

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding cites this paper.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding SparQ Attention: Bandwidth-Efficient LLM Inference

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.247507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.247507Z digest=sha256:1f0bd73a60992dbc6ad647fd8bc42378fb209235752950baecbe01e7f306b37b

Observation cb719406-d1f1-4827-914d-bb50068c3795 · inbound

ReasonCache: Accelerating Large Reasoning Model Serving through KV Cache Sharing cites this paper.

ReasonCache: Accelerating Large Reasoning Model Serving through KV Cache Sharing SparQ Attention: Bandwidth-Efficient LLM Inference

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:17:00.840258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T03:14:05.109011Z digest=sha256:376c18586077045d5d23ab6c0d07db54913ffbff6b261557434b45b658c26493

Observation f0bd9784-91ec-47c6-9db7-33bf7abcd744 · inbound

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs cites this paper.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs SparQ Attention: Bandwidth-Efficient LLM Inference

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:55.010618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:55.010618Z digest=sha256:861aea6d618852ad121ba737078f00bb56a4016251fd4ae85e04605bd537abcf

Observation a5967116-cafb-43bb-bcb4-87eacc2ed466 · inbound

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services cites this paper.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services SparQ Attention: Bandwidth-Efficient LLM Inference

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:58:35.755161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:4aad34d0f3b741d3aed487f81888403d4f5cc432fb30c427061e2abb3f092052

Observation 16dac94c-7bab-4c2c-b5a8-46d3266dc0df · inbound

Salca: A Sparsity-Aware Hardware Accelerator for Efficient Long-Context Attention Decoding cites this paper.

Salca: A Sparsity-Aware Hardware Accelerator for Efficient Long-Context Attention Decoding SparQ Attention: Bandwidth-Efficient LLM Inference

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:11:18.789798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T17:56:39.124969Z digest=sha256:7bc8dfb7fafae149ef0e77a582d1b3ac1e731a12cb24b2e875b42a1d474975b8

Observation 72938cef-bb62-4e1a-8dbd-1dd05c411287 · inbound

Sparse Attention as a Range Searching Problem: Towards an Inference-Efficient Index for KV Cache cites this paper.

Sparse Attention as a Range Searching Problem: Towards an Inference-Efficient Index for KV Cache SparQ Attention: Bandwidth-Efficient LLM Inference

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:15:50.537412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T01:15:35.871863Z digest=sha256:cff596192505225907ba388c6edb6c164ecf30e44eded17f53568edc73bfd7a2

Observation 2cebaacf-f2ed-48c5-9729-2ec4ee0c4f7d · inbound

An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference cites this paper.

An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference SparQ Attention: Bandwidth-Efficient LLM Inference

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:05:54.081442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T02:56:28.828593Z digest=sha256:d115d75a9d34ec130bf40a26a18a82c96f3d1ec733710c888c3bcc22c2d05918

Observation 5cfa8952-9a26-4efd-85b7-4615c3927039 · inbound

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving cites this paper.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving SparQ Attention: Bandwidth-Efficient LLM Inference

Reference 101

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:6f90ef9f967071e3e96a6035f3ed4b61ea4aa1dcc3d667721580f362aa06be5b