Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2312.04985.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T13:46:11.419712Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 9892386f-9661-4b1d-80ec-97f60e102f1a · inbound
Top-Theta Attention: Sparsifying Transformers by Compensated Thresholding SparQ Attention: Bandwidth-Efficient LLM Inference
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 870ee56a-545b-4041-89b2-ff0acca5654d · inbound
Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification SparQ Attention: Bandwidth-Efficient LLM Inference
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 733b74c5-244a-427c-924e-d71070f9fb19 · inbound
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM SparQ Attention: Bandwidth-Efficient LLM Inference
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8094898b-a09f-4271-ac6e-1e47a32bf967 · inbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling SparQ Attention: Bandwidth-Efficient LLM Inference
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9d52bb4-15be-4955-bc9d-3d23740d9950 · inbound
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference SparQ Attention: Bandwidth-Efficient LLM Inference
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40aab145-4630-4637-9fd4-00bad92266b0 · inbound
Cartridges: Lightweight and general-purpose long context representations via self-study SparQ Attention: Bandwidth-Efficient LLM Inference
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e514596-25c8-486d-96ff-73b94b3258e6 · inbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding SparQ Attention: Bandwidth-Efficient LLM Inference
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb719406-d1f1-4827-914d-bb50068c3795 · inbound
ReasonCache: Accelerating Large Reasoning Model Serving through KV Cache Sharing SparQ Attention: Bandwidth-Efficient LLM Inference
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f0bd9784-91ec-47c6-9db7-33bf7abcd744 · inbound
AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs SparQ Attention: Bandwidth-Efficient LLM Inference
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5967116-cafb-43bb-bcb4-87eacc2ed466 · inbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services SparQ Attention: Bandwidth-Efficient LLM Inference
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 16dac94c-7bab-4c2c-b5a8-46d3266dc0df · inbound
Salca: A Sparsity-Aware Hardware Accelerator for Efficient Long-Context Attention Decoding SparQ Attention: Bandwidth-Efficient LLM Inference
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 72938cef-bb62-4e1a-8dbd-1dd05c411287 · inbound
Sparse Attention as a Range Searching Problem: Towards an Inference-Efficient Index for KV Cache SparQ Attention: Bandwidth-Efficient LLM Inference
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2cebaacf-f2ed-48c5-9729-2ec4ee0c4f7d · inbound
An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference SparQ Attention: Bandwidth-Efficient LLM Inference
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5cfa8952-9a26-4efd-85b7-4615c3927039 · inbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving SparQ Attention: Bandwidth-Efficient LLM Inference
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.