Pith. sign in

Paper Citation Record · LEDGER

Self-attention Does Not Need $O(n^2)$ Memory

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2112.05682.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2112.05682 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:01:36.169940Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

18
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b6a711ae-3328-4a71-93e3-8b9e79ecf0ac · inbound

FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness cites this paper.

FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness Self-attention Does Not Need $O(n^2)$ Memory

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T16:22:08.919757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T16:22:08.801066Z digest=sha256:7c83e0c429948c8a14b695f7e7d35eacd6ca8649b5a8838dfd021fbcff760696

Observation 972429b4-e267-4ac3-9c40-c444ea8b69f1 · inbound

A Comprehensive Overview of Large Language Models cites this paper.

A Comprehensive Overview of Large Language Models Self-attention Does Not Need $O(n^2)$ Memory

Reference 128

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:28:39.429937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T20:28:38.900026Z digest=sha256:e372fbf8f37d9fe6cc6e627b9c8a9eb376f80137998a287ff16a1203f0b6cfcd

Observation d752eee2-551f-47db-94aa-fcf3b9ece64c · inbound

FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning cites this paper.

FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning Self-attention Does Not Need $O(n^2)$ Memory

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T02:39:44.880343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T02:39:44.770344Z digest=sha256:2cc19f04a00916aa2747288666739ff88d6353d06c15bc9d5616e245beafe921

Observation 10b33196-1bd4-498c-be96-9d339123afc1 · inbound

Baichuan 2: Open Large-scale Language Models cites this paper.

Baichuan 2: Open Large-scale Language Models Self-attention Does Not Need $O(n^2)$ Memory

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-24T06:54:03.598810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-24T06:51:02.531751Z digest=sha256:c4499e39a0d6775fd23b72dba0d3bab158e06f7ba43017e9570deb0f5e276bbc

Observation 2e9a9a31-4350-484f-8963-3297da5c085e · inbound

Ring Attention with Blockwise Transformers for Near-Infinite Context cites this paper.

Ring Attention with Blockwise Transformers for Near-Infinite Context Self-attention Does Not Need $O(n^2)$ Memory

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T19:28:28.282101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T19:28:28.201789Z digest=sha256:f13a5d6b2a25c138008708c0f84f5c79b55bacb0bbc7ef71ea87673f266063bf

Observation d4475732-bcaf-4502-a2e0-0e72730aac4d · inbound

The Falcon Series of Open Language Models cites this paper.

The Falcon Series of Open Language Models Self-attention Does Not Need $O(n^2)$ Memory

Reference 97

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:46:10.152199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-16T09:46:09.701440Z digest=sha256:f7ad2559296658daa5651051867278317a3d3c52b09569dc1c0c451f5c9475e6

Observation 04abeaa0-441c-474a-a619-144c1b0a10b8 · inbound

FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision cites this paper.

FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision Self-attention Does Not Need $O(n^2)$ Memory

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T19:45:36.409336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T19:45:36.337956Z digest=sha256:261affad90a77f48a7c2341b74ee67a1daf10e6c0c3f6611bf15411b61c9ce98

Observation 81b46da2-31a4-4759-8c9c-ba8f692f9a17 · inbound

Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models cites this paper.

Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models Self-attention Does Not Need $O(n^2)$ Memory

Reference 230

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T06:38:37.006933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T06:38:36.517935Z digest=sha256:b2a38479ed951b53802a071f619c4687f3e4a82b97fac0c8c58ca87af37a9cdd

Observation ea0f0ed2-e6b4-43a2-9b2e-39585cda9e9d · inbound

Transformer Neural Processes - Kernel Regression cites this paper.

Transformer Neural Processes - Kernel Regression Self-attention Does Not Need $O(n^2)$ Memory

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-23T08:37:44.960168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-23T08:36:55.970834Z digest=sha256:df779a1016224fc16f6d9467bef94c5b856852009b44810814f2a6cd707efa35

Observation b8637c49-e944-40d8-90df-d2cfae424361 · inbound

BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching cites this paper.

BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching Self-attention Does Not Need $O(n^2)$ Memory

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-23T16:58:11.950592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T16:57:46.645061Z digest=sha256:b40224bdbe3c9f1102487af131984497ee04f50451b017de8bf286b18b195f9a

Observation cc33bdb7-acf2-469e-94b6-81f3970129cd · inbound

Flex Attention: A Programming Model for Generating Optimized Attention Kernels cites this paper.

Flex Attention: A Programming Model for Generating Optimized Attention Kernels Self-attention Does Not Need $O(n^2)$ Memory

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:27:16.668684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T21:27:16.615071Z digest=sha256:514ba4b634288eb68bdb490890a474e5a885631633f0846743142563abb67c7b

Observation d1acaf38-48d1-4533-b92d-d1ccf2ea4737 · inbound

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection cites this paper.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Self-attention Does Not Need $O(n^2)$ Memory

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.169940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.169940Z digest=sha256:47c00ceca19f71890a6a466c9e799b6152a117e598ab617aa90be592e2b953dd

Observation d1bdae36-d87c-4e8e-bc7c-ff74eab4a9a4 · inbound

FLASH-D: FlashAttention with Hidden Softmax Division cites this paper.

FLASH-D: FlashAttention with Hidden Softmax Division Self-attention Does Not Need $O(n^2)$ Memory

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:43.089177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:43.089177Z digest=sha256:2491f603f4a096d5205b57717666ea88a5c2580edd94bf7729b492023e0f5294

Observation 9f593de3-9fa0-436f-8301-0952ca9795bd · inbound

Low-Cost FlashAttention with Fused Exponential and Multiplication Hardware Operators cites this paper.

Low-Cost FlashAttention with Fused Exponential and Multiplication Hardware Operators Self-attention Does Not Need $O(n^2)$ Memory

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:46.917536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:46.917536Z digest=sha256:fcd4815a3c3cbd310315d25fdef324d82b16875a979471436899306a57be7ce5

Observation 643622ba-c6d2-42da-98a4-a59551229b5a · inbound

TransAct V2: Lifelong User Action Sequence Modeling on Pinterest Recommendation cites this paper.

TransAct V2: Lifelong User Action Sequence Modeling on Pinterest Recommendation Self-attention Does Not Need $O(n^2)$ Memory

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:21.904574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:21.904574Z digest=sha256:18bce3492747412548649a55d1155d9b18330ee4ffbe4efc116dc2fb120e790a

Observation 89ad099d-36a3-4b2d-bd40-0f1fcc7e8357 · inbound

HMAR: Efficient Hierarchical Masked Auto-Regressive Image Generation cites this paper.

HMAR: Efficient Hierarchical Masked Auto-Regressive Image Generation Self-attention Does Not Need $O(n^2)$ Memory

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:47:52.364862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:47:52.364862Z digest=sha256:2ad528d1f8da1fc9766b625b5ad336a59333d681e21defb5c22120770b0f5ffd

Observation 38941a85-0e27-44a0-bfcd-2a070ef9bc96 · inbound

Inter2Former: Dynamic Hybrid Attention for Efficient High-Precision Interactive cites this paper.

Inter2Former: Dynamic Hybrid Attention for Efficient High-Precision Interactive Self-attention Does Not Need $O(n^2)$ Memory

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:43.755282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:43.755282Z digest=sha256:037c0d5127b965aa448dd9962286ff257b832090c0c11d6b11d9cf53733d01b6

Observation 7216fa82-e54e-4411-b98e-efde2995ecbb · inbound

Local Representative Token Guided Merging for Text-to-Image Generation cites this paper.

Local Representative Token Guided Merging for Text-to-Image Generation Self-attention Does Not Need $O(n^2)$ Memory

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:46.317905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:46.317905Z digest=sha256:700f98bc17be4d164391310385ab08c145c69fbc8ceb5b7358499fdac439046c

Observation 638ae11d-b20d-40e9-b01c-805952011d9b · inbound

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers cites this paper.

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers Self-attention Does Not Need $O(n^2)$ Memory

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:11:14.081783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:11:14.081783Z digest=sha256:2ce4b51c0a83d714db3b106a1182a2e42270c81e4d93824b3064c0259a370db0

Observation 99e518d1-e991-443a-a8e0-5e07929bd86f · inbound

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions cites this paper.

Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions Self-attention Does Not Need $O(n^2)$ Memory

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T21:38:54.293572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:38:54.293572Z digest=sha256:2f7d73618542f6d16fb1b871ae40ab685f21bd6473d96c5c73f544a8867ccfa5

Observation 0b3b49d8-a9f9-41e2-bda2-7a53544e8126 · inbound

NEST: Nested Event Stream Transformer for Sequences of Multisets cites this paper.

NEST: Nested Event Stream Transformer for Sequences of Multisets Self-attention Does Not Need $O(n^2)$ Memory

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:20:47.663417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T09:20:04.153122Z digest=sha256:c13b8f412aff5b58e54f250d9f132c7a1b8c973cd5e451fa3439fc2786d50b95

Observation ee937aa5-6b7e-42bf-9640-f141a66bb86c · inbound

Drift-Resilient Temporal Priors for Visual Tracking cites this paper.

Drift-Resilient Temporal Priors for Visual Tracking Self-attention Does Not Need $O(n^2)$ Memory

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T19:53:11.800078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T19:50:03.094100Z digest=sha256:ff3e3f4ae83ca05aac4c222ed0930a5e8bd3d34a12fcb1043423b65d68b0ad61

Observation d4707d6d-ce4b-4c31-b737-880614fbfba8 · inbound

Dispatch-Aware Ragged Attention for Pruned Vision Transformers cites this paper.

Dispatch-Aware Ragged Attention for Pruned Vision Transformers Self-attention Does Not Need $O(n^2)$ Memory

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:35:18.742860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T11:34:19.658008Z digest=sha256:68eddfd3f5bbb00504c0f887083d7f7de1ef0b2251419cfe26966e23b6c6ff2a

Observation 205bb844-086f-434c-8e74-121fbf1e6989 · inbound

Dispatch-Aware Ragged Attention for Pruned Vision Transformers cites this paper.

Dispatch-Aware Ragged Attention for Pruned Vision Transformers Self-attention Does Not Need $O(n^2)$ Memory

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:52:32.334786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T07:48:16.052462Z digest=sha256:608ac4a4c0cc0f96b07b7016854c1ec735a87cc8422141e457fcebbdaa53758c

Observation 923a333d-45d7-41d8-b935-f4c342be3ec1 · inbound

HieraSparse: Hierarchical Semi-Structured Sparse KV Attention cites this paper.

HieraSparse: Hierarchical Semi-Structured Sparse KV Attention Self-attention Does Not Need $O(n^2)$ Memory

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T07:16:54.124178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T07:15:19.184970Z digest=sha256:1e1c14ef04b21d6ac747f49c0916cb611a92702d3add40e43b7b379210108162

Observation 65597e69-dfee-4232-8119-d19a8694aae1 · inbound

The Recurrent Transformer: Greater Effective Depth and Efficient Decoding cites this paper.

The Recurrent Transformer: Greater Effective Depth and Efficient Decoding Self-attention Does Not Need $O(n^2)$ Memory

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:21:04.643546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T22:03:35.260373Z digest=sha256:797676191c13f5e697e3746154db406793830d2246a5378fd405e3b65ecb1c75

Observation 9c77f844-1a35-43dd-864d-5f32381df7d2 · inbound

ELSA: Exact Linear-Scan Attention for Fast and Memory-Light Vision Transformers cites this paper.

ELSA: Exact Linear-Scan Attention for Fast and Memory-Light Vision Transformers Self-attention Does Not Need $O(n^2)$ Memory

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:11:17.606339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T06:33:10.730973Z digest=sha256:78553962ef37e600e5ff14997ce3f1d3e72c4c2fd2cdf8f30c1c814631cfd603

Observation 9d47a2eb-ad46-43c1-a937-23bf0f14a43b · inbound

Context Memorization for Efficient Long Context Generation cites this paper.

Context Memorization for Efficient Long Context Generation Self-attention Does Not Need $O(n^2)$ Memory

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T10:43:12.653023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T10:39:09.720412Z digest=sha256:c548df3f8f64c6c636ae0c060c70e1b00c854c5a8c4520df07a07fd2c11e16fd

Observation a1e99ec1-568c-452f-ac75-a04e0eb6ee8b · inbound

Prefilling-dLLM: Predictive Prefilling for Long-Context Inference in Diffusion Language Models cites this paper.

Prefilling-dLLM: Predictive Prefilling for Long-Context Inference in Diffusion Language Models Self-attention Does Not Need $O(n^2)$ Memory

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:57:38.691461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T13:30:19.620692Z digest=sha256:4f679c47dd59000531638a7a4d141d0f3825f35c67260b0e8a9f5ac9bd633e01

Observation 100734a7-637b-4833-800d-0d27d27b588a · inbound

Design-CP: Context Parallelism for Design of Protein Nanoparticles cites this paper.

Design-CP: Context Parallelism for Design of Protein Nanoparticles Self-attention Does Not Need $O(n^2)$ Memory

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-12T02:29:52.764344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T02:29:52.764344Z digest=sha256:5df4bd1da386df10edee80748eb66679ca28e651385e7d3f1caa3a057485bdbc

Observation 868777fa-406b-4c51-8148-393509166bad · inbound

Intrinsic and Triangulation-Agnostic Attention: A Simple and Powerful Approach for Learning on Meshes cites this paper.

Intrinsic and Triangulation-Agnostic Attention: A Simple and Powerful Approach for Learning on Meshes Self-attention Does Not Need $O(n^2)$ Memory

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-31T05:06:55.085464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:06:55.085464Z digest=sha256:1ddea84a42580a4df524be8bc3e019432bf261ac107197eb230b6e3d412962a5