Pith. sign in

Paper Citation Record · LEDGER

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels

As of 21 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 1 inbound Pith citation observation for arXiv:2606.07713.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.07713 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T22:32:56.444150Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T13:08:09.010538Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact7
  • verified fuzzy0
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3cce840d-df48-4bb0-a082-325f5c2d61b8 · outbound

This paper cites FATHOM: Fast attention through optimizing memory.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels FATHOM: Fast attention through optimizing memory

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:ae769124fdfcd6406ffeade379aa0f5b5a2f9452bc239d91684caf9593cefba4

Observation a35b2311-1aa3-49ac-a113-eb678af4a77a · outbound

This paper cites MOSA: Matrix optimized self-attention hardware accelerator for mobile device.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels MOSA: Matrix optimized self-attention hardware accelerator for mobile device

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:d90a437e7ba7ff6798c2bf71bb3d1c38f877c67537f1b90bc778323a9b8e79cb

Observation 2ff700ae-f76d-4b9c-b841-57e88d6cbaff · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-02T16:37:09.301147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:50b2a59d468b052dfd8f652d505f3330f36d199cce72cb5ba27589ddfd211b06

Observation 9e6bbb37-4863-4ed9-a3e4-4f72e56f162b · outbound

This paper cites FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-02T16:37:09.295708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:68feef2c405282ea1c08e4ece0a901d4c96cf583ae1a5c2e5a47ea854dcf04cb

Observation 6605622a-6c48-4fd0-b93a-e5e1ab789318 · outbound

This paper cites Hardware considerations for tensor implementation and analysis using the field programmable gate array.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels Hardware considerations for tensor implementation and analysis using the field programmable gate array

Reference 5

Resolution
verified exact
doi, observed 2026-06-27T22:41:23.898840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:5ae2fb9ad613bd2fb252a3ad1f028ef4dbfd9edd540d7d34b45f46a5cf1e5c29

Observation bc413ab4-7cbb-4e9f-86cf-7049052967c2 · outbound

This paper cites Realizing mathematics of arrays operations as custom architecture hardware-software co-design solutions.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels Realizing mathematics of arrays operations as custom architecture hardware-software co-design solutions

Reference 6

Resolution
verified exact
doi, observed 2026-06-27T22:41:23.900855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:9516b3e2d1d0bdbd1d7c1e86f7b48382e4fdf993cc376081399fde21b236712e

Observation 16216584-1b2d-4b85-8a07-c0fa99c83335 · outbound

This paper cites Processing in memory for mathematics of arrays operations.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels Processing in memory for mathematics of arrays operations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:1dd236f62b95fd4e49cdcef85a54fc40a3d611864b2c9e6c5a4b978020811787

Observation 892184b8-f20f-4656-958a-3a9e40760422 · outbound

This paper cites A Fast Optimization View: Reformulating Single Layer Attention in LLM Based on Tensor and SVM Trick, and Solving It in Matrix Multiplication Time.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels A Fast Optimization View: Reformulating Single Layer Attention in LLM Based on Tensor and SVM Trick, and Solving It in Matrix Multiplication Time

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:37:09.298326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:0c34426210900a787188ffde978c1f611e3528acc667a1a3990df216f4a41e31

Observation 1a516b0c-d999-4a5d-a791-1a5bb6b91bc9 · outbound

This paper cites an unresolved cited work.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:28a94eb22d8e53f99dfcbd919c97041a3edb2a435eaeb3aa299948e6645c5b48

Observation d1055bb5-dbff-4801-914b-ad13cf3a9eb9 · outbound

This paper cites A ^3 : Accelerating attention mechanisms in neural networks with approximation.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels A ^3 : Accelerating attention mechanisms in neural networks with approximation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:dfa5f073681ab55cdc1ccf93f362c19e79687d423abd423c57b27b51a547fb67

Observation 318cae26-df65-49f6-b0e8-cce56e9e9e3a · outbound

This paper cites Acceleration of fully connected layers on FPGA using the Strassen matrix multiplication.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels Acceleration of fully connected layers on FPGA using the Strassen matrix multiplication

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:e15de2a86ccd2cba1428cf2c415b52353feeb4fae9a3a3e2241fc3bd7bd63916

Observation 1b2f47cc-08a3-4c42-acf9-80b3ac870984 · outbound

This paper cites Design and Implementation of an FPGA-Based Hardware Accelerator for Transformer.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels Design and Implementation of an FPGA-Based Hardware Accelerator for Transformer

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:37:09.293187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:6e97ca3e6ca43dfc897e7b764ad5a059f78ba2d8e819ac11d0304ef104ffd331

Observation f7372ef7-b175-4c53-93e5-2bf4e08abd07 · outbound

This paper cites Research on matrix multiplication optimization and deployment method for heterogeneous platforms.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels Research on matrix multiplication optimization and deployment method for heterogeneous platforms

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:1b20da4fe3bfe6c4cea3405a25c2ef4dabf88f4d7e91999947755b6b3bdfe17f

Observation 2a5736dd-15d8-49d6-86c5-9d5b4fb35d16 · outbound

This paper cites Array access and performance regarding numerical algorithms.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels Array access and performance regarding numerical algorithms

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:0a8aafa60337992008a7326da0af77d57c82e8ff1e98100a1f5f766997a842b6

Observation e41d9201-4449-4f98-9682-a16cd1455ac3 · outbound

This paper cites A Mathematics of Arrays.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels A Mathematics of Arrays

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:cd5e6462e55dff578e8664468bd4d1afb13bca7bbe52e5a80579a06804a5ab27

Observation de69877b-f1c1-409c-9965-6f0457587a40 · outbound

This paper cites From array algebra to energy efficiency on GPUs: Data and hardware shapes with dimension-lifting to optimize memory-processor layouts.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels From array algebra to energy efficiency on GPUs: Data and hardware shapes with dimension-lifting to optimize memory-processor layouts

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:37:09.280982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:d2ca8c0e14d382692d205974210d7203561c23acccbc2805e98a1eec8949d2f2

Observation dc789cdc-e395-4774-b16e-ecbd46091da6 · outbound

This paper cites Towards automatic, predictable and high-performance parallel code generation.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels Towards automatic, predictable and high-performance parallel code generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:0fb5bd6f136a05be5ad75cece8f10595644d7d35c1b424855f329d29afe752d7

Observation 82f8b2be-9d72-4e49-927f-43f4e20e1efe · outbound

This paper cites New mathematics for computer performance: array algebra and cost functions.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels New mathematics for computer performance: array algebra and cost functions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:8a9af7891d7203de56c8969d5991d4897f7e059882148b930a07c9a6079c67f9

Observation 288b6404-5124-4727-8712-cc421e9ab0a0 · outbound

This paper cites OPTIMUS: Optimized matrix multiplication structure for transformer neural network accelerator.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels OPTIMUS: Optimized matrix multiplication structure for transformer neural network accelerator

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:f4c621c16ffc9048211883440b566981067626aa222b73039d87d3926610da0c

Observation 411e3766-809e-4e95-b3ab-8e7a5647c639 · outbound

This paper cites FACT: FFN-attention co-optimized transformer architecture with eager correlation prediction.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels FACT: FFN-attention co-optimized transformer architecture with eager correlation prediction

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:fa2ae61295b2759ce37bcc3496794abff642b9fed29a4de2095da66ff42c04fa

Observation ebbeb976-a8bf-4c79-a53a-404cc2581dbf · outbound

This paper cites High-performance Gemmini-based matrix multiplication accelerator for deep learning workloads.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels High-performance Gemmini-based matrix multiplication accelerator for deep learning workloads

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:8935808d5f6a610fa5f201c5417cd29e15d075d328a0e2942d0a075d42b3f6bb

Observation b09d5781-1cb5-4a9b-9b76-5dc65048c819 · outbound

This paper cites Design and implementation of a BRAM-banked double-buffered matrix multiplication accelerator for transformer models on edge FPGAs.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels Design and implementation of a BRAM-banked double-buffered matrix multiplication accelerator for transformer models on edge FPGAs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:9b30b6db4cecf1c0487dc678fb336f6b81d82c0a0b70d458216322c1bad857d9

Observation b706a879-9be6-464d-a411-9e548f1f4de6 · outbound

This paper cites Improving the performance of DGEMM with MoA and cache-blocking.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels Improving the performance of DGEMM with MoA and cache-blocking

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:cc38f9e82f40e586d604d40ab8551275524aff86bc6d3b44c7727535ccb4dc03

Observation 2cedba33-dad1-4d9a-9f92-2e71ebd5d115 · outbound

This paper cites Threaded multicore GEMM with MoA and cache-blocking.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels Threaded multicore GEMM with MoA and cache-blocking

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:b3bf52d913f4a6ff7b5746ffaab90f0080036d8874142ce1f0ee393c6fd6752e

Observation 01cf1fa2-2842-4610-afb9-c30081b9588a · outbound

This paper cites Attention is all you need.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels Attention is all you need

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:d41c0125a0eeb71494f1100006bdc21b13948271be27170b512a9c6d595b666d

Observation 357f9832-6c53-4e8a-ad71-fc104d91e53f · outbound

This paper cites Hardware friendly transformer optimization with dynamic attention matrix fusion.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels Hardware friendly transformer optimization with dynamic attention matrix fusion

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:f1a317ccd7cda255644c8cb0e8e9d81f705b651a9f23bb33fbde6087c2c2a5c4

Pith citing papers

Observation 487c3671-2490-403b-b5b6-ecbcad1095be · inbound

MoA-Structured Decode Attention DNF Derivation, KV-Cache Accumulation, GQA/MQA, and OpenACC Kernel cites this paper.

MoA-Structured Decode Attention DNF Derivation, KV-Cache Accumulation, GQA/MQA, and OpenACC Kernel Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T13:08:09.010538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:08:09.010538Z digest=sha256:ddc0c68dc5af2b705df0630cd65cd96726dc197e43477db3e132185c16dc7f85