Pith. sign in

Paper Citation Record · LEDGER

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators

As of 21 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2505.18824.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18824 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:29:29.161352Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy22
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6baf57a8-36f6-4226-92a5-69aa9eaf98e4 · outbound

This paper cites On the computational complexity of self-attention,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators On the computational complexity of self-attention,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:34.233199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:29:26.197619Z digest=sha256:b2c5858340fb6d7e23fefd4702a5d7507a9c4989d3af0f706ed823e36b665daf

Observation 2da04287-b6d6-4094-9ce7-7ba8329aa926 · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:26.256211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:26.256211Z digest=sha256:f3c7c652f0c97fa5d4ac6c73a78a1a90b68f33f6f5e8d5b460b4fc29b9e87bf3

Observation b0cfb9b3-f0e9-4eff-95a5-e7f8e75b694e · outbound

This paper cites Data movement is all you need: A case study on optimizing transformers,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators Data movement is all you need: A case study on optimizing transformers,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:34.117193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:29:26.366113Z digest=sha256:3bcd2a0b4bbed5c969debf5cb3435c5428a16de021553d8bf5dad4f35ecbb2e4

Observation b9234b8b-47e6-4fe6-9478-77b618a41694 · outbound

This paper cites FlashAttention: Fast and memory-efficient exact attention with IO-awareness,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators FlashAttention: Fast and memory-efficient exact attention with IO-awareness,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:33.965439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:29:26.504090Z digest=sha256:bf96ff638c3dc213bbcbf39d599b173b4231e05ac7537da337ca8aa8f783cf1a

Observation 4b1303b3-74ce-4003-a6e8-97045b74dd87 · outbound

This paper cites FlashAttention-2: Faster attention with better parallelism and work partitioning,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators FlashAttention-2: Faster attention with better parallelism and work partitioning,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:26.631529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:26.631529Z digest=sha256:526db3b7e40a559f267d816499dede8a154cd89350d8f66bb950e618a887de8e

Observation ff115aa5-8c1a-4601-a0ef-4a94f2508275 · outbound

This paper cites FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:26.731577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:26.731577Z digest=sha256:abc23bdf8d55deae3719e2b4f5c7886a50f910019e44ddf2b4457f9065b3df4f

Observation ead9f0f6-6a3c-4d01-8439-511de6f2d9fe · outbound

This paper cites SambaNova SN40L: Scaling the AI memory wall with dataflow and composition of experts,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators SambaNova SN40L: Scaling the AI memory wall with dataflow and composition of experts,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:33.841153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:29:26.876797Z digest=sha256:3375e564dbe07ba7a422b01348341c9ff63cc53ba7e2ddd868b63b80ac8bfbd6

Observation 555d3de4-b705-4e0b-9c89-cd24aeb2e39d · outbound

This paper cites Wafer-scale AI: GPU impossible performance,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators Wafer-scale AI: GPU impossible performance,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:33.637534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:29:27.012739Z digest=sha256:1a80b98ebf6487653e620ed15bd963aa198ca277d06513872971d5b54ada0cba

Observation 3ea76046-4229-47bc-9ca7-5f44509bda34 · outbound

This paper cites Blackhole & TT-Metalium: The standalone AI computer and its programming model,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators Blackhole & TT-Metalium: The standalone AI computer and its programming model,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:33.468948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:29:27.140730Z digest=sha256:ae6a039e4bdcda7ce9ea963a571015ee967cdd8e92aa603f36b66fe87095ac2c

Observation f5bedae7-44b5-429c-85de-efbe08a69ee3 · outbound

This paper cites 16.2 rngd: A 5nm tensor-contraction processor for power-efficient inference on large language models,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators 16.2 rngd: A 5nm tensor-contraction processor for power-efficient inference on large language models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:33.303301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:29:27.260624Z digest=sha256:785a76f4c1cb9f9356abe98709fed20c4c20c2f5a7f9e1ab55613d0556c76cfe

Observation 47282922-b877-4fad-93a0-853e5b5f66bb · outbound

This paper cites Attention in SRAM on Tenstorrent Grayskull.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators Attention in SRAM on Tenstorrent Grayskull

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:27.394330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:27.394330Z digest=sha256:0a9797eeb0557ac486ddb88afc4d66642ef2d9210c5b5260f6a54b74a53eb84a

Observation a08011de-6c7d-4a4d-834e-bc71041f9341 · outbound

This paper cites FLAT: An optimized dataflow for mitigating attention bottlenecks,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators FLAT: An optimized dataflow for mitigating attention bottlenecks,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:33.164993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:29:27.554302Z digest=sha256:45066924ab8893703409cfa8f9376b7cdb62d1df45db40938d2c0bf93a2925b2

Observation 2bff695c-0266-4cd1-ad94-120f6114f8ee · outbound

This paper cites Fusemax: Leveraging extended einsums to optimize attention accelerator design,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators Fusemax: Leveraging extended einsums to optimize attention accelerator design,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:33.034879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:29:27.667462Z digest=sha256:48026296b8b4748341e7d98b1f31715c2c127936bbba4e0ec857f6ae3c4b29f6

Observation bb749651-27e1-4ec3-ba8e-74bbc933b1ae · outbound

This paper cites Gemini: Mapping and architecture co-exploration for large- scale DNN chiplet accelerators,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators Gemini: Mapping and architecture co-exploration for large- scale DNN chiplet accelerators,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:32.881056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:29:27.784853Z digest=sha256:f81b4a9ac54eb15d715a5726493844bc6c6f5665004283089a20f4d30f41022b

Observation b3bf5695-480a-476c-99a9-b3a58ff851e5 · outbound

This paper cites DOJO: The microarchitecture of Tesla’s exa-scale computer,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators DOJO: The microarchitecture of Tesla’s exa-scale computer,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:32.650726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:29:27.883144Z digest=sha256:0f7f85b8ae1bc3925fc590d823d2c6b23fd1bf25125b4d1a88c7d50f1e3e7d4a

Observation 05faa99a-a993-4ae1-af69-9b1534aa2b64 · outbound

This paper cites Collective communication,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators Collective communication,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:32.384160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:29:28.022282Z digest=sha256:d7636d0ee05d560eec47acc2d868a62afee9f3634727ef126e0073625762f5ec

Observation 0a933052-8573-4252-86f6-55e8dd9aedf4 · outbound

This paper cites Towards the ideal on-chip fabric for 1-to-many and many-to-1 communication,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators Towards the ideal on-chip fabric for 1-to-many and many-to-1 communication,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:32.102965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:29:28.127464Z digest=sha256:885667609ab6663e2655e9f7b1e9a90e2e6d0111c58301e98622bc005f390248

Observation decf9542-b5a8-411a-8e29-25f720ce5dcc · outbound

This paper cites GVSoC: a highly configurable, fast and accurate full- platform simulator for RISC-V based IoT processors,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators GVSoC: a highly configurable, fast and accurate full- platform simulator for RISC-V based IoT processors,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:31.742199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:29:28.250138Z digest=sha256:38f17fc141f8bee0a8ca7c284516dbbb3af6c497e0eebcecde6e6232d29cb4a0

Observation 8d73a607-c1f1-4753-be67-0d6e73a8d4f1 · outbound

This paper cites Snitch: A tiny pseudo dual-issue processor for area and energy efficient execution of floating-point intensive workloads,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators Snitch: A tiny pseudo dual-issue processor for area and energy efficient execution of floating-point intensive workloads,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:31.429126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:29:28.357031Z digest=sha256:ce775373bc66e0728c75ac2060e18eb1c06388dc95a3f0a71c0c7c7d15b27286

Observation 7c482cb8-b5c9-422f-a249-418a79543af7 · outbound

This paper cites Spatz: Clustering compact RISC-V-based vector units to maximize computing efficiency,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators Spatz: Clustering compact RISC-V-based vector units to maximize computing efficiency,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:31.190544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:29:28.485665Z digest=sha256:eb29e18351b54217412a2d5f88f16d1219a7f0d5570b5d2dfa5f14ab92c953c7

Observation c91be4fd-2578-4fca-9f12-9b05c8bff576 · outbound

This paper cites A high-performance, energy-efficient modular DMA engine architecture,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators A high-performance, energy-efficient modular DMA engine architecture,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:30.874902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:29:28.604764Z digest=sha256:5243cbec1cdc88b75a5c7fdc9597c79e730343190f596dd19398412e458b828d

Observation 1c013230-326d-4b0a-b056-7555a1d7a0ba · outbound

This paper cites RedMule: A mixed-precision matrix–matrix oper- ation engine for flexible and energy-efficient on-chip linear algebra and TinyML training acceleration,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators RedMule: A mixed-precision matrix–matrix oper- ation engine for flexible and energy-efficient on-chip linear algebra and TinyML training acceleration,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:30.533129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:29:28.720003Z digest=sha256:4f9168c40bbccc4f9f44b27ef4fd973155bebe0ee192451c5ee95430a085819b

Observation c09561ed-bd44-40a2-9f50-dbc1fa446dcd · outbound

This paper cites FlooNoC: A 645-Gb/s/link 0.15-pJ/B/hop open-source NoC with wide physical links and end-to-end AXI4 parallel multistream support,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators FlooNoC: A 645-Gb/s/link 0.15-pJ/B/hop open-source NoC with wide physical links and end-to-end AXI4 parallel multistream support,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:30.278817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:29:28.830392Z digest=sha256:b4bc01fbcce76b6098ea48b35cdd4f99a12b97b1e94aab3a19a17a7716bd2673

Observation f01a241a-6578-4f5a-a103-07e2a786ee47 · outbound

This paper cites DRAMSys: a flexible DRAM subsystem design space exploration framework,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators DRAMSys: a flexible DRAM subsystem design space exploration framework,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:29.966934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:29:28.952136Z digest=sha256:1993a9a8e01a44547c1673c966865c251b699a8abe50d8bf03aa3b8e4f6a1508

Observation d0180082-2dda-450a-a015-df5fd2fa1707 · outbound

This paper cites SUMMA: Scalable universal matrix multiplication algorithm,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators SUMMA: Scalable universal matrix multiplication algorithm,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:29.651742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:29:29.047456Z digest=sha256:64f3b09a842f174abb9ff27d5316139ee027863cde940aa3e4ba4c1c4db80ae4

Observation 6d21fc12-e934-4bfc-924e-d552d33a1e7e · outbound

This paper cites MI300X vs H100 vs H200 Benchmark Part 1: Training,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators MI300X vs H100 vs H200 Benchmark Part 1: Training,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:29.425767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:29:29.161352Z digest=sha256:4275aceca697b5e3a4239d720aeb575db3cc60f9e654d99d4b9d37645dfb9395

Pith citing papers

No inbound Pith citation observations are available.