Pith. sign in

Paper Citation Record · LEDGER

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU

As of 20 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2607.09385.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.09385 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-13T03:24:08.548348Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8f797521-39f6-4039-a6c6-7b6c7474d6f3 · outbound

This paper cites AIOS: LLM Agent Operating System,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU AIOS: LLM Agent Operating System,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:72e91e30ea10f510c1d50990bddb5c854222229d84dbb3bdc9c5662d311679c3

Observation 6cabc7fb-7d3e-4737-bb29-1bfe8855bd16 · outbound

This paper cites Intelligence per Watt: Measuring Intelligence Efficiency of Local AI.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU Intelligence per Watt: Measuring Intelligence Efficiency of Local AI

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:3847b7a3f25ff79e22589db17089c9ba2628cdca8b5395ff669716e986584f5b

Observation c86430ee-29e0-452a-bc8e-e86670bc3a88 · outbound

This paper cites A survey on privacy risks and protection in large language models,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU A survey on privacy risks and protection in large language models,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:37a41c6cfb460baa33c3801c4abb89524e624a1f39eea4db22b809eaba33872b

Observation 0169cce9-deab-44a9-8f4c-7623e2f15737 · outbound

This paper cites AMD XDNA NPU in Ryzen AI Processors,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU AMD XDNA NPU in Ryzen AI Processors,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:025596295eba193a0a86fb79002dbd662d5693372b84de0651e0cf9213b26183

Observation 791d64a9-ac84-444e-b942-900736d56b01 · outbound

This paper cites FastAttention: Extend FlashAttention2 to NPUs and Low-resource GPUs.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU FastAttention: Extend FlashAttention2 to NPUs and Low-resource GPUs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:db213e4e1bc919f43648c2ac48f20cadf4e0c265f06b99feda6a44ec1d8e9843

Observation 51625f01-9fe2-453c-8512-c5041da42fcb · outbound

This paper cites NITRO: LLM Inference on Intel Laptop NPUs.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU NITRO: LLM Inference on Intel Laptop NPUs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:52fed9874be71d119a993f2558040e8c4129f01d5f661365a67abecb70847704

Observation af5092ae-d8e8-4f46-86e4-7fe8a3f2fb01 · outbound

This paper cites LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:2a407e9ec2f38b05b4045fa666016659249b861e785d7c2f498c60177ed5ecb2

Observation 2e0449c7-df1e-4cf9-bb57-244104506728 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:eba68b37a04ac3969e45e0547aa22ab27dcd08b103397e2079ed4bea62635cf2

Observation ec5e0bda-8a02-41f0-904c-6d37dd09f835 · outbound

This paper cites ITA: An Energy-Efficient Attention and Softmax Accelerator for Quantized Transformers,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU ITA: An Energy-Efficient Attention and Softmax Accelerator for Quantized Transformers,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:54c00617f51d9c3d2e6dd6a3b2bbcfe324b3045a7d8d44a8cc8270eaa43ec912

Observation 49393d3c-451d-4193-8683-07a989579fdb · outbound

This paper cites PACE: An Optimal Piecewise Polynomial Approximation Unit for Flexible and Efficient Transformer Non- linearity Acceleration.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU PACE: An Optimal Piecewise Polynomial Approximation Unit for Flexible and Efficient Transformer Non- linearity Acceleration

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:9b703d43f7a63e6a164eb60760131a90e7ad50d692118e1ca37584dc0f62ba76

Observation 298e2b47-4141-406d-9e07-8406b95e6788 · outbound

This paper cites Efficiency, Expressivity, and Extensibility in a Close- to-Metal NPU Programming Interface,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU Efficiency, Expressivity, and Extensibility in a Close- to-Metal NPU Programming Interface,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:c3551ac288257585cecc50f29095cdeeb3a13a54435a2fb40759045263e1211d

Observation c5179c24-4ad6-4893-8a0c-ee762a67cb6d · outbound

This paper cites Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:60771eee60d239122be76847741a22a26a673a2e9cd1d582c9c87d733ee3f619

Observation 04dc20c6-24a3-41c4-a0f9-208596bc28e4 · outbound

This paper cites Dato: A Task-Based Programming Model for Dataflow Accelerators.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU Dato: A Task-Based Programming Model for Dataflow Accelerators

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:fb1639dcf81f5f66846bb8498a3d6fad1b8470daa672acc373786604a99bfc6d

Observation e84660a6-0c51-43d8-9400-dd4e96c4dbab · outbound

This paper cites The Llama 3 Herd of Models.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:6f5c628159662fd9df575a8ee11799218b777d22e2d82d55ecf95a683a17f745

Observation 81227a4a-1e43-4901-bdfe-527f98b79a34 · outbound

This paper cites Attention is All you Need,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU Attention is All you Need,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:3d06c2c60516106fd87143203b1b9f8a5d741b7680937df6705a3a2a627cb8e9

Observation e7938933-4197-4fa2-9022-0f06f642c0fb · outbound

This paper cites Online normalizer calculation for softmax.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU Online normalizer calculation for softmax

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:6d3081920e706c4ae720150efab499757107025b4b8e0c53341c1898b34bd3cb

Observation 7cf0515c-fe3e-4696-b449-af178db110ce · outbound

This paper cites Qualcomm Hexagon DSP: An architecture optimized for mobile multimedia and communications,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU Qualcomm Hexagon DSP: An architecture optimized for mobile multimedia and communications,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:774eeb1ef5aade25c2ee74d004be551271866632cfbdf3da6f6fb8db575a354f

Observation bf284211-aa75-478b-ba22-d8d79a5c40a4 · outbound

This paper cites ASCEND-CC: Confidential Computing on Heterogeneous NPU for Emerging Generative AI Workloads,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU ASCEND-CC: Confidential Computing on Heterogeneous NPU for Emerging Generative AI Workloads,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:8c2cd5833be5da646e2f097211441ae68efa0eb6f2569e24820bbc2c5b9c22db

Observation 89afd327-e73c-4ae8-a86b-eec2d7666a7d · outbound

This paper cites SPARTA: Spatial Acceleration for Efficient and Scalable Horizontal Diffusion Weather Stencil Computation,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU SPARTA: Spatial Acceleration for Efficient and Scalable Horizontal Diffusion Weather Stencil Computation,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:a75ee0165558f964c58bccec6014dcf460b160e30f4bb035a967961df08688ec

Observation e9be7f1d-72f1-4fb9-8a13-f64e5799deea · outbound

This paper cites You Only Look Once: Unified, Real-Time Object Detection,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU You Only Look Once: Unified, Real-Time Object Detection,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:99c05edd384a4942e7e56160502b7a35b1fc95e99020c4841d4611e492331fc7

Observation 939f641c-a518-4d11-a121-0072a340699f · outbound

This paper cites 14.5 Envision: A 0.26-to-10TOPS/W subword- parallel dynamic-voltage-accuracy-frequency-scalable Convolutional Neural Network processor in 28nm FDSOI,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU 14.5 Envision: A 0.26-to-10TOPS/W subword- parallel dynamic-voltage-accuracy-frequency-scalable Convolutional Neural Network processor in 28nm FDSOI,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:e5073a62afc6f164dde781305c94b206bb63524bfe730655c7c80cfa93dafbae

Observation c3e47a6b-c8dd-4345-ac33-9ceea2c88789 · outbound

This paper cites Basic Linear Algebra Subprograms for Fortran Usage,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU Basic Linear Algebra Subprograms for Fortran Usage,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:0660d7c47dab3042b4f5dd556f3295e956ad7fc3d271c74f736addbbe938bb28

Observation 231500f8-a40c-4e67-a058-00385c9cff44 · outbound

This paper cites AKG: automatic kernel generation for neural processing units using polyhedral transformations,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU AKG: automatic kernel generation for neural processing units using polyhedral transformations,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:eb2c9911d33aa370d0799827bead1d7da8774f1bf2160481d48a0d7df21699a2

Observation 99b59607-afdf-41b0-a2ad-0073f3ef1f0b · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU PyTorch: An Imperative Style, High-Performance Deep Learning Library,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:6f87eb47f2912036c40bd91ba1f6d39d6089616c502bf4ebbb7661ae258a4f9e

Pith citing papers

No inbound Pith citation observations are available.