Pith. sign in

Paper Citation Record · LEDGER

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers

As of 9 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 2 inbound Pith citation observations for arXiv:2507.16676.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.16676 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:11:14.757058Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T04:14:11.793614Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T01:06:24.646981Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1726cef6-75a7-4ce1-837c-824687e063c7 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:11:15.258961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:11:12.856969Z digest=sha256:8e4755e8c20b5fa5d39373c5a0bff55f3e79785cf64027f1a4b6323d10906a88

Observation 4da5dd94-c367-4018-ba04-582409e69339 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers LLaMA: Open and Efficient Foundation Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:11:12.939944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:11:12.939944Z digest=sha256:86e4ad352ee075cadee924d69fc068f68e64b1dfd75ff3bd27c7262f103cde20

Observation a1afa995-413d-4e94-b78c-a71235cd3281 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:11:15.241717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:11:12.994346Z digest=sha256:43e0cfdecba78d30db961992436a24180b8f0f222c7734c40fa428829a901cc0

Observation 20d921e8-6919-4c0d-9d9e-0a0a27c4412e · outbound

This paper cites Attention is all you need,.

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers Attention is all you need,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:11:13.029505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:11:13.029505Z digest=sha256:c49c4620e7ae9b6884b98cd38e82a344ce9b12ba35a1763008a0f829ebf4dcf6

Observation a09977e7-0f39-4512-bbe8-e4077bc9fbe9 · outbound

This paper cites BERT: Pre-training of deep bidirectional transformers for language understanding,.

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers BERT: Pre-training of deep bidirectional transformers for language understanding,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:11:15.213824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:11:13.142377Z digest=sha256:ff2aada8dad5dd5372552ecb4b70a0b276020a9f3090ada832cad978e513f37d

Observation 2a891ef4-de55-4400-92cc-4e2116c50824 · outbound

This paper cites ALBERT: A lite bert for self-supervised learning of language represen- tations,.

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers ALBERT: A lite bert for self-supervised learning of language represen- tations,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:11:15.195794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:11:13.230159Z digest=sha256:eb61752f9d0a24c6babf48b32799d34a6e662d19134378120a8e9a6887fe4beb

Observation 1537dfb6-9156-4d47-bb92-c1090cd701a8 · outbound

This paper cites Longformer: The Long-Document Transformer.

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers Longformer: The Long-Document Transformer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:11:13.311438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:11:13.311438Z digest=sha256:492843bbf3f671049d1c9a9abddc894158100a2450cbd0e421ab72bc55fe1986

Observation 67ea9787-9010-4110-b306-2a138b8e6fe1 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers Generating Long Sequences with Sparse Transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:11:13.465021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:11:13.465021Z digest=sha256:3545a1acabbbb857e1ff7362531b7cc41e8160d18c6eb2f25a48f9d8cab3287f

Observation c3642d0b-c345-4a11-bca0-c8daf9967d3b · outbound

This paper cites Transformers are rnns: Fast autoregressive transformers with linear attention,.

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers Transformers are rnns: Fast autoregressive transformers with linear attention,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:11:15.174047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:11:13.606383Z digest=sha256:9344ac7cc0dd352055a4a3c74fbca72fe6647fa8441f0c8c07b7afc51c9ced20

Observation 2b1bdcbb-ad84-4906-8dea-3dd5314f9036 · outbound

This paper cites Linformer: Self-Attention with Linear Complexity.

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers Linformer: Self-Attention with Linear Complexity

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:11:13.710772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:11:13.710772Z digest=sha256:a8630d9779e7a2e5a2d6acda07635b034ba81708ebd992f4fd5bdbfd0456ad65

Observation e2a02d83-144b-4fc3-805c-e5fa7ff7d6b0 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with IO-awareness,.

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers Flashattention: Fast and memory-efficient exact attention with IO-awareness,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:11:15.159661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:11:13.849399Z digest=sha256:ff6ba4e163b5e575e313815e47be61f07e387c603506b51d0e2d2dc579591b10

Observation 224d45e6-ed93-4230-ba39-554804dfef5c · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:11:13.979517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:11:13.979517Z digest=sha256:34a7b41b800a8e77097c08bd3c636c787fb0972605b8c3ec84e09a9a01a50887

Observation 638ae11d-b20d-40e9-b01c-805952011d9b · outbound

This paper cites Self-attention Does Not Need $O(n^2)$ Memory.

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers Self-attention Does Not Need $O(n^2)$ Memory

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:11:14.081783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:11:14.081783Z digest=sha256:2ce4b51c0a83d714db3b106a1182a2e42270c81e4d93824b3064c0259a370db0

Observation f871cac0-fd3d-409c-915f-bb332ca2a661 · outbound

This paper cites Testability and dependability of ai hardware: Survey, trends, challenges, and perspectives,.

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers Testability and dependability of ai hardware: Survey, trends, challenges, and perspectives,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:11:15.146191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:11:14.223539Z digest=sha256:0c072f01f732b4d934058ce59ddee6fca318af4160a93f970473d1d2b66a54f7

Observation 71bc6a85-2426-4b23-83df-16ba079075cb · outbound

This paper cites Radiation-induced soft errors in advanced semiconductor technologies,.

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers Radiation-induced soft errors in advanced semiconductor technologies,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:11:15.133086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:11:14.334947Z digest=sha256:199efdfb2b6b5ef5ec624d506a39869c4fb2c53ba10e05ca048761603455cb59

Observation 7ca68cc3-8aee-44dc-b5f6-ffa33f6aa9dc · outbound

This paper cites Designing reliable systems from unreliable components: the challenges of transistor variability and degradation,.

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers Designing reliable systems from unreliable components: the challenges of transistor variability and degradation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:11:15.118044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:11:14.499102Z digest=sha256:d7faece379f51ce4634741805cb5e5335f256169939ce83c8cb5a67f694307a4

Observation ee666352-902a-450a-bdd1-06b1f0d9933f · outbound

This paper cites Koren and C.

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers Koren and C

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:11:15.103951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:11:14.643518Z digest=sha256:397080538d337ebc046e2ea4f16fcc08e01528fff2cbf11666f28bfdfc8d31bd

Observation 0cbad548-a087-4880-8e34-edd838c1435d · outbound

This paper cites Algorithm-based fault tolerance for matrix operations,.

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers Algorithm-based fault tolerance for matrix operations,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:11:15.089713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:11:14.700405Z digest=sha256:a2922fd958418dd47a2b411a0760e27c6811cb875b95b4e0d33db8cf366ad89c

Observation 7bb21ce6-8621-46fb-a2d0-12c872287a08 · outbound

This paper cites Towards practical algorithm based fault tolerance in dense linear algebra,.

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers Towards practical algorithm based fault tolerance in dense linear algebra,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:11:15.074743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:11:14.704838Z digest=sha256:826aa124b7424a61d1177c9fdfee1fbc926798c183d5bdf305a8232888e8ce10

Observation e996bb4f-9dc7-4655-bc83-9b40101f8eac · outbound

This paper cites Low-cost online convolution checksum checker,.

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers Low-cost online convolution checksum checker,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:11:15.059694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:11:14.708914Z digest=sha256:f55180e75d47824560bf4c5bfe4721f654314c26ce6063eb650cfaddfc471539

Observation 36c45af1-b32d-4f3d-a4bb-903cd44c9458 · outbound

This paper cites Making convolutions resilient via algorithm-based error detection techniques,.

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers Making convolutions resilient via algorithm-based error detection techniques,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:11:15.046116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:11:14.713172Z digest=sha256:6917f7ce3f43df65aba3d9f83bae321dac3f8f69121211dd010c82ffaec8c19d

Observation 191a8307-dc79-4808-9aa7-2b79db652095 · outbound

This paper cites GCN-ABFT: Low-cost online error checking for graph convolutional networks,.

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers GCN-ABFT: Low-cost online error checking for graph convolutional networks,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:11:15.032332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:11:14.717974Z digest=sha256:0d4b29b32914f280950ce1dae1d9f16e4a25abd1d5d21b4bbf64a94aae7f95ec

Observation cd30ca60-d1c9-4606-802d-66762c458f80 · outbound

This paper cites ApproxABFT: Approximate Algorithm-Based Fault Tolerance for Neural Network Processing.

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers ApproxABFT: Approximate Algorithm-Based Fault Tolerance for Neural Network Processing

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:11:14.722039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:11:14.722039Z digest=sha256:d9436cf1cf0d0c3af681bad35a8ba89ebfdcb63c218ee40a4be3c012dc0646e5

Observation 8023e98d-5b9e-4789-aca6-5a691cffcec9 · outbound

This paper cites ATTNChecker: Highly-optimized fault tolerant attention for large language model train- ing,.

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers ATTNChecker: Highly-optimized fault tolerant attention for large language model train- ing,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:11:15.017561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:11:14.726171Z digest=sha256:d9b987adad5d4a8b4763a700a50aa12f22286de7dd260a5e3d6f87b69eac0a24

Observation 5d8a004a-ae2c-4111-9990-a4fac118f290 · outbound

This paper cites Error resilient transformers: A novel soft error vulnerability guided approach to error checking and suppression,.

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers Error resilient transformers: A novel soft error vulnerability guided approach to error checking and suppression,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:11:14.998636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:11:14.730302Z digest=sha256:c5d6dda803f78c87b3fa6e9d2ef0da7e06d584df7efb59098f1e1d4cc2e51945

Observation aa88b97b-615f-45d0-9b79-48bd054a1690 · outbound

This paper cites Language models are unsupervised multitask learn- ers,.

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers Language models are unsupervised multitask learn- ers,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:11:14.983410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:11:14.734026Z digest=sha256:5dd63e0d2db280360fa83d4f7c09aa558292ef73fa32fe8a08b88035a823cc0e

Observation 11c98e64-7fa8-4204-b16c-3efb622cb0ab · outbound

This paper cites Attention is all you need,.

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers Attention is all you need,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:11:14.969704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:11:14.738353Z digest=sha256:3a5c24ba1d4fe00dd8fc330ff9a486c99d32df8f2bd548ff564f1aadffca2665

Observation d201a12b-bd95-42cb-ba83-5c45b9d0acbc · outbound

This paper cites Mnnfast: a fast and scalable system architecture for memory-augmented neural networks,.

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers Mnnfast: a fast and scalable system architecture for memory-augmented neural networks,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:11:14.955836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:11:14.742149Z digest=sha256:a98c6e7af5ad534424a88afd6905b970ae53b79a045ef6d86ee9fe0ffc07a948

Observation 7d34483d-868d-4fac-a524-80541c150029 · outbound

This paper cites Online normalizer calculation for softmax.

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers Online normalizer calculation for softmax

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:11:14.746427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:11:14.746427Z digest=sha256:846e524c0dd76fc3d442cddc5b0e187400c60614f26ebd066ae5e4e3dfc1245d

Observation b25e2d2f-7a4a-4b63-a6d5-3c82c9e06d6b · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T15:11:14.752400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:11:14.752400Z digest=sha256:0600af8597816b07c3a3bf25ba0df20bdcbf9366693eda480f4d286af0a0ce03

Observation 8973ae6c-3569-4362-b91e-d8841794b4e3 · outbound

This paper cites Automatic detection of floating- point exceptions,.

Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers Automatic detection of floating- point exceptions,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:11:14.940425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:11:14.757058Z digest=sha256:5cd0799b75f7cb01af8ae7494fea43d76b5e3b645a50dba4c43291c9964b6696

Pith citing papers

Observation 5ada3721-6945-4e6a-a34f-318d1719c932 · inbound

Not All Errors Are Equal: A Systematic Study of Error Propagation in Large Language Model Inference cites this paper.

Not All Errors Are Equal: A Systematic Study of Error Propagation in Large Language Model Inference Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:06:24.648870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T12:36:23.320194Z digest=sha256:28264b8b74e45cc67281c1e4286bd2bdee6d714e232d3f4ed63fe040214ee9ca

Observation 6658f018-4ac8-44a5-ac88-3c0f9c2152f5 · inbound

Self-Verifying Measurement Records: Hash-Linked Evidence Graphs for Hardware Benchmarking cites this paper.

Self-Verifying Measurement Records: Hash-Linked Evidence Graphs for Hardware Benchmarking Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-01T17:05:50.955968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T04:14:11.793614Z digest=sha256:926b96f799ea956358943edef8120825b2da4c264813210ebe40b9659de0f92d