Pith. sign in

Paper Citation Record · LEDGER

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention

As of 6 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 5 inbound Pith citation observations for arXiv:2510.04212.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.04212 v3

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T10:01:56.131253Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T15:20:16.164286Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact23
  • verified fuzzy4
  • unresolved0
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch6

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 84ef2b0d-6d00-4360-8707-47ea35c43375 · outbound

This paper cites Scalify: scale propagation for efficient low-precision LLM training.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Scalify: scale propagation for efficient low-precision LLM training

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:02:31.677451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:865b5d979a3730578857fc630d8a33383276cebb65ca834016b89d83f384f882

Observation c8ce8386-332b-4824-a827-d4101c1410d5 · outbound

This paper cites u-$\mu$P: The Unit-Scaled Maximal Update Parametrization.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention u-$\mu$P: The Unit-Scaled Maximal Update Parametrization

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:02:31.733884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:f7d0e0e7e6c231d67f5d69fbfe323cfa7c08bb5bb538c80634067af44005df45

Observation acad3c74-c1c2-4dc6-bcbf-cf9f912f814c · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T10:02:31.810103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:307bedb9069d3fbd1a615565af8d3ac89ceda9ba2c76992741ad7a97950d6979

Observation 2747e26a-663e-407e-9239-46432204b867 · outbound

This paper cites Scaling FP8 training to trillion-token LLMs.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Scaling FP8 training to trillion-token LLMs

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:02:31.710568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:494af26d3974c354ca9ba8e2a69492d9b36ddd9f0d52afc2b5bf84a990b27da7

Observation a45bdfba-62a7-40cd-8e79-c5196ab742e5 · outbound

This paper cites Aaron Gokaslan, Vanya Cohen, Ellie Pavlick, and Stefanie Tellex.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Aaron Gokaslan, Vanya Cohen, Ellie Pavlick, and Stefanie Tellex

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T10:02:31.806824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:d9525e0e0675248c285a5cca4c9752578ad70a32e25126320e389165a93bcf92

Observation 0a266c88-bd65-4dac-bdbe-67731ecbcadc · outbound

This paper cites Is Flash Attention Stable?.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Is Flash Attention Stable?

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T10:02:31.719792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:bcfa1661ac16203f5bd92635f2d1e94cae47e1d013d6c662fd7f46c604de582b

Observation 51ef9cb5-1de1-4065-b39d-0102ff8bb07f · outbound

This paper cites Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-30T01:18:51.459998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:061bb09bb156468b5cd1754bdfafa4c761521b02fe1bb4c4dae86ef8b9743cf8

Observation cc065a55-4627-413b-9d94-158782c9e583 · outbound

This paper cites Query-Key Normalization for Transformers.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Query-Key Normalization for Transformers

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:02:31.673927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:05a43373b48ace52e34334c5a18761e9027633eeab56774c8ec3e895e67f3026

Observation 8f6b5e54-03fc-4290-a857-7034d2da5616 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Training Compute-Optimal Large Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:02:31.752653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:c096f2408874a68ad3f31c5096d6456c9b70c6f3bcbe2b53950b0e5cda1a184e

Observation 88aff18f-3fb0-4d42-b513-6972e031ff0f · outbound

This paper cites SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:02:31.706582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:ab3747ade45fc0060e03dffb1154b1f0937fa7568b98f6dceab94d8ab405ef59

Observation 52e00fa9-ec25-4b75-8a7a-7be7ccc302f3 · outbound

This paper cites A Study of BFLOAT16 for Deep Learning Training.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention A Study of BFLOAT16 for Deep Learning Training

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:02:31.738309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:ec84bb7257397151d2c0e97ed7deb5f0da542b76cb478a30f3573e3ab7a282ba

Observation 53f66d0e-6c7b-41f6-96be-f52c7003805a · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Kimi K2: Open Agentic Intelligence

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:02:31.756129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:68ba873917ffa784bf24ffb3f522837a79b55f2cc03031fa163b811ce7e30e17

Observation 0d9e66a6-2abb-4607-b5bf-4c19b90c21b2 · outbound

This paper cites DeepSeek-V3 Technical Report.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention DeepSeek-V3 Technical Report

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:02:31.702310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:7b4009ecc032b16e15cb8f4c769b47cf74856867a146e39ae1bf8dc921da95fc

Observation 3e77396a-2006-4532-94df-1611f448be03 · outbound

This paper cites Mixed Precision Training With 8-bit Floating Point.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Mixed Precision Training With 8-bit Floating Point

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:02:31.684327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:46c60c3d92a1b3e84ba7c050a989674e01557e814d159b3a0f46f42f4e592d1c

Observation 7b607e40-1368-4393-880c-4fb464af087f · outbound

This paper cites Mixed Precision Training.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Mixed Precision Training

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:02:31.654634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:19e264d0ad1cdb821ba4d8f51d9063bdfa2ffa3d654cdb3397990d170c0f32a2

Observation 267a6714-9fda-4947-8fd6-a76fb131a427 · outbound

This paper cites FP8 Formats for Deep Learning.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention FP8 Formats for Deep Learning

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:02:31.723095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:2df308e9e1263b8844af7e9ea514e47bac667e500c15ded6afe438a2a3b8f91c

Observation 89f76f7b-7d55-4d47-bbd3-2d7e6cf46b53 · outbound

This paper cites A Theory on Adam Instability in Large-Scale Machine Learning.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention A Theory on Adam Instability in Large-Scale Machine Learning

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T10:02:31.691756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:5ac48236469f0f378f3c27941b18cc309662fb671c17e84b81db5659ecd471f4

Observation 1e1dbe79-4ae7-41dc-939e-c2a6900e0c92 · outbound

This paper cites nanoGPT Issue.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention nanoGPT Issue

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T10:02:31.803305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:21cbcf7b6309c2499fc9175f666893026bdb285448f6f0ce88296df7bda06520

Observation 93c4ca0f-0f0f-492e-9f4b-9839f0659c4e · outbound

This paper cites 8-bit Numerical Formats for Deep Neural Networks.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention 8-bit Numerical Formats for Deep Neural Networks

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:02:31.639481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:ec2f9a4b3e67696c00313d738f1b1b5cd3ee26b7d0959eefa3b70108c8b58f24

Observation 65c71bf6-d9cd-4002-9fb2-6a0f447ec470 · outbound

This paper cites FP8-LM: Training FP8 Large Language Models.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention FP8-LM: Training FP8 Large Language Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T10:02:31.727622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:18a1899c18063c4bf54232ef80b959c052b170e3867616dedd1561655d4cf069

Observation 56d5338f-e883-4b21-a256-22388e7fa288 · outbound

This paper cites Training and inference of large language models using 8-bit floating point.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Training and inference of large language models using 8-bit floating point

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:02:31.696207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:e096eccdb5d1afdd05a724154fbb89861932feaf6321b08992a802a608f48dcb

Observation bf83668c-6f1c-4e1e-b0be-bd887d7e2282 · outbound

This paper cites Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:02:31.643403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:421cc86408679be0a8c617b8474da6fe27d448e6dad51f66a6954d504e14c7c8

Observation 35f406cc-da44-44be-9695-590dda5c8203 · outbound

This paper cites Qwen3 Technical Report.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Qwen3 Technical Report

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T10:02:31.647731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:457264cac764a52cb2554977f8039b8055ef0bf58081f6a6f3cbcd902f522326

Observation f223dff7-69e1-434c-8aa5-d4a7d67c0139 · outbound

This paper cites Scaling Language Models: Methods, Analysis & Insights from Training Gopher.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:02:31.746740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:140477c9f498748cad7874ea6dd8ee6fe9d18718d8436b68d66333a431a2b727

Observation 22a56c48-f98b-43e8-84e5-c06205cf2fcc · outbound

This paper cites Methods of improving LLM training stability.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Methods of improving LLM training stability

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:02:31.742547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:62b0898569c919bd8d651f0146665bae74f402d75e09e760926327a72ffc5dd3

Observation ab409d59-1954-493a-bdfc-a17037221042 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention LLaMA: Open and Efficient Foundation Language Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:02:31.651207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:2b54738a6fc2748c510919d8948d4c2f622886f26f06726f13d1de799dc16445

Observation 0ee60707-4b4f-478a-ab91-f55567e38442 · outbound

This paper cites Training LLMs with MXFP4.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Training LLMs with MXFP4

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:02:31.760017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:d376e850647e437b7dcde90b6d0feccb77f7f56eadd21ed82d9d73713517a43c

Observation 7eee6e6f-7efe-42c7-a26d-c059d5f17b33 · outbound

This paper cites Optimizing Large Language Model Training Using FP4 Quantization.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Optimizing Large Language Model Training Using FP4 Quantization

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T00:00:17.330307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:1e634f213c839b0472be71a387b5dbf1eb97c437ec210e48ed42cf490771ab7c

Observation 8f603f96-661d-4b81-a8c9-86a00ab49069 · outbound

This paper cites Mitchell Wortsman, Tim Dettmers, Luke Zettlemoyer, Ari S.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Mitchell Wortsman, Tim Dettmers, Luke Zettlemoyer, Ari S

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T10:02:31.799841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:1aba9168b11b44af3d61f69b638bcd624033f645eda533f57ae51cf01df8d27a

Observation 7b776de5-23a8-4a73-bd64-03fed27455ee · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Efficient Streaming Language Models with Attention Sinks

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T10:02:31.687763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:aed54cf8b58857310791d0774c2f83c71222cbad2d9029e56d9791db89a1bded

Observation ab5518ee-fc94-4220-a46e-72af56d69016 · outbound

This paper cites Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:02:31.664367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:83c5f3d71ad7c76bae5004596231437132e39db41812bf185c37a78470b01a22

Observation 94f1e29d-3a57-407c-8f9e-e46d13956016 · outbound

This paper cites A Spectral Condition for Feature Learning.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention A Spectral Condition for Feature Learning

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T10:02:31.669691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:9c30834c9aa74f2dd361528d204467c141a18a1f311e2c614296b947a2242c5f

Observation 2afeead0-f78f-4c56-b46a-2fd1adf104bb · outbound

This paper cites Towards Efficient Pre-training: Exploring FP4 Precision in Large Language Models.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Towards Efficient Pre-training: Exploring FP4 Precision in Large Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:02:31.635398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:77c1418c26b56695f4a939f76d1edf8d500844d1b5dc3ca359b6e5316d28d97e

Observation 7d863fef-6e0c-4126-91d7-be66321cb39a · outbound

This paper cites gradient spikes.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention gradient spikes

Reference 35

Resolution
malformed identifier
raw_fallback, observed 2026-05-18T10:02:31.796249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:3b9111f938a23814f515c8fd310c9f340077628858fb33eb49c3ee0302aafa1a

Observation a5f9a6f2-e73d-488c-b77f-1eed422f3313 · outbound

This paper cites Seg en à st 're ich s ho hem S oh ne / Un ser m Kaiser Ferdinand !.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Seg en à st 're ich s ho hem S oh ne / Un ser m Kaiser Ferdinand !

Reference 36

Resolution
malformed identifier
raw_fallback, observed 2026-05-18T10:02:31.793196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:8a13928bba38dd52ef267216e06aae0f8cc0703cc878eb78f2f7b7a85ae6c7b1

Pith citing papers

Observation d8e3f90e-889b-4d2a-94c3-39e22dfc17bc · inbound

Reference Traces for Auditing Invisible Weight Updates and Guiding Exact-Budget Protection cites this paper.

Reference Traces for Auditing Invisible Weight Updates and Guiding Exact-Budget Protection Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T15:32:26.691504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:32:26.691504Z digest=sha256:5deeb0b756b3c92b05ccc8998014f6fcbd6abc62de42b3439d01f0b4a85164c8

Observation ab61e964-02b0-4497-874c-767d0fae2e8b · inbound

Reference Traces for Auditing Invisible Weight Updates and Guiding Exact-Budget Protection cites this paper.

Reference Traces for Auditing Invisible Weight Updates and Guiding Exact-Budget Protection Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T07:54:47.068638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:54:47.068638Z digest=sha256:33203b5972d54a50b0108cfe35d8048ce6648450bd615877f29e0c3111b0a53c

Observation 55e5c0eb-4757-49d1-9eb8-3d238c96c5ab · inbound

An Isotropy-Preserving Spectral Cap for Muon: Theory and Three Case Studies cites this paper.

An Isotropy-Preserving Spectral Cap for Muon: Theory and Three Case Studies Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:35.779079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:35.779079Z digest=sha256:5eededfeb57bdc32ccc82c33dfcfa16983dfbc9940b35396e98eea1e576fc869

Observation f11b71b5-ce65-43a5-9ff8-5b4734344cb7 · inbound

Automated Numerical Stability Analysis of Deep Learning Operators cites this paper.

Automated Numerical Stability Analysis of Deep Learning Operators Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T02:20:39.957037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:20:39.957037Z digest=sha256:af11a6a3afccfef13ee81eb3680ac7ed99da104c4dbb97d2ee77bbc04508bed9

Observation 55a7f5a8-c489-48d9-a2be-bfe3e7067fcb · inbound

One QK Channel, Many Sources: Guarding Low-Precision Attention Collapse cites this paper.

One QK Channel, Many Sources: Guarding Low-Precision Attention Collapse Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T15:20:16.164286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T15:20:16.164286Z digest=sha256:0b9bfb33f10a6d47f1ddca4651f8b374895e549c319b8d7dce42e76243cb7952