Pith. sign in

Paper Citation Record · LEDGER

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training

As of 9 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 1 inbound Pith citation observation for arXiv:2502.05967.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05967 v3

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:19:29.257758Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-19T16:35:40.231293Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T16:37:39.800519Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 96ddc2c1-2900-4902-9aa8-8785742dc3e0 · outbound

This paper cites Scaling FP 8 training to trillion-token LLM s.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training Scaling FP 8 training to trillion-token LLM s

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.723781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.124327Z digest=sha256:df5e78d6d91402bd0ca28a050ad0bb6dfea940b9eee91df0d0e25d4756996953

Observation be115354-0bf3-490b-8643-6dd7d0517722 · outbound

This paper cites Calibrating the Mosaic evaluation Gauntlet , 4 2024.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training Calibrating the Mosaic evaluation Gauntlet , 4 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.713434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.128906Z digest=sha256:0242076ac2c9f3029beec22891dc5e79048244257a4d0170824f12ccd8a06b66

Observation 1ae901c9-233d-4631-bc4f-e16d4ce3175f · outbound

This paper cites Unit scaling: Out-of-the-box low-precision training.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training Unit scaling: Out-of-the-box low-precision training

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.703031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.132981Z digest=sha256:d0014625b6aaacf819e80fb8def9afffa3861e85613be8ab9403f018c0cfc295

Observation 4b7d9766-8df9-4506-8ef9-042a19fff08b · outbound

This paper cites Y., Deiseroth, B., Cruz-Salinas, A.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training Y., Deiseroth, B., Cruz-Salinas, A

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.691447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.137768Z digest=sha256:6425d372c84afb91ce41b3a81c4e09dde596b1207009d4395c6a6fe733082c10

Observation e5c59901-c6f5-45d5-a2d8-1fa0f4027f2a · outbound

This paper cites and Berger, R.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training and Berger, R

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.681667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.147562Z digest=sha256:c7a76a2411bb468e86e7dcf5624df9e8352af14d3d3b6f04512eafc78ad4ae94

Observation 63e97931-b0fe-4eb0-b0b5-cdfa73cab07c · outbound

This paper cites an unresolved cited work.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-08T17:19:29.670824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.151332Z digest=sha256:c218171b3a38cd844a54770c52183e6c7b35df8b9839525fdd16bd415a652c83

Observation 83163bd4-76ba-4dc6-822d-0ef0c3b474a1 · outbound

This paper cites LLM .int8(): 8-bit matrix multiplication for transformers at scale.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training LLM .int8(): 8-bit matrix multiplication for transformers at scale

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.660262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.156732Z digest=sha256:3a8160d6abeb882c5c469c4dc4bf262bbc5868dce9946199c32f9142a3c60d36

Observation c5b080e9-8866-464e-9a1e-05eae69138c8 · outbound

This paper cites Blazingly fast LLM evaluation for in-context learning, 2 2023.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training Blazingly fast LLM evaluation for in-context learning, 2 2023

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.649863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.161453Z digest=sha256:104010d1b0bf520b05b38e64035cbc7a12fcc096f65dc36be8b15aa134c814b3

Observation 131c9f01-2a4d-4373-bdc1-e7c3d7d42eef · outbound

This paper cites Flex Attention: A Programming Model for Generating Optimized Attention Kernels.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training Flex Attention: A Programming Model for Generating Optimized Attention Kernels

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T17:19:29.164938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:19:29.164938Z digest=sha256:29e9d7248a4ff5804e0b1de74f47dcbc96f579079da405a2ee41055b8f67ca90

Observation 5b033ee9-601f-4d05-bac6-400ec649f6a7 · outbound

This paper cites FBGEMM: Enabling High-Performance Low-Precision Deep Learning Inference.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training FBGEMM: Enabling High-Performance Low-Precision Deep Learning Inference

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T17:19:29.168874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:19:29.168874Z digest=sha256:0d24043eb3ea076413ec9d5d1e4e6fbfc1dfd95591cffc24549ef5325eeb961d

Observation c2caf857-8d98-465e-843d-2c3babedd853 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training Adam: A Method for Stochastic Optimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T17:19:29.172567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:19:29.172567Z digest=sha256:86c5fe9eb33ece680f43f7c6f69465d4cab1d5cfc0026c0adf40d674853ae049

Observation 71dac2fd-01c7-4d32-a163-a40994b02812 · outbound

This paper cites An Empirical Study of $\mu$P Learning Rate Transfer.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training An Empirical Study of $\mu$P Learning Rate Transfer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T17:19:29.177838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:19:29.177838Z digest=sha256:4d58394e102e68c7dfca06c5ded361164ecfa8d67614c75a77ad1a75d571899a

Observation 38016c2e-d8c8-4112-9998-e32ca5ea3471 · outbound

This paper cites Swin transformer v2: Scaling up capacity and resolution.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training Swin transformer v2: Scaling up capacity and resolution

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T17:19:29.182203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:19:29.182203Z digest=sha256:25ca50bf15674b22f9a422509c8b12319c1f7ad48d6cc73732691f23ca21bae8

Observation aa95d758-2d0a-41e1-88ab-269b8b08569e · outbound

This paper cites Mixed precision training.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training Mixed precision training

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.632410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.185719Z digest=sha256:e1b96a3a5bf007f03db468ace903e40676397408cbb4ae9647123bffe9ae6828

Observation 1ae13680-7316-4c76-bd0e-eeb389e9106d · outbound

This paper cites FP8 Formats for Deep Learning.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training FP8 Formats for Deep Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T17:19:29.189352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:19:29.189352Z digest=sha256:90f1470dc6fc647b60e614c1356c16b8b31437640424646f8e2ca95dd967d61f

Observation 58904f1e-1df3-4abd-b755-600b094d9e95 · outbound

This paper cites I., Alizadeh-Vahid, K., Mehta, S., del Mundo, C.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training I., Alizadeh-Vahid, K., Mehta, S., del Mundo, C

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.621493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.193163Z digest=sha256:304975b0848bbe10d648abb5a8460ef17b804eea23a1ec0f09d4f6087a135c29

Observation 3f90a058-88e0-48f3-8eff-0275d559a572 · outbound

This paper cites Composer.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training Composer

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.610565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.196804Z digest=sha256:9355e46c4ebab682710e5bf187027d1723391a1d60dda1bd1407cf9740016be5

Observation 3d7f6d4f-077d-4fb7-b60d-eb2a0c67a44e · outbound

This paper cites LLM F oundry.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training LLM F oundry

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.599524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.200330Z digest=sha256:b656952e5703b162aa54f079d47273768b41c401a38fb409f7d7c76ff2379103

Observation 80a15c84-ba1c-4aa0-aba8-25a7aef724fb · outbound

This paper cites Streaming.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training Streaming

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.588442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.203719Z digest=sha256:cd2e10b82740d6fb04835a010b5a43c220523f9b2c5593312db0596d96e25b16

Observation 55d2c8bb-f9de-4b44-b4fb-ddc3c061e630 · outbound

This paper cites Asynchronous multiply-and-accumulate instruction: wgmma.mma\_async.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training Asynchronous multiply-and-accumulate instruction: wgmma.mma\_async

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.577873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.208054Z digest=sha256:d2e4a9e8d59eb116e50983d2486d27bf6de16e2294178fb4b839be21ccc3a6e7

Observation c417dd04-99ef-49cd-92b8-c135a238169d · outbound

This paper cites Transformer E ngine, 2023.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training Transformer E ngine, 2023

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.567179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.212306Z digest=sha256:193a469ed0456131e4f1106eb73ebc5fe0b78ec7714948f3a0acc336b41b2565

Observation 4404d4ff-b480-467d-903c-1de8557f7243 · outbound

This paper cites cuBLAS : cublasLtMatmul().

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training cuBLAS : cublasLtMatmul()

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.556156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.216823Z digest=sha256:ab17b1a16bdfd95c36fbe93bc89d4cdd0460879ead2c6bc20e1a6bdfdd169f5f

Observation f9c9cec1-849b-4b84-ae3a-4f0d1c8534cc · outbound

This paper cites 2 OLMo 2 Furious.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training 2 OLMo 2 Furious

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T17:19:29.220412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:19:29.220412Z digest=sha256:42180383522f03903d8b4ae724eb15ebd205831cdac7870cebbfb1efc6bbcdc1

Observation 4d7e4bda-9da5-49cc-993a-5c765416cdb3 · outbound

This paper cites V., Cui, X., Zhang, W., and Gopalakrishnan, K.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training V., Cui, X., Zhang, W., and Gopalakrishnan, K

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.544769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.224893Z digest=sha256:487185b347abdfce7b5ddd0363257f2ad09ffaf56ec610da8ebc760fa1b301c3

Observation 9235daec-c6e0-4fe5-befc-e41a88f096f5 · outbound

This paper cites T., and Cox, D.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training T., and Cox, D

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T17:19:29.229551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:19:29.229551Z digest=sha256:b608c97112afcabe7e4cc6d6ad0ffb62a1bf398d4c064e928debc7eafb7dc220

Observation 2247b83d-277f-4b2f-b595-915f8d23d1e0 · outbound

This paper cites N., Kaiser, L.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training N., Kaiser, L

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T17:19:29.233664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:19:29.233664Z digest=sha256:d06c2ab909415dd2246e9b0a79ab71aa81ab1a518cdfdae8694007a7c7af1ad4

Observation 90af06df-b171-4857-bc55-902473b9d5df · outbound

This paper cites How to set AdamW's weight decay as you scale model and dataset size.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training How to set AdamW's weight decay as you scale model and dataset size

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T17:19:29.237555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:19:29.237555Z digest=sha256:9ecf4309527db2d73419e69daa931ba642430d4807860d2f077d57320ec50ce7

Observation 5990abfe-fed8-4212-8098-42e725d1572a · outbound

This paper cites J., Xiao, L., Everett, K.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training J., Xiao, L., Everett, K

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.526729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.241813Z digest=sha256:f8f64175d69a7c2e23deaabc7ed7ed0303adfd2af4580662a5325a8094ebcd65

Observation f43f067b-788c-4ce3-b69e-e1547344345b · outbound

This paper cites J., Babuschkin, I., Sidor, S., Liu, X., Farhi, D., Ryder, N., Pachocki, J., Chen, W., and Gao, J.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training J., Babuschkin, I., Sidor, S., Liu, X., Farhi, D., Ryder, N., Pachocki, J., Chen, W., and Gao, J

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.515309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.245461Z digest=sha256:b896cc901a37fd7a8bb9bcfaae0bcc463ac2280b64f1c2357ccf408811e172e7

Observation b974ecab-14f8-4001-bdb9-dc2704304119 · outbound

This paper cites A Spectral Condition for Feature Learning.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training A Spectral Condition for Feature Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T17:19:29.249207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:19:29.249207Z digest=sha256:d410e64d093cdf15a9fbe454a7b5d75c03b375f6e7b7ca54952eebd04f917e90

Observation 5c200309-8b24-494a-9c8d-de876a55f057 · outbound

This paper cites Tensor programs VI : Feature learning in infinite depth neural networks.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training Tensor programs VI : Feature learning in infinite depth neural networks

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.504154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.253110Z digest=sha256:d08583ed788cdf025700e9fa4ec0d00a5f049f4d1ac426a033f7b02cf8f8f99a

Observation 7c282597-b668-4e7a-aae6-89acc8a195f0 · outbound

This paper cites write newline.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training write newline

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T17:19:29.257758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:19:29.257758Z digest=sha256:359e80e0da9327e874bb99743dd98fecd6f22b03812152a3202f1aeaedf683f1

Pith citing papers

Observation 480a76eb-61a5-447c-bacd-b121edb009f7 · inbound

GQA-{\mu}P: The maximal parameterization update for grouped query attention cites this paper.

GQA-{\mu}P: The maximal parameterization update for grouped query attention $\mu$nit Scaling: Simple and Scalable FP8 LLM Training

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:37:39.801944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:35:40.231293Z digest=sha256:4fdefcc446bbda982d78ac655e706fadbf4bed1d77588cdf511ca7c1aceef3f8