Pith. sign in

Paper Citation Record · LEDGER

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training

As of 10 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 1 inbound Pith citation observation for arXiv:2502.05967.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05967 v3

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:19:29.257758Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-19T16:35:40.231293Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T16:37:39.800519Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 96ddc2c1-2900-4902-9aa8-8785742dc3e0 · outbound

This paper cites Scaling FP 8 training to trillion-token LLM s.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training Scaling FP 8 training to trillion-token LLM s

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.723781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.124327Z digest=sha256:46699bd3ca8eea9687811b9202fa32a0ceb3b71065610b6df8cc9c34d5ac01b4

Observation be115354-0bf3-490b-8643-6dd7d0517722 · outbound

This paper cites Calibrating the Mosaic evaluation Gauntlet , 4 2024.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training Calibrating the Mosaic evaluation Gauntlet , 4 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.713434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.128906Z digest=sha256:36efb53087a237f6fec70444de0c06aebd129dd0f7357c3c11fa7757af5ad0ae

Observation 1ae901c9-233d-4631-bc4f-e16d4ce3175f · outbound

This paper cites Unit scaling: Out-of-the-box low-precision training.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training Unit scaling: Out-of-the-box low-precision training

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.703031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.132981Z digest=sha256:349f06bfc5f4a3a0878d085f871b21be505234b2597df3cf2ce83e06c09f7efa

Observation 4b7d9766-8df9-4506-8ef9-042a19fff08b · outbound

This paper cites Y., Deiseroth, B., Cruz-Salinas, A.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training Y., Deiseroth, B., Cruz-Salinas, A

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.691447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.137768Z digest=sha256:239a99601515bf3ff1e861b98e73c0c7422ce74923cef44a61439e33b3428ccb

Observation e5c59901-c6f5-45d5-a2d8-1fa0f4027f2a · outbound

This paper cites and Berger, R.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training and Berger, R

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.681667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.147562Z digest=sha256:18710b1da4e6e380e6a0198e336067e6d00f74c8402c1c2178e973d0df235b56

Observation 63e97931-b0fe-4eb0-b0b5-cdfa73cab07c · outbound

This paper cites an unresolved cited work.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-08T17:19:29.670824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.151332Z digest=sha256:8f52b9032ad3c65f14f182612fca16c11e57644bfdc24f9e7a7b803715d7786c

Observation 83163bd4-76ba-4dc6-822d-0ef0c3b474a1 · outbound

This paper cites LLM .int8(): 8-bit matrix multiplication for transformers at scale.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training LLM .int8(): 8-bit matrix multiplication for transformers at scale

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.660262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.156732Z digest=sha256:d37b7c7461f88a80c0ad5dccbff125670c131ac49346fe951ebbc2f0244be0e3

Observation c5b080e9-8866-464e-9a1e-05eae69138c8 · outbound

This paper cites Blazingly fast LLM evaluation for in-context learning, 2 2023.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training Blazingly fast LLM evaluation for in-context learning, 2 2023

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.649863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.161453Z digest=sha256:3e8e2f8cb976f57560ecf30e49c3d3f93c3a532822528c2c1921ad0d9dd9d89e

Observation 131c9f01-2a4d-4373-bdc1-e7c3d7d42eef · outbound

This paper cites Flex Attention: A Programming Model for Generating Optimized Attention Kernels.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training Flex Attention: A Programming Model for Generating Optimized Attention Kernels

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T17:19:29.164938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:19:29.164938Z digest=sha256:29e9d7248a4ff5804e0b1de74f47dcbc96f579079da405a2ee41055b8f67ca90

Observation 5b033ee9-601f-4d05-bac6-400ec649f6a7 · outbound

This paper cites FBGEMM: Enabling High-Performance Low-Precision Deep Learning Inference.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training FBGEMM: Enabling High-Performance Low-Precision Deep Learning Inference

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T17:19:29.168874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:19:29.168874Z digest=sha256:0d24043eb3ea076413ec9d5d1e4e6fbfc1dfd95591cffc24549ef5325eeb961d

Observation c2caf857-8d98-465e-843d-2c3babedd853 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training Adam: A Method for Stochastic Optimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T17:19:29.172567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:19:29.172567Z digest=sha256:86c5fe9eb33ece680f43f7c6f69465d4cab1d5cfc0026c0adf40d674853ae049

Observation 71dac2fd-01c7-4d32-a163-a40994b02812 · outbound

This paper cites An Empirical Study of $\mu$P Learning Rate Transfer.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training An Empirical Study of $\mu$P Learning Rate Transfer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T17:19:29.177838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:19:29.177838Z digest=sha256:0d02e23b518ec683e796fd198c7fbeeb21ee90f819efc850e2ef1d3c61663e93

Observation 38016c2e-d8c8-4112-9998-e32ca5ea3471 · outbound

This paper cites Swin transformer v2: Scaling up capacity and resolution.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training Swin transformer v2: Scaling up capacity and resolution

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T17:19:29.182203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:19:29.182203Z digest=sha256:25ca50bf15674b22f9a422509c8b12319c1f7ad48d6cc73732691f23ca21bae8

Observation aa95d758-2d0a-41e1-88ab-269b8b08569e · outbound

This paper cites Mixed precision training.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training Mixed precision training

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.632410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.185719Z digest=sha256:a37a5c56ef6098cf751ee737f6563694a56f618cffa911074e512e6ec447482a

Observation 1ae13680-7316-4c76-bd0e-eeb389e9106d · outbound

This paper cites FP8 Formats for Deep Learning.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training FP8 Formats for Deep Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T17:19:29.189352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:19:29.189352Z digest=sha256:90f1470dc6fc647b60e614c1356c16b8b31437640424646f8e2ca95dd967d61f

Observation 58904f1e-1df3-4abd-b755-600b094d9e95 · outbound

This paper cites I., Alizadeh-Vahid, K., Mehta, S., del Mundo, C.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training I., Alizadeh-Vahid, K., Mehta, S., del Mundo, C

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.621493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.193163Z digest=sha256:882fb51a2697e7bf55a3ac3b2f1c990837b5a84cc08c8bfa99f19d446361a5b3

Observation 3f90a058-88e0-48f3-8eff-0275d559a572 · outbound

This paper cites Composer.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training Composer

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.610565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.196804Z digest=sha256:b7709410505acb6bd9b474fd0d17a3e7e3b6a465e16fe28c76e00573c9432444

Observation 3d7f6d4f-077d-4fb7-b60d-eb2a0c67a44e · outbound

This paper cites LLM F oundry.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training LLM F oundry

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.599524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.200330Z digest=sha256:d06ef753d9514e6bcee5d2b8344e777bcc2dd37d47748c4bc81cd7851d200e12

Observation 80a15c84-ba1c-4aa0-aba8-25a7aef724fb · outbound

This paper cites Streaming.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training Streaming

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.588442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.203719Z digest=sha256:47a676695cbabf400c77627ca8fd6d7d3c7f216cc7057f9524ff2944a8e96768

Observation 55d2c8bb-f9de-4b44-b4fb-ddc3c061e630 · outbound

This paper cites Asynchronous multiply-and-accumulate instruction: wgmma.mma\_async.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training Asynchronous multiply-and-accumulate instruction: wgmma.mma\_async

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.577873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.208054Z digest=sha256:018c23a02a3df4dca1fc1aef851aa480224a28989894d91eb4265ebca46ab071

Observation c417dd04-99ef-49cd-92b8-c135a238169d · outbound

This paper cites Transformer E ngine, 2023.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training Transformer E ngine, 2023

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.567179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.212306Z digest=sha256:75cbd5aaa384bafa2be142b39b3c519755b122cd2b9f71079a8a7d7ab7c39aeb

Observation 4404d4ff-b480-467d-903c-1de8557f7243 · outbound

This paper cites cuBLAS : cublasLtMatmul().

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training cuBLAS : cublasLtMatmul()

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.556156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.216823Z digest=sha256:d17a9e5809452ced4ad4b78d456eaf652608f3bc97d8179ce80e310d901c49e0

Observation f9c9cec1-849b-4b84-ae3a-4f0d1c8534cc · outbound

This paper cites 2 OLMo 2 Furious.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training 2 OLMo 2 Furious

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T17:19:29.220412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:19:29.220412Z digest=sha256:42180383522f03903d8b4ae724eb15ebd205831cdac7870cebbfb1efc6bbcdc1

Observation 4d7e4bda-9da5-49cc-993a-5c765416cdb3 · outbound

This paper cites V., Cui, X., Zhang, W., and Gopalakrishnan, K.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training V., Cui, X., Zhang, W., and Gopalakrishnan, K

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.544769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.224893Z digest=sha256:b3b96eb78dddda98ac12e6afe376d05dd12f038a3340c990a90365ca16a68b8c

Observation 9235daec-c6e0-4fe5-befc-e41a88f096f5 · outbound

This paper cites T., and Cox, D.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training T., and Cox, D

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T17:19:29.229551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:19:29.229551Z digest=sha256:b608c97112afcabe7e4cc6d6ad0ffb62a1bf398d4c064e928debc7eafb7dc220

Observation 2247b83d-277f-4b2f-b595-915f8d23d1e0 · outbound

This paper cites N., Kaiser, L.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training N., Kaiser, L

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T17:19:29.233664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:19:29.233664Z digest=sha256:d06c2ab909415dd2246e9b0a79ab71aa81ab1a518cdfdae8694007a7c7af1ad4

Observation 90af06df-b171-4857-bc55-902473b9d5df · outbound

This paper cites How to set AdamW's weight decay as you scale model and dataset size.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training How to set AdamW's weight decay as you scale model and dataset size

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T17:19:29.237555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:19:29.237555Z digest=sha256:9ecf4309527db2d73419e69daa931ba642430d4807860d2f077d57320ec50ce7

Observation 5990abfe-fed8-4212-8098-42e725d1572a · outbound

This paper cites J., Xiao, L., Everett, K.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training J., Xiao, L., Everett, K

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.526729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.241813Z digest=sha256:6c813e21757aea8769f5b25c505411435868edc564715646c702cb3374f4693c

Observation f43f067b-788c-4ce3-b69e-e1547344345b · outbound

This paper cites J., Babuschkin, I., Sidor, S., Liu, X., Farhi, D., Ryder, N., Pachocki, J., Chen, W., and Gao, J.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training J., Babuschkin, I., Sidor, S., Liu, X., Farhi, D., Ryder, N., Pachocki, J., Chen, W., and Gao, J

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.515309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.245461Z digest=sha256:94a93fcc33b7684050b8f2a7373748729ce3c615f49ca0071c5be2ac2bd2cba9

Observation b974ecab-14f8-4001-bdb9-dc2704304119 · outbound

This paper cites A Spectral Condition for Feature Learning.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training A Spectral Condition for Feature Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T17:19:29.249207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:19:29.249207Z digest=sha256:d410e64d093cdf15a9fbe454a7b5d75c03b375f6e7b7ca54952eebd04f917e90

Observation 5c200309-8b24-494a-9c8d-de876a55f057 · outbound

This paper cites Tensor programs VI : Feature learning in infinite depth neural networks.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training Tensor programs VI : Feature learning in infinite depth neural networks

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:19:29.504154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T17:19:29.253110Z digest=sha256:6cf0edf74ffab605bfc68b7bf90949e19994716557da7670e6aee282e01ad46e

Observation 7c282597-b668-4e7a-aae6-89acc8a195f0 · outbound

This paper cites write newline.

$\mu$nit Scaling: Simple and Scalable FP8 LLM Training write newline

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T17:19:29.257758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:19:29.257758Z digest=sha256:359e80e0da9327e874bb99743dd98fecd6f22b03812152a3202f1aeaedf683f1

Pith citing papers

Observation 480a76eb-61a5-447c-bacd-b121edb009f7 · inbound

GQA-{\mu}P: The maximal parameterization update for grouped query attention cites this paper.

GQA-{\mu}P: The maximal parameterization update for grouped query attention $\mu$nit Scaling: Simple and Scalable FP8 LLM Training

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:37:39.801944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T16:35:40.231293Z digest=sha256:db0beb42825f3b81c51b0fa3fae9be741da2463a49547a0efee74fe890e53300