Pith. sign in

Paper Citation Record · LEDGER

Recipes for Pre-training LLMs with MXFP8

As of 11 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 9 inbound Pith citation observations for arXiv:2506.08027.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08027 v2

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:12:47.881130Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:54:30.832174Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:26:13.675967Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved30
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e77c830a-1820-49fd-ae46-2e13d579e519 · outbound

This paper cites Ocp microscaling (mx) specification.

Recipes for Pre-training LLMs with MXFP8 Ocp microscaling (mx) specification

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:49.996595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:12:45.489420Z digest=sha256:97d420c8e9485f12f25ea86dcab6e802782effdf769ae8878d1a6e9f7de2a80b

Observation d2778806-4b3d-4b72-8898-02fb0cef4683 · outbound

This paper cites URL https://resources.nvidia.com/ en-us-blackwell-architecture.

Recipes for Pre-training LLMs with MXFP8 URL https://resources.nvidia.com/ en-us-blackwell-architecture

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:49.848468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:12:45.576876Z digest=sha256:8168e992197f79163c06614a40ea6141595c56c638c92c766f196e2d57d13ef5

Observation ea954705-fd04-42aa-92d2-feba37489914 · outbound

This paper cites Microscaling Data Formats for Deep Learning.

Recipes for Pre-training LLMs with MXFP8 Microscaling Data Formats for Deep Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:45.790981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:45.790981Z digest=sha256:7ee89f7630c410782b1f03395aaa6615d5833ab43d8902c73e3767fbde378a0a

Observation c91acf53-9320-43cb-890b-4de84268977c · outbound

This paper cites With Shared Microexponents, A Little Shifting Goes a Long Way.

Recipes for Pre-training LLMs with MXFP8 With Shared Microexponents, A Little Shifting Goes a Long Way

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.032229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.032229Z digest=sha256:192d183a255701644a0d0595bc169a80765551cf4a258136caac54706369b7ed

Observation 978acedf-b124-4ddd-9747-37dc484cf5b0 · outbound

This paper cites VS-Quant: Per-vector Scaled Quantization for Accurate Low-Precision Neural Network Inference.

Recipes for Pre-training LLMs with MXFP8 VS-Quant: Per-vector Scaled Quantization for Accurate Low-Precision Neural Network Inference

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.101980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.101980Z digest=sha256:454e2b59aa641f7e2e7e0f8f4f9384b2c3cae815d0b2f142153d9f9f93e02635

Observation 5004dd90-74fe-4df4-9bd4-241d6565ebf4 · outbound

This paper cites IEEE Std 754-2008 , pages 1–70, 2008.

Recipes for Pre-training LLMs with MXFP8 IEEE Std 754-2008 , pages 1–70, 2008

Reference 6

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T12:12:48.576575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:12:46.111933Z digest=sha256:c0da8be5a02c7504d21baee43274f46bd6b8951f322dab8a3591195e9ff09dfb

Observation 97beda48-3844-43fc-b362-5c5cab22791b · outbound

This paper cites FP8 Formats for Deep Learning.

Recipes for Pre-training LLMs with MXFP8 FP8 Formats for Deep Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.118074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.118074Z digest=sha256:79e900041c2e2ad385b406f954b6f7f4e74ff90e6c7e0e91210fe7fd6a5b6ab0

Observation f0aed45f-126b-44e2-9d42-acb6f5b6ea54 · outbound

This paper cites Nemotron-4 15B Technical Report.

Recipes for Pre-training LLMs with MXFP8 Nemotron-4 15B Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.131912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.131912Z digest=sha256:8c7d50c0720927e3d39eae136e51975b24470e4b83f05c4c634ce1b82c9fb2cc

Observation b5193df3-7e9b-400a-9b83-9666c5fe7341 · outbound

This paper cites Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models.

Recipes for Pre-training LLMs with MXFP8 Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.195871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.195871Z digest=sha256:4b2867f48cca54802470a48e6d2fef03dcd1463b4a40f9a0031a4ac0e08717ec

Observation c911db37-3acf-4e46-ad04-4fe11bdbcc4d · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Recipes for Pre-training LLMs with MXFP8 Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.269897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.269897Z digest=sha256:0de9a87bcaabb03eec21b02a36a477b9331d65ec7ad7f8766571ee13d946ae8c

Observation 13320e44-8583-430e-80bf-3fb3072fbb25 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Recipes for Pre-training LLMs with MXFP8 Measuring Massive Multitask Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.342345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.342345Z digest=sha256:e404987c2ba6a2ded63410ad72fe0d65e28714ee268e87af6ca67883609c1a44

Observation d51b2a97-8a5a-4e00-acf8-e89e75597ef9 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Recipes for Pre-training LLMs with MXFP8 Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.423646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.423646Z digest=sha256:4b1cdc5d8a5ee194b9652233993467e942a10365b1d017cc7f2135b2916fba0b

Observation aa12abb8-26b2-4213-9838-7d379e44d8e6 · outbound

This paper cites RACE: Large-scale ReAding Comprehension Dataset From Examinations.

Recipes for Pre-training LLMs with MXFP8 RACE: Large-scale ReAding Comprehension Dataset From Examinations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.493873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.493873Z digest=sha256:ed8795f2c8875d5fb4e83bce1477e68dcac14d79cda70c7fa4cefdf77fb7d7bf

Observation eb1c1a33-5e07-4ef4-b502-57f205bd6a5f · outbound

This paper cites PIQA: Reasoning about Physical Commonsense in Natural Language.

Recipes for Pre-training LLMs with MXFP8 PIQA: Reasoning about Physical Commonsense in Natural Language

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.582966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.582966Z digest=sha256:d6007053310e4db8598cb9fdafeb564b1e4843368bb5dadda50e1d271a745875

Observation 576341e0-8c74-421d-84d1-797c1d9cb0b9 · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale, 2019.

Recipes for Pre-training LLMs with MXFP8 Winogrande: An adversarial winograd schema challenge at scale, 2019

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.617897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.617897Z digest=sha256:840f5b35bd557f914387a1486680376dcddd33df49a23b7401dc48ded700629c

Observation ac0d46a1-cc6d-418c-a915-b00240ae8c9c · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Recipes for Pre-training LLMs with MXFP8 HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.646932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.646932Z digest=sha256:1ad601c4b39b2be44db4b6892465110c851b95c21b5fd228861cbf036a3e7ff6

Observation 8158c597-b84e-42aa-b16f-721c584852f3 · outbound

This paper cites Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering.

Recipes for Pre-training LLMs with MXFP8 Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.664591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.664591Z digest=sha256:922deb236edb27253c1cfd9a957acb6e06985a903c0d758418643d830babe80c

Observation 47481e3d-4ead-439e-b2a9-1a932ff3223a · outbound

This paper cites Socialiqa: Com- monsense reasoning about social interactions, 2019.

Recipes for Pre-training LLMs with MXFP8 Socialiqa: Com- monsense reasoning about social interactions, 2019

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:49.697169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:12:46.684428Z digest=sha256:70d33271ebee54ff62381311c140126bbe2879a4ba98a92cb5ad5333256027e7

Observation 27417057-aa1c-4011-b874-ef31b5cef6b2 · outbound

This paper cites CommonsenseQA: A question answering challenge targeting commonsense knowledge.

Recipes for Pre-training LLMs with MXFP8 CommonsenseQA: A question answering challenge targeting commonsense knowledge

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.703553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.703553Z digest=sha256:6bf08302aaddf92145b3f6d104c7119028163ff8938e60c7545c2480c1e9449f

Observation 4324b597-0cae-48ec-97f6-5ad225c24c3c · outbound

This paper cites The Llama 3 Herd of Models.

Recipes for Pre-training LLMs with MXFP8 The Llama 3 Herd of Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.730511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.730511Z digest=sha256:4a1486238aff2a7b5365b3781873180d1fc8162b1a8c57e3361c100294005355

Observation 8c86409a-67f0-46ea-a62c-353d2b508dcc · outbound

This paper cites 8-bit numerical formats for deep neural networks, 2022.

Recipes for Pre-training LLMs with MXFP8 8-bit numerical formats for deep neural networks, 2022

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:49.576743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:12:46.736164Z digest=sha256:8a1d18cbab29a87ab53975b5edcadc1bbe5e40ed9c715a0c5f65b2a8c57a1e25

Observation c7d49f70-6b50-4e37-9bf3-15da04787a62 · outbound

This paper cites Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J.

Recipes for Pre-training LLMs with MXFP8 Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.754388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.754388Z digest=sha256:eb7c155ea8b7c9dadbcbb4f4ea1557eb2f949b1bc378d91e7f5b7e25797ab5c3

Observation 8abe41fd-7c90-4bf4-af27-f82a37759078 · outbound

This paper cites Ocp 8-bit floating point specification (ofp8).

Recipes for Pre-training LLMs with MXFP8 Ocp 8-bit floating point specification (ofp8)

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:49.395984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:12:46.890232Z digest=sha256:ff3f8ca6e21e25b2b7b229bc97f8e99dab76cb15598abae11589f3e792ba35a3

Observation 073bdb9f-4131-4710-9eb0-861788e9ff79 · outbound

This paper cites Smoothquant: Accurate and efficient post-training quantization for large language models,.

Recipes for Pre-training LLMs with MXFP8 Smoothquant: Accurate and efficient post-training quantization for large language models,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:49.253153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:12:46.946348Z digest=sha256:7bae5e1cadd15e21e267df8393552bb611fa40f407411bcf34181ae37bdee80d

Observation 97cacf03-4797-4ebe-bce0-b18e0910712d · outbound

This paper cites QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving.

Recipes for Pre-training LLMs with MXFP8 QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:47.074402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:47.074402Z digest=sha256:943f5a36125d06ac6b91c9dd7f428696fce9ffcc8bdd9a4d571965f0dae5368f

Observation 28f760c5-22f8-4eb8-bbd2-b581dedb1e4c · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

Recipes for Pre-training LLMs with MXFP8 GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:47.159556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:47.159556Z digest=sha256:383fa82946da4f0e4634d01550a96257fb62d2a3f0de4ea8f068dbe328b23519

Observation 5a0ae3d2-6aab-4a30-8b62-7852f67537a9 · outbound

This paper cites AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration.

Recipes for Pre-training LLMs with MXFP8 AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:47.221979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:47.221979Z digest=sha256:2e590f638d0def70f1bcecf8431feb4b9d6ccc4958fd256f03349b29b69a1c38

Observation 6873f7de-f463-4129-8d2a-ad77c575243b · outbound

This paper cites Scaling FP8 training to trillion-token LLMs.

Recipes for Pre-training LLMs with MXFP8 Scaling FP8 training to trillion-token LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:47.282884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:47.282884Z digest=sha256:847e8d9f6752003df9a5b5275c19fef1edd5cdf66f6615aedc54fd480680f4e8

Observation b808aec9-2b83-4fe6-b3cd-43066f20400f · outbound

This paper cites The llama 4 herd: The beginning of a new era of natively multimodal ai innova- tion.

Recipes for Pre-training LLMs with MXFP8 The llama 4 herd: The beginning of a new era of natively multimodal ai innova- tion

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:49.124436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:12:47.338504Z digest=sha256:b21e25665ee8a0116f36e9788c8c2e3b6244f57f8cd792729fdef73159c03d1b

Observation 2ad6dd0c-2c4d-450a-8a9c-19ee36aabc99 · outbound

This paper cites Training LLMs with MXFP4.

Recipes for Pre-training LLMs with MXFP8 Training LLMs with MXFP4

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:47.436064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:47.436064Z digest=sha256:220925d5745100a564efea89bd6cdbdfaf1add3654d16e71e600740950ba0f29

Observation faf32e08-be52-45c7-a1cb-f11c842a2909 · outbound

This paper cites Optimizing Large Language Model Training Using FP4 Quantization.

Recipes for Pre-training LLMs with MXFP8 Optimizing Large Language Model Training Using FP4 Quantization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:47.457958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:47.457958Z digest=sha256:1447cb10b81cd66c9980fcf151b7d54acf1f9bd26eb34dc16df48fa37cf376e1

Observation 40625447-23d6-4b9f-b3ca-81594a08b54c · outbound

This paper cites Transformer engine.

Recipes for Pre-training LLMs with MXFP8 Transformer engine

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:48.996727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:12:47.463802Z digest=sha256:d95b2c774376a380a1281499888c2b94273d2746fcf977bc6f8b48b6e101546b

Observation 0f338323-0bc3-41c1-b8b8-2a77fa8253e5 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Recipes for Pre-training LLMs with MXFP8 DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:47.507189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:47.507189Z digest=sha256:ef6430f14aae8eabbc629ccf0bb576ad9c594df098a609a409b15dbcd6dde98f

Observation cefd3c8d-4ce7-45c9-99ff-acb5919ea260 · outbound

This paper cites Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective.

Recipes for Pre-training LLMs with MXFP8 Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:47.569722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:47.569722Z digest=sha256:c88a500d8b8ae9d71aa37af04a959a6508133099adc11223c8fcad12d741e8cb

Observation caef3419-ee7f-44e3-ac4c-c9d34222be05 · outbound

This paper cites Nemotron-4 340B Technical Report.

Recipes for Pre-training LLMs with MXFP8 Nemotron-4 340B Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:47.627819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:47.627819Z digest=sha256:4174f9d1e2a9720bbf18e0528dd427f49e6906ad12aefb8d1106df3ddbf181f3

Observation be684594-0e13-4231-98fc-0ab4833e81eb · outbound

This paper cites an unresolved cited work.

Recipes for Pre-training LLMs with MXFP8 Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:12:48.896269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:12:47.734191Z digest=sha256:986d265ef1311e125a7cb0ada485e7ff88073895c9b8bea2e1ef405adedda67f

Observation 802c0213-cdba-4429-bcf3-8cbef9468f7f · outbound

This paper cites an unresolved cited work.

Recipes for Pre-training LLMs with MXFP8 Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:12:48.785664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:12:47.807734Z digest=sha256:36c91578cf4262f375b622854f22a4444571168e5328429694e90a40e7a858bb

Observation 91e9d15e-83dd-480e-89f7-a4fca49d9d1b · outbound

This paper cites By construction amax/destmax never exceeds 2127 (which is the largest value representable in UE8M0) with FP8, FP6 or FP4 formats.

Recipes for Pre-training LLMs with MXFP8 By construction amax/destmax never exceeds 2127 (which is the largest value representable in UE8M0) with FP8, FP6 or FP4 formats

Reference 40

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:12:48.698298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T12:12:47.881130Z digest=sha256:ee8f0d8a84d2dfa1a4dd8dbcc6b53f7e0da9eeed2cb46f839c50a31542d01f28

Observation ff2704d2-c5dc-4343-a25d-6be97f90f558 · outbound

This paper cites SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models.

Recipes for Pre-training LLMs with MXFP8 SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:47.016932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:47.016932Z digest=sha256:3c8dd91b16c5717f71a7171b94c49b58bfdc6e2e776ee8708bce5414910d7256

Observation 80d5696e-8679-4c52-a166-52d569aefcda · outbound

This paper cites DeepSeek-V3 Technical Report.

Recipes for Pre-training LLMs with MXFP8 DeepSeek-V3 Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.816128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.816128Z digest=sha256:f786ea5ac5d4ad70ca2e9094fb07ae8c7df7b2662a6232d67c34f5d09f62dcc4

Pith citing papers

Observation d238393d-c84c-4717-8e97-4a6fc1071923 · inbound

A Comprehensive FP8 Training Recipe for Reasoning-Enhanced Language Models cites this paper.

A Comprehensive FP8 Training Recipe for Reasoning-Enhanced Language Models Recipes for Pre-training LLMs with MXFP8

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T14:51:29.915567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:51:29.915567Z digest=sha256:e61ff3f0d428a8be866ddb56b2cc3d672de161934231511e88eec47bc3c6c2be

Observation 886ff8dc-5d2d-408b-bfc2-f9ebd0f55fcf · inbound

Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling cites this paper.

Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling Recipes for Pre-training LLMs with MXFP8

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:23:52.769839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:23:01.845123Z digest=sha256:87f23244e24508059318d4f6084f2153519428e2052796a097c39cb1e78e5567

Observation 5bc29d3f-438b-4a33-b922-6dfd36cb0907 · inbound

StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models cites this paper.

StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models Recipes for Pre-training LLMs with MXFP8

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:15:22.229993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T12:10:44.802059Z digest=sha256:6b72224b48702184a4328908dec04b5185138af4f86a36868bd6abe271d6e71f

Observation a50e61fd-5ca4-4c24-8e7c-b3e0301cdfbb · inbound

OSP-Next: Efficient High-Quality Video Generation with Sparse Sequence Parallelism, HiF8 Quantization, and Reinforcement Learning cites this paper.

OSP-Next: Efficient High-Quality Video Generation with Sparse Sequence Parallelism, HiF8 Quantization, and Reinforcement Learning Recipes for Pre-training LLMs with MXFP8

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:13:27.553824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T13:06:02.201607Z digest=sha256:e923d993e522d8d09eec99d13b7ca12ec3ab10ddf0f6a13c7173867ae45008de

Observation 81579add-036b-46f1-b340-2027e20a9d29 · inbound

Stochastic Rounding Increases Small Singular Values cites this paper.

Stochastic Rounding Increases Small Singular Values Recipes for Pre-training LLMs with MXFP8

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:26:13.679444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T21:02:34.955394Z digest=sha256:79fb81527c59905f19039ae8740927c0b0794e80e99c242997cc5e64cec5fa46

Observation cc3f4a9d-f5f0-4d3c-9c6d-a1906fda7d6a · inbound

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales cites this paper.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Recipes for Pre-training LLMs with MXFP8

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:18.277482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:18.277482Z digest=sha256:bdb0e49cc1c6e92c25277516ae21ec1df862f318728c4bb0edaa51df78327660

Observation 35ce8b5d-d972-4296-9c31-5c3f40a2dbc8 · inbound

ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning cites this paper.

ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning Recipes for Pre-training LLMs with MXFP8

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-31T23:13:13.763909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:13:13.763909Z digest=sha256:7c5d8ba75715264fb07221655e8cbd7b1c8b8e8077cf8e73d898838b2f8de7c6

Observation 6bc22111-1ccd-450d-9be1-e93ba5b19779 · inbound

Stable FP4 Training via Transposition-Invariant Block Quantization cites this paper.

Stable FP4 Training via Transposition-Invariant Block Quantization Recipes for Pre-training LLMs with MXFP8

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-31T05:04:00.365982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:04:00.365982Z digest=sha256:24f7e2db24d1a9a6248c7a212d3ff6f5c9a7471f122341101c413a69bcb06cea

Observation 4524645d-eafa-4379-bd0b-c0b96d1facf0 · inbound

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference cites this paper.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Recipes for Pre-training LLMs with MXFP8

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.832174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.832174Z digest=sha256:751d2f3b944eaf2b1779ce1bccdd84f920a22b497cc396c0c1d2991b3e4ba8e0