Pith. sign in

Paper Citation Record · LEDGER

FP8-LM: Training FP8 Large Language Models

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2310.18313.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.18313 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:30:56.455931Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 763e5068-7994-42db-ac48-6d78988c0002 · inbound

NVILA: Efficient Frontier Visual Language Models cites this paper.

NVILA: Efficient Frontier Visual Language Models FP8-LM: Training FP8 Large Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.015165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:94506b6697771fda990449e2e6ddb3524a4e2a9519b603f22c6a4ab6916db680

Observation 15b4d8c2-04c8-4e6c-92e6-dc57740d77a2 · inbound

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache cites this paper.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache FP8-LM: Training FP8 Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.957098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.957098Z digest=sha256:c45dbe2d0ce5a88f360d536a06faa978b9d3690a76d02d443499f7c459687702

Observation 85021c8f-eaee-4d09-89cd-2ba63c29690b · inbound

Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities cites this paper.

Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities FP8-LM: Training FP8 Large Language Models

Reference 123

Resolution
unresolved
no resolver link, observed 2026-08-16T04:30:56.455931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:30:56.455931Z digest=sha256:c2dc2e295847593b8997bc00330a4d80512eeac3fe1966426edcd208534b197a

Observation fe5b0c0e-2db7-4bbd-8de5-f6fc9c49b100 · inbound

Resource-Efficient Language Models: Quantization for Fast and Accessible Inference cites this paper.

Resource-Efficient Language Models: Quantization for Fast and Accessible Inference FP8-LM: Training FP8 Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T21:55:07.817270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:55:07.817270Z digest=sha256:420cd75aaf8f38bdac1f74af5eda10e05f5aaeac0f43faaeb73ecebb4187b49f

Observation bb6d3808-fbf6-4c1d-acc2-5901fa13b8f7 · inbound

Gaussian Weight Sampling for Scalable, Efficient and Stable Pseudo-Quantization Training cites this paper.

Gaussian Weight Sampling for Scalable, Efficient and Stable Pseudo-Quantization Training FP8-LM: Training FP8 Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:33.088734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:33.088734Z digest=sha256:e3be3a6f6fbed4a572fd1f56b1c66edf83233d72f0719219ada6d9250f47b28e

Observation 8fce1a7a-1080-4413-84fb-cb17fd1a81d6 · inbound

Scaling Law for Quantization-Aware Training cites this paper.

Scaling Law for Quantization-Aware Training FP8-LM: Training FP8 Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:08.205307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:08.205307Z digest=sha256:fa2c99d465cff7e26435a28bd34948bdfc33a85375913ad94d19ed82d765a82b

Observation cadca501-3ecf-453d-89b4-51f13c108a8d · inbound

FP4 All the Way: Fully Quantized Training of LLMs cites this paper.

FP4 All the Way: Fully Quantized Training of LLMs FP8-LM: Training FP8 Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:38.586190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:38.586190Z digest=sha256:34c9fc47cd86f5a1f7f0c8f00da2a71e3f63f3176ed1f009192b0b4522267d37

Observation 0a7074d7-04b2-49b1-89f0-021c0ca4f351 · inbound

Log-Normal Multiplicative Dynamics for Stable Low-Precision Training of Large Networks cites this paper.

Log-Normal Multiplicative Dynamics for Stable Low-Precision Training of Large Networks FP8-LM: Training FP8 Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T19:09:44.432131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:09:44.432131Z digest=sha256:1a7c7ca53a50e99174ae1ea3dc8888d4e5ebf0eb45b5dc82885513de48f77a8d

Observation 75d247f4-3fc0-40e0-a2dd-0d17ef07c062 · inbound

Thunder-LLM: Efficiently Adapting LLMs to Korean with Minimal Resources cites this paper.

Thunder-LLM: Efficiently Adapting LLMs to Korean with Minimal Resources FP8-LM: Training FP8 Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T23:58:09.794501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:58:09.794501Z digest=sha256:1083d736873314467154c8a25e246a2583b1f4e6d21849567138bef9b84fd895

Observation d2a5babf-e9c6-4c60-bc28-991320d69f9b · inbound

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models cites this paper.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models FP8-LM: Training FP8 Large Language Models

Reference 125

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.329029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.329029Z digest=sha256:419ece3199d4cc6a216ad9f7f1ef91a9d3de90d56758be55778b549444006f1b

Observation 4e004260-9480-4950-98a4-5d9a0ca18c8b · inbound

Compute Requirements for Algorithmic Innovation in Frontier AI Models cites this paper.

Compute Requirements for Algorithmic Innovation in Frontier AI Models FP8-LM: Training FP8 Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:50.556519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:50.556519Z digest=sha256:44273f953fe27a72a87828a78065b8355df5e02ba94a3477e0c1525af72898e4

Observation 285f31be-c662-4247-a280-03a878174585 · inbound

Foundation Models for Clean Energy Forecasting: A Comprehensive Review cites this paper.

Foundation Models for Clean Energy Forecasting: A Comprehensive Review FP8-LM: Training FP8 Large Language Models

Reference 182

Resolution
unresolved
no resolver link, observed 2026-08-06T11:06:14.282566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:06:14.282566Z digest=sha256:9902183d533c8a087f799e34577586cf43d6338e9be7258c75b35afec82b7dcf

Observation c0fb46db-12d3-4cbe-93b1-181efd3dac45 · inbound

A Comprehensive FP8 Training Recipe for Reasoning-Enhanced Language Models cites this paper.

A Comprehensive FP8 Training Recipe for Reasoning-Enhanced Language Models FP8-LM: Training FP8 Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T14:51:29.942693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:51:29.942693Z digest=sha256:98b43738181f3284a8444493ee42e9231fb10b645cd010c0ecb88cb29bb20761

Observation 65c71bf6-d9cd-4002-9fb2-6a0f447ec470 · inbound

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention cites this paper.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention FP8-LM: Training FP8 Large Language Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T10:02:31.727622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:20f78bfb2e27199bd0c6d6a61a7b701a307ea5d9d737d4e85222f23798253ff3

Observation ef1bf59f-ab7a-4a94-9df9-3c917d86f672 · inbound

Beyond Sunk Costs: Boosting LLM Pre-training Efficiency via Orthogonal Growth of Mixture-of-Experts cites this paper.

Beyond Sunk Costs: Boosting LLM Pre-training Efficiency via Orthogonal Growth of Mixture-of-Experts FP8-LM: Training FP8 Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:40:36.473045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T20:36:22.974054Z digest=sha256:236e54c6ea6f27a485edb484d1f09f527e1fcd530598d8fc905c90eda01b7537

Observation 02604264-df50-490b-a011-9e1fbbb6c0c3 · inbound

BOOST: BOttleneck-Optimized Scalable Training Framework for Low-Rank Large Language Models cites this paper.

BOOST: BOttleneck-Optimized Scalable Training Framework for Low-Rank Large Language Models FP8-LM: Training FP8 Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:21:21.497637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T23:19:02.358348Z digest=sha256:1bf5d0ed1a2db840217c523e43b5b06629cc48b545a65487a7cd1f591abac79c

Observation 8ed3bdc4-b757-4ea0-bb14-1e4392b6b0dd · inbound

DynamiQ: Accelerating Gradient Synchronization using Compressed Multi-hop All-reduce cites this paper.

DynamiQ: Accelerating Gradient Synchronization using Compressed Multi-hop All-reduce FP8-LM: Training FP8 Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T03:13:23.907650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:13:23.907650Z digest=sha256:cf498d86f24e9ec67bc42ec62fb5cede2baa013357e510c0c0b240b6b582a1ea

Observation 58616300-6525-49dd-b99b-01fd55b1c059 · inbound

AdaHOP: Fast and Accurate Low-Precision Training via Outlier-Pattern-Aware Rotation cites this paper.

AdaHOP: Fast and Accurate Low-Precision Training via Outlier-Pattern-Aware Rotation FP8-LM: Training FP8 Large Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:53:15.729174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T20:51:39.641700Z digest=sha256:0a22baf5fb29f4644604e3f66f818c24cff59deb2bfe2717902f8254ef3161f3

Observation 45540f5b-6106-4a87-858e-470fccff2dcb · inbound

StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models cites this paper.

StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models FP8-LM: Training FP8 Large Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:15:22.211750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T12:10:44.802059Z digest=sha256:907451acb292fc222a2f5ae957d3987f0abdf71cbfb6f021362d20bb1dc9894c

Observation 3f491aad-732c-48fa-91d4-07731a328bf2 · inbound

AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMs cites this paper.

AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMs FP8-LM: Training FP8 Large Language Models

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:31:17.298359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-09T19:46:13.015064Z digest=sha256:97ef969356fb6e13334de310639a8bb16a6c0ab9aba1b4f6127db9d3e54c7209

Observation 594a3859-6f27-4b3e-a234-8c216e94e9c0 · inbound

AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMs cites this paper.

AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMs FP8-LM: Training FP8 Large Language Models

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:21:30.430366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T05:17:09.793360Z digest=sha256:b73804ae5b108b53f1fba9a6948d74612b6530204697c99d04cca884f65d8809

Observation 33d5b4f0-9487-4082-ae95-b199113bb720 · inbound

AIS: Adaptive Importance Sampling for Quantized RL cites this paper.

AIS: Adaptive Importance Sampling for Quantized RL FP8-LM: Training FP8 Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:14:52.614389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T03:13:14.384567Z digest=sha256:edcb0e157a31da950677d9d33a28680421b0ff432ff0feed5a7c401f076ceb49

Observation d985f985-09be-439a-bf12-7a044b18d526 · inbound

Max-Window Scale Estimation for Near-Lossless HiF8 W8A8 Quantization-Aware Training cites this paper.

Max-Window Scale Estimation for Near-Lossless HiF8 W8A8 Quantization-Aware Training FP8-LM: Training FP8 Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:04:00.895244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T23:02:56.306685Z digest=sha256:f79d171ae97118fcbdd010c2cd8d8857720e9ccf60a98014bb2a232f34a55432

Observation 4cfb0e83-6296-4c81-ac6e-30f89ffacec9 · inbound

Addressing Variable Heterogeneity in Distributed Multimodal Training with Entrain cites this paper.

Addressing Variable Heterogeneity in Distributed Multimodal Training with Entrain FP8-LM: Training FP8 Large Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:23:17.995056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T10:20:12.860927Z digest=sha256:7a3c8ecdc47616c40b5c1b0fbf4eec1f673748286a76bbc7bc0b01f65fbd292e

Observation fd2b5d8b-a31b-43ec-a486-9079da3e292c · inbound

GNMR: Runtime Stability Control for Low-Precision Large Language Model Training cites this paper.

GNMR: Runtime Stability Control for Low-Precision Large Language Model Training FP8-LM: Training FP8 Large Language Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:42:36.061916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T18:53:47.187437Z digest=sha256:b1abddb2819c18455d1b401e9a52088c58cba034f65fca86b41798d6dc8a6b8d

Observation a19828fe-eb46-46d2-a111-215fe384bfff · inbound

Achieving Cloud-Grade SLOs for Local Mixture-of-Experts Inference through CPU-GPU Hybrid Design cites this paper.

Achieving Cloud-Grade SLOs for Local Mixture-of-Experts Inference through CPU-GPU Hybrid Design FP8-LM: Training FP8 Large Language Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-03T07:27:44.410480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T12:06:46.806138Z digest=sha256:a7be4341e3f7b9ff2df997b0785f839bbf0d7b289c54d9f9b6ec53ac24c97fac

Observation 0a318c2e-7a83-4f3b-8fbf-d728b7ed2f33 · inbound

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining cites this paper.

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining FP8-LM: Training FP8 Large Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-09T09:16:06.500582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:0b263d6a327927c87b2f0f1cdc4a9eaeceefa8c1cc957e4d2ae02dad04df8633