Pith. sign in

Paper Citation Record · LEDGER

Scaling FP8 training to trillion-token LLMs

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2409.12517.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.12517 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:23:40.929720Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T07:27:44.431731Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c4b9be6e-0317-4925-8dc5-709c656c726e · inbound

NVILA: Efficient Frontier Visual Language Models cites this paper.

NVILA: Efficient Frontier Visual Language Models Scaling FP8 training to trillion-token LLMs

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.040130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:fccdb9421fede2d02c7c19da7d1c9f6ce0eb01764ae6ed434105ff725373bddf

Observation f804b551-985d-46b4-9dc7-9f50336ff2f0 · inbound

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture cites this paper.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture Scaling FP8 training to trillion-token LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:40.929720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:40.929720Z digest=sha256:af49dd92c79e776a3dcc942ff66b403780695b78fb035def00c96f556bc3b95f

Observation 67aba63f-2798-4df1-9f36-1e7f316667f5 · inbound

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache cites this paper.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Scaling FP8 training to trillion-token LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.922252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.922252Z digest=sha256:084ca4e626f1721a7dcf209854f0c357e5d787f9932ddd37222ac7f05e20f204

Observation 4cac8f33-ef05-41ca-aeeb-e6fc175a5a72 · inbound

Scaling Law for Quantization-Aware Training cites this paper.

Scaling Law for Quantization-Aware Training Scaling FP8 training to trillion-token LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:06.104682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:06.104682Z digest=sha256:0b8495900e480614abc5f18cf5005a86c5895ba8243da009f3b886d60b8eed1a

Observation f59b69c3-0296-407c-b2e7-12e45dd8d740 · inbound

FP4 All the Way: Fully Quantized Training of LLMs cites this paper.

FP4 All the Way: Fully Quantized Training of LLMs Scaling FP8 training to trillion-token LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:38.069250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:38.069250Z digest=sha256:fc9ff18b8e1c98de029042f0583c183403246db889664e99019dd044e16bfe06

Observation 6873f7de-f463-4129-8d2a-ad77c575243b · inbound

Recipes for Pre-training LLMs with MXFP8 cites this paper.

Recipes for Pre-training LLMs with MXFP8 Scaling FP8 training to trillion-token LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:47.282884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:47.282884Z digest=sha256:847e8d9f6752003df9a5b5275c19fef1edd5cdf66f6615aedc54fd480680f4e8

Observation 87602fef-a9ad-4628-a28a-e57900c9ba43 · inbound

Characterization and Mitigation of Training Instabilities in Microscaling Formats cites this paper.

Characterization and Mitigation of Training Instabilities in Microscaling Formats Scaling FP8 training to trillion-token LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:47:52.213010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:47:52.213010Z digest=sha256:5fe565d4546967899430276345a2e1cf5815e358616c3372aee94748648a066e

Observation d5b58414-1fa1-4c1b-b30d-a0aa36c878c4 · inbound

Thunder-LLM: Efficiently Adapting LLMs to Korean with Minimal Resources cites this paper.

Thunder-LLM: Efficiently Adapting LLMs to Korean with Minimal Resources Scaling FP8 training to trillion-token LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:58:09.687331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:58:09.687331Z digest=sha256:9e7d4fde1672b025ed92860c2c0df54d2938213467f8aef2d566a0f08b9c0249

Observation 0ef52389-2791-4e33-9958-2e8aac0f444e · inbound

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models cites this paper.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Scaling FP8 training to trillion-token LLMs

Reference 128

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.337409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.337409Z digest=sha256:dadb4bb765070d1d54f0a8e7e09b958850cf92c189723f8dfe09fef1974684cc

Observation a17c0e55-6fb2-471e-9476-971700957f1f · inbound

A Comprehensive FP8 Training Recipe for Reasoning-Enhanced Language Models cites this paper.

A Comprehensive FP8 Training Recipe for Reasoning-Enhanced Language Models Scaling FP8 training to trillion-token LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T14:51:29.281759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:51:29.281759Z digest=sha256:4b2fbc707b0a860f85db6f3cc0dbbe3f45efdc7e6f9ba02e7cfa03d5c12a58b6

Observation 2747e26a-663e-407e-9239-46432204b867 · inbound

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention cites this paper.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Scaling FP8 training to trillion-token LLMs

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:02:31.710568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:ed88921a8789539bcc956d36e4b4b75c86bf720dde00a178a7120df85c8605c0

Observation 16d8a70c-f9d1-4dae-b920-fb03d4795f40 · inbound

From Detection to Recovery: Operational Analysis on LLM Pre-training with 504 GPUs cites this paper.

From Detection to Recovery: Operational Analysis on LLM Pre-training with 504 GPUs Scaling FP8 training to trillion-token LLMs

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:11:15.382428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T02:10:43.006238Z digest=sha256:2cc83f757f238a0688f17de8664369e0aa32ecf22d002e5c47685a04d468b06d

Observation ceea8870-8af1-47ab-941f-aaac89e01fb9 · inbound

From Detection to Recovery: Operational Analysis on LLM Pre-training with 504 GPUs cites this paper.

From Detection to Recovery: Operational Analysis on LLM Pre-training with 504 GPUs Scaling FP8 training to trillion-token LLMs

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:35:46.637821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T22:57:31.085889Z digest=sha256:6579166d3c2c64b7facf2360d974eb94a8fb635df9794b5996df595c4808654c

Observation 8732e7ae-e736-45f3-b579-645a99dd585c · inbound

LoKA: Low-precision Kernel Applications for Recommendation Models At Scale cites this paper.

LoKA: Low-precision Kernel Applications for Recommendation Models At Scale Scaling FP8 training to trillion-token LLMs

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:06:28.075687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:33:41.411292Z digest=sha256:3fd6c664a3849865e667788487a22a1e86a29c03cc908687547664b01342d552

Observation aee8c9c7-516b-4276-869f-b4e7a917cc26 · inbound

LoKA: Low-precision Kernel Applications for Recommendation Models At Scale cites this paper.

LoKA: Low-precision Kernel Applications for Recommendation Models At Scale Scaling FP8 training to trillion-token LLMs

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:59:46.257916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T04:55:01.973832Z digest=sha256:b22964a15c3428ee692e12f5895f7a032bdf65790eca4dc88d28d2de28630e1d

Observation ec427814-b049-427b-889c-79e1561f93b9 · inbound

Expand More, Shrink Less: Shaping Effective-Rank Dynamics for Dense Scaling in Recommendation cites this paper.

Expand More, Shrink Less: Shaping Effective-Rank Dynamics for Dense Scaling in Recommendation Scaling FP8 training to trillion-token LLMs

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:30:22.650040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:29:49.474591Z digest=sha256:b1158dfdffbe6a8f4bbc604d5e6852ef39dc9e19d62cb58674e358a70ba24fc0

Observation 04e1bcac-80e3-4b73-b29a-26833937f021 · inbound

Achieving Cloud-Grade SLOs for Local Mixture-of-Experts Inference through CPU-GPU Hybrid Design cites this paper.

Achieving Cloud-Grade SLOs for Local Mixture-of-Experts Inference through CPU-GPU Hybrid Design Scaling FP8 training to trillion-token LLMs

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T07:27:44.433192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T12:06:46.806138Z digest=sha256:065ffccfc5fd1b02562b74b54756790e9266ae5e70131d669862008174bcd5ad

Observation 7ae5b301-63ae-49e1-8d4e-0c59f242f88d · inbound

Full-Stack FP4: Stable LLM Pretraining with Quantized Projections, Optimizers, and Attention cites this paper.

Full-Stack FP4: Stable LLM Pretraining with Quantized Projections, Optimizers, and Attention Scaling FP8 training to trillion-token LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-11T19:17:59.044982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:17:59.044982Z digest=sha256:8ad643914983e484dec448992cb6cfd471b019de5b991e2fec4e167d5f6b4426

Observation fc699b5a-1423-467d-977a-07cd651a73b9 · inbound

One QK Channel, Many Sources: Guarding Low-Precision Attention Collapse cites this paper.

One QK Channel, Many Sources: Guarding Low-Precision Attention Collapse Scaling FP8 training to trillion-token LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T15:20:15.926212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T15:20:15.926212Z digest=sha256:24e844925e655292553e59ccdf312afc685e720aca93577f4a4d3762dcdc59a8