Pith. sign in

Paper Citation Record · LEDGER

MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2402.15627.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.15627 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T00:27:31.152088Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

24
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c9623257-d2ec-477c-912e-e570e9abb52d · inbound

PipeFusion: Patch-level Pipeline Parallelism for Diffusion Transformers Inference cites this paper.

PipeFusion: Patch-level Pipeline Parallelism for Diffusion Transformers Inference MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-24T01:18:42.456861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T01:17:11.261301Z digest=sha256:e3afe7dc54a87a88048cce0e48e9ceab9854781f02ab675d1af022d106ec132b

Observation 1e7d53a5-b4b8-4a87-ba54-3f17715a0d3d · inbound

HybridFlow: A Flexible and Efficient RLHF Framework cites this paper.

HybridFlow: A Flexible and Efficient RLHF Framework MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:53:38.974569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T07:53:38.715353Z digest=sha256:26aeff2e28b5b4f98c98dbd3b7e66577cd9e987fb93ac85e185b58ce864f138e

Observation fd958756-76f8-458d-9489-a4cfacf4dcc3 · inbound

InfiniteHBD: Building Datacenter-Scale High-Bandwidth Domain for LLM with Optical Circuit Switching Transceivers cites this paper.

InfiniteHBD: Building Datacenter-Scale High-Bandwidth Domain for LLM with Optical Circuit Switching Transceivers MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T00:27:31.152088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T00:27:31.152088Z digest=sha256:d47e3c24a79341e9c890e11ffb12ae4e7c95d0aff684dc7a7bde024a8b6ff724

Observation 5efa6d26-87f3-48ad-9a39-2f07ad1c15a1 · inbound

Goku: Flow Based Video Generative Foundation Models cites this paper.

Goku: Flow Based Video Generative Foundation Models MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T21:07:32.315237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T21:07:32.315237Z digest=sha256:702f20716cbe280e88db0df1018282b7b7d6e25f29cabeb5900c2de7fdf8838d

Observation 3096847c-53a4-4aa4-b1ac-49ce794eef91 · inbound

DeepCEE: Efficient Cross-Region Model Distributed Training System under Heterogeneous GPUs and Networks cites this paper.

DeepCEE: Efficient Cross-Region Model Distributed Training System under Heterogeneous GPUs and Networks MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:10.526279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:10.526279Z digest=sha256:56a4015505b56fff1b862637f5338783242a7e8fe6c000212693f5b4fd2bb09c

Observation 8c2bf174-ace9-4cdd-a052-9d1b2be0ea38 · inbound

Evolving HPC services to enable ML workloads on HPE Cray EX cites this paper.

Evolving HPC services to enable ML workloads on HPE Cray EX MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:14.380662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:14.380662Z digest=sha256:b6be1629c58685f1a72789067ed85f2fd8393f84901503dc56b68bb4655dd3ba

Observation 9b54b7f8-9df3-4bc9-922d-d99dbeffcb90 · inbound

BlueLM-2.5-3B Technical Report cites this paper.

BlueLM-2.5-3B Technical Report MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:20:53.803437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:20:53.803437Z digest=sha256:bcfe9befbe3462d0fc4a156b193a82321707af5461bbaea6e8c7956730f476ce

Observation 8e8bd21b-d892-4068-8334-cff07a1e3c6d · inbound

Towards Experiment Execution in Support of Community Benchmark Workflows for HPC cites this paper.

Towards Experiment Execution in Support of Community Benchmark Workflows for HPC MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T11:57:01.397736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:57:01.397736Z digest=sha256:3c3065d368d6ecf5f285577373ee5bf2ed6f5d6c61b9104d8738875911ae0c75

Observation a0ac690c-53b7-4fe9-a033-64c61ee2c191 · inbound

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection cites this paper.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:41:50.682623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:1b58a5466b35537f016dd3b3f97c072f2fa0c7284377c2028ba0be5423361351

Observation e6b0403f-0105-4ae9-9b66-2307729c82ea · inbound

TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM Training cites this paper.

TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM Training MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:06:20.803716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T01:36:41.804171Z digest=sha256:7bd9ba9c715bc75fe3c43ff13f691dd08d2c16197f9864cea011790d0be3ee6d

Observation 1419dc6c-744b-40fa-b6f4-b5082c8d4f49 · inbound

MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production cites this paper.

MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:06:14.978055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T02:04:07.344134Z digest=sha256:0b18c29d498d75bdb8d3b3724d4198c387a1f44ff344c414369d7e4cf94045f9

Observation 5d072b5f-85c8-4d9e-b9b7-2b701f313c91 · inbound

Instant GPU Efficiency Visibility at Fleet Scale cites this paper.

Instant GPU Efficiency Visibility at Fleet Scale MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:43:54.770415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T02:43:09.077121Z digest=sha256:e98d93a9ee863e1ad7ee18458446e03f2926ad7e1fcc35205c7a6d1649928baf

Observation 5bbf6308-359d-4b6c-aba3-505a1ec8a955 · inbound

The Cost and Network Limits of Space-Based AI Compute cites this paper.

The Cost and Network Limits of Space-Based AI Compute MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T04:42:46.167810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:42:46.167810Z digest=sha256:2a052da6403cb68ecba1d624542f309576fad72b7739d2ef3d54887d1f680999