Pith. sign in

Paper Citation Record · LEDGER

Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 49 inbound Pith citation observations for arXiv:2104.04473.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2104.04473 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 49 of 49 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:17:19.030628Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

9
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a4d62a99-e85b-4e0d-86d5-17278d726ffe · inbound

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model cites this paper.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-24T12:14:26.541851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:be42a51b537443c123275e966a3391fb489d4d85a5b2a4ff8ae27f8d4028de39

Observation 50c00f4f-b34c-4082-acab-2ff8b53dd1d2 · inbound

OPT: Open Pre-trained Transformer Language Models cites this paper.

OPT: Open Pre-trained Transformer Language Models Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 142

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T20:53:17.533584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T20:53:16.720145Z digest=sha256:694cfea5aeaec19934b87660ebf38be55a2cf544a366baf378e003685bb6151f

Observation 52a26e0c-be81-4f2f-9c4c-ef851eb31be0 · inbound

Multi-matrix Factorization Attention cites this paper.

Multi-matrix Factorization Attention Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.901462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.901462Z digest=sha256:3ea6a4d59e51cd44e2e463ebf94696ed099c024494ad09f1fff92360918ae957

Observation e815fc39-0ed3-46ec-a5df-920a95bf6b17 · inbound

Automatically Planning Optimal Parallel Strategy for Large Language Models cites this paper.

Automatically Planning Optimal Parallel Strategy for Large Language Models Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:59:59.146149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:59:59.146149Z digest=sha256:bbf0f0867ee2f78cc867448ca5a53aff9e0687b19df2b7fe53fdd6e03cd19506

Observation bcbdff8a-f26d-4bc1-b441-fb0214fe0104 · inbound

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning cites this paper.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.123613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.123613Z digest=sha256:cba9e78e62383fe9800cb99f9bda5d53e4d23a3cc748eefe1b29b401ea8876aa

Observation 9da226f6-618d-430e-9a3f-1de9a9a1b1f6 · inbound

Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass cites this paper.

Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:20.118465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:20.118465Z digest=sha256:48ce10a9f4d390eb6af034a6a212b373a5d09450f055fca883c8e7a98384c4c1

Observation fb0670d6-8efc-4d2e-b247-9226f418704c · inbound

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers cites this paper.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:02.642128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:20:02.642128Z digest=sha256:4f58e37f5217f0e35644167c5eee875993d589267a73817e7f0135c4e99bdb7b

Observation 38141650-6fcb-4e1f-9623-450cc0ed660b · inbound

Trends in AI Supercomputers cites this paper.

Trends in AI Supercomputers Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T11:17:19.030628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:17:19.030628Z digest=sha256:0ac7509814294b554f044be67d41f9baab7e16dbc7a6a6e9955672035c6afc9f

Observation 0301b650-4d67-42df-9fa1-fd627987e3a5 · inbound

Taming the Titans: A Survey of Efficient LLM Inference Serving cites this paper.

Taming the Titans: A Survey of Efficient LLM Inference Serving Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:10.027522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:49:10.027522Z digest=sha256:12af1cd331519beb509946650f3584964fef670fb18f0e6744c769f8a20a36a5

Observation 7dc5a2f3-85a1-4c6f-a8c2-299c257b3ebc · inbound

Hetu v2: A General and Scalable Deep Learning System with Hierarchical and Heterogeneous Single Program Multiple Data Annotations cites this paper.

Hetu v2: A General and Scalable Deep Learning System with Hierarchical and Heterogeneous Single Program Multiple Data Annotations Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T05:32:40.828417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:32:40.828417Z digest=sha256:c1dbade10bcb929e1298b470e24dda1217d0a125fb27e1aab43e20adea146793

Observation b98d2e69-c40f-4c0f-a376-a80a2b2b4410 · inbound

Hardware-Efficient Attention for Fast Decoding cites this paper.

Hardware-Efficient Attention for Fast Decoding Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:31.424098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:31.424098Z digest=sha256:2f2cb549c60223c26af8ee0db002cdad653655319719b69c18ac00a9ad2c6add

Observation 2b1adefe-1e6a-4e8c-8fe1-fb2f309c00e3 · inbound

Speeding up Model Loading with fastsafetensors cites this paper.

Speeding up Model Loading with fastsafetensors Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:04.791744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:04.791744Z digest=sha256:4ef850f7f236dbae9a97b94319c41d00a355b880fb2b50392aeaa5531eed7a81

Observation 2416c498-850c-4e9e-9e1a-2891e812dea0 · inbound

Technical Report of TeleChat2, TeleChat2.5 and T1 cites this paper.

Technical Report of TeleChat2, TeleChat2.5 and T1 Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:22.531285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:43:22.531285Z digest=sha256:606eba07e0952ba918d8c18f5e8bf41bb3a8fe32d5c2d055a86cba731291ee27

Observation 434e3dfa-8e73-4bda-9894-afc73a21f4fb · inbound

Efficient and Scalable Agentic AI with Heterogeneous Systems cites this paper.

Efficient and Scalable Agentic AI with Heterogeneous Systems Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:16:14.564374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:16:14.564374Z digest=sha256:439103acb6ebedb4621d1300afd7c6adf9ce43aae5986a9257957f71ac41a3da

Observation 98915eea-eaf6-4490-9f89-fa2b7b9db928 · inbound

SpikingBrain: Spiking Brain-inspired Large Models cites this paper.

SpikingBrain: Spiking Brain-inspired Large Models Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:51:45.811840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T18:51:06.243305Z digest=sha256:831c6f76ea234165aa876f4c3a75374ba393746cd6eb00e2aac92daeb6c01305

Observation fdb78c4e-ce1a-4b43-aee2-89eaf78d9c06 · inbound

Towards EnergyGPT: A Large Language Model Specialized for the Energy Sector cites this paper.

Towards EnergyGPT: A Large Language Model Specialized for the Energy Sector Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:42:47.536706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T17:39:17.456350Z digest=sha256:d0b24ae5891a8bce24a1220fea30b67f11b7cd9129c6d91b0465abcc81afa959

Observation 5016bb78-fabd-40f9-90ce-84c305794c98 · inbound

Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining cites this paper.

Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T04:49:03.022290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T04:46:33.641714Z digest=sha256:6ce7ec59e0a31f133d8c0c25fe80945a01ba4f88750723133af6bb5897dfcbbb

Observation 24b707cb-bacc-4ff0-948d-eaa90040a26f · inbound

Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning cites this paper.

Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:43:37.893546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-16T22:43:01.937642Z digest=sha256:5ca4a60b100f965964728af64fc4e1a8c7fc5a464d9fc3db07f30f5b2d6ef608

Observation 19c9e4ea-19e8-4fd7-82ab-82d278e94d87 · inbound

Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency cites this paper.

Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:09.562026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:09.562026Z digest=sha256:692fcdb0f8ccb6b6c6ec600c15b2d1e5bd9d854451e810c684e8332053bb0957

Observation 328733c0-12be-44c1-8352-58b97e7d75af · inbound

Kimi K2.5: Visual Agentic Intelligence cites this paper.

Kimi K2.5: Visual Agentic Intelligence Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-10T16:09:05.345041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T16:09:05.225767Z digest=sha256:23642cd78011c7cfba2605c255238baaa4624bb483a7b3a691e844d32899aa8f

Observation 177b40ef-d995-4946-a80f-645b9db7e617 · inbound

Attention Residuals cites this paper.

Attention Residuals Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:39:04.522153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T06:39:04.312270Z digest=sha256:7fa0edacad4de9561d131991162eb5790f1e1c2b449bdc14b71e472258d8ba68

Observation c9817349-2284-48aa-8094-16363d79dabe · inbound

AEGIS: Scaling Long-Sequence Homomorphic Encrypted Transformer Inference via Hybrid Parallelism on Multi-GPU Systems cites this paper.

AEGIS: Scaling Long-Sequence Homomorphic Encrypted Transformer Inference via Hybrid Parallelism on Multi-GPU Systems Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:23:09.283474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T19:21:26.158870Z digest=sha256:ccb1416cad0a3377dd47e991710442c128ae1f5fb7849b0cde64c79f535011f5

Observation 6a0912e7-1778-41aa-9a95-863b705e390c · inbound

An Engineering Journey Training Large Language Models at Scale on Alps: The Apertus Experience cites this paper.

An Engineering Journey Training Large Language Models at Scale on Alps: The Apertus Experience Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:20:30.878284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T14:16:02.816822Z digest=sha256:43c7f9270bf78fe486d7317a37a01cfbc4ff41ee05fe7b84cb96bf93ce805913

Observation b02ba0ef-36db-49ff-9907-d9d5d8aa0421 · inbound

Nautilus: An Auto-Scheduling Tensor Compiler for Efficient Tiled GPU Kernels cites this paper.

Nautilus: An Auto-Scheduling Tensor Compiler for Efficient Tiled GPU Kernels Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:24:21.492412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T10:21:05.219519Z digest=sha256:dec26f72c8bf8adf1c15dde6c5fee0c6faa38086bee35e978c48e46442e7d0fa

Observation b98fd996-021b-4bf4-bd1b-e32939204f3b · inbound

GPUOS: A GPU Operating System Primitive for Transparent Operation Fusion cites this paper.

GPUOS: A GPU Operating System Primitive for Transparent Operation Fusion Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:51:04.200107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T04:30:44.728872Z digest=sha256:4a9b2d0127f3493cceefc169f4487f4301fc34b265c1dd7ecbe7b735797508ea

Observation fbf8fc09-41b9-4bdb-b3b7-64885625d174 · inbound

Efficient Training on Multiple Consumer GPUs with RoundPipe cites this paper.

Efficient Training on Multiple Consumer GPUs with RoundPipe Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:31:26.784462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-07T10:37:22.251566Z digest=sha256:1f577a871ad19ff338b36dcef1748d602fce2bb50be4706dc001c1f5ee61b124

Observation 4459bbc7-997e-4ae8-bc91-8ddb8ad82895 · inbound

A Scalable Recipe on SuperMUC-NG Phase 2: Efficient Large-Scale Training of Language Models cites this paper.

A Scalable Recipe on SuperMUC-NG Phase 2: Efficient Large-Scale Training of Language Models Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:45:56.331268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T02:21:17.841592Z digest=sha256:6dfe19b950912422fa046dca558e7def70287fc6f161778ea553c0c07d5b3cbe

Observation d8715644-38f9-4b6f-9f37-30e957cda438 · inbound

Navigating LLM Valley: From AdamW to Memory-Efficient and Matrix-Based Optimizers cites this paper.

Navigating LLM Valley: From AdamW to Memory-Efficient and Matrix-Based Optimizers Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:41:45.669017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T04:01:32.057022Z digest=sha256:a179706cfcb182024842a184760a3116c1ca68a7ab162da2decd0477bfcc5313

Observation 5e29785d-d1b4-451b-88e8-0ea4b476cde0 · inbound

MinT: Managed Infrastructure for Training and Serving Millions of LLMs cites this paper.

MinT: Managed Infrastructure for Training and Serving Millions of LLMs Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:27:51.896687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-14T19:25:12.407148Z digest=sha256:9eef15cee0bbc089d49bd68368f2844096226823af422a41ea69ae9ef004e9f0

Observation 66d0774e-14e5-4333-ad20-0313dfe62708 · inbound

MinT: Managed Infrastructure for Training and Serving Millions of LLMs cites this paper.

MinT: Managed Infrastructure for Training and Serving Millions of LLMs Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:05:06.246888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T21:47:00.295144Z digest=sha256:e1741af7fc24eb9d459b5b20cca7d00feefd9b8e02a1a2289f354eac6c97ab73

Observation 98a325aa-8eea-4233-abb5-1b7f161c9428 · inbound

Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference cites this paper.

Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T14:13:21.202970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T14:11:58.106397Z digest=sha256:97c9bf5fd396e25a4dc23adf0dc67e5cfd04fc4acfe00973fb0db5317623bb9c

Observation 74ea8c6e-298e-4381-a6e7-301a64b72879 · inbound

Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference cites this paper.

Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:59:55.795270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T08:55:31.298030Z digest=sha256:f360b5fd55c56defd1a6a05142500360b774dddf9c0ba8aa6d91102d387b9984

Observation 79497a54-12f9-423d-aafb-370042858c82 · inbound

A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability cites this paper.

A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:38:09.493332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T07:35:32.225708Z digest=sha256:9e7edc227a8339682fb60f55ba2d9b2c3db18929858d034b78b1ed9d7f165ed5

Observation 12c7c42b-e7ee-4c9e-a26f-e5612c9b589e · inbound

Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks cites this paper.

Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:44:48.308088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T15:42:21.405913Z digest=sha256:998dc2efff86f4b11d6df13bec9e38bb37b52534b5afdb6ec7c6481e3fd0bd9f

Observation 18dd1c7d-1d54-4fb3-b41d-d0ad1f22525a · inbound

Heterogeneous Parallelism for Multimodal Large Language Model Training cites this paper.

Heterogeneous Parallelism for Multimodal Large Language Model Training Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:43:50.412459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T18:42:09.591282Z digest=sha256:54f2448d7e013ca5c1d4f1f03ddae65c82f4d8254161e2ed07c40cc771db71dc

Observation 702f9e44-561d-495f-a5ef-ae30dc9bc281 · inbound

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention cites this paper.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.872984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:1a56541de55f66c84a853d93d60857a5379df244c57010a794b90d656a910f07

Observation 5c6e3757-cf93-446d-827b-e7eeb1a5facb · inbound

Model Multiplicity for Adversarial Detection in Small Language Model Training on Edge Devices cites this paper.

Model Multiplicity for Adversarial Detection in Small Language Model Training on Edge Devices Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T19:37:19.322887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T21:26:08.955055Z digest=sha256:745a1aac0249bd1a727d80fe960c82039d4cda6ee4ec3e3b4d226566d5217bce

Observation 8ac83b28-42d8-4410-8f26-1dd54aa5ebbe · inbound

Piper: A Programmable Distributed Training System cites this paper.

Piper: A Programmable Distributed Training System Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-03T07:57:44.655438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T11:34:02.562929Z digest=sha256:7c81449881dd8b352843e16a4f24a5a2bb5ea63064c562f2d16552c75879054b

Observation 132e3e92-7590-4d9e-87c2-8672827fad05 · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 225

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:09:46.461495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T08:09:57.542558Z digest=sha256:33caa836d686d7ce824359d130a9ea8e3124ab500fd19d93d370bb6e04d54cdb

Observation db29d1e9-7eb4-4784-8506-b0ee07621a42 · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 211

Resolution
unresolved
no resolver link, observed 2026-08-02T10:27:18.508664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:27:18.508664Z digest=sha256:1343c808d036ef38d4fddd710b8bf1477573a9687f06519416b9a073de3e4077

Observation 78b96fde-1b39-415a-b5e3-11a9ee3918da · inbound

PHOENIX: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint cites this paper.

PHOENIX: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:28:43.916831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-03T17:26:07.870260Z digest=sha256:17e25367d3a6e08de7a8700d9ffde7cb30ec955f16869048c9a8f2abdb77ab1f

Observation e857b2bb-5386-442e-8cc5-6b2d1d92bfee · inbound

PHOENIX: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint cites this paper.

PHOENIX: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T08:40:29.554554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T08:40:29.554554Z digest=sha256:95f8c586c8ff9b7ed9145485ad2c85a47ff2bdc9ed4ae5ea6c317a622b3086e3

Observation e919841b-30d3-4e66-a2d1-e594d06b6753 · inbound

Design-CP: Context Parallelism for Design of Protein Nanoparticles cites this paper.

Design-CP: Context Parallelism for Design of Protein Nanoparticles Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-12T02:29:52.764344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T02:29:52.764344Z digest=sha256:bf3c31de4222bfecc9c51c80b91b5ff90fca30539e889d05903ae951b4593f4d

Observation 7ed52c58-f009-466d-870c-16cd34446578 · inbound

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining cites this paper.

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-09T09:16:06.525993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:82156de5a1e0b9c0834881c43fdfd4cf58067134023bca08c9f9e0e1089e2912

Observation 00d01956-3547-4166-88a1-c427c5c4440e · inbound

LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget cites this paper.

LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-02T00:41:45.468565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:41:45.468565Z digest=sha256:8d923fe93aebb5410417f07e47e811edcf07ff0f18590dc1bb708934722c445e

Observation 73494bae-8a1d-4c9f-a9ac-3a3c3d90d1a0 · inbound

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix cites this paper.

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-01T17:32:37.511773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:32:37.511773Z digest=sha256:a16c49f34fad056a61ee098811c627377f242543e0355805c4a06964a95fcc61

Observation af076d3f-fb73-4ce8-9211-432cca962ff0 · inbound

MALT: Lightweight Curvature-Aware Muon via Diagonal Preconditioning cites this paper.

MALT: Lightweight Curvature-Aware Muon via Diagonal Preconditioning Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T05:39:25.283911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:39:25.283911Z digest=sha256:2dc45f8cd804dc443964b35bdb6b53fc5bb9dc29af80a4a6ce0cce618ca6120c

Observation e931b501-2f67-4399-aabb-423c49c75522 · inbound

Memory-Efficient Activation Checkpointing with Sliding Window and Hirschberg's Algorithm for 0/1 Knapsack Solving in PyTorch cites this paper.

Memory-Efficient Activation Checkpointing with Sliding Window and Hirschberg's Algorithm for 0/1 Knapsack Solving in PyTorch Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-14T04:31:49.781442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:31:49.781442Z digest=sha256:85aa2cd90bd939215a8540e36356beb39deb55acf252134eae2327ff3f0a3d96

Observation 1a8ef394-e7e8-4f03-adaf-c0e050621e9d · inbound

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model cites this paper.

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-14T14:53:28.934067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T14:53:28.934067Z digest=sha256:39b56818e73ee33d0dde878ce6052d1a82bd690589e4d058a79cf7b6c7150645