Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 49 inbound Pith citation observations for arXiv:2104.04473.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:17:19.030628Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
9
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation a4d62a99-e85b-4e0d-86d5-17278d726ffe · inbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 50c00f4f-b34c-4082-acab-2ff8b53dd1d2 · inbound
OPT: Open Pre-trained Transformer Language Models Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 142
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 52a26e0c-be81-4f2f-9c4c-ef851eb31be0 · inbound
Multi-matrix Factorization Attention Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e815fc39-0ed3-46ec-a5df-920a95bf6b17 · inbound
Automatically Planning Optimal Parallel Strategy for Large Language Models Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcbdff8a-f26d-4bc1-b441-fb0214fe0104 · inbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9da226f6-618d-430e-9a3f-1de9a9a1b1f6 · inbound
Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb0670d6-8efc-4d2e-b247-9226f418704c · inbound
Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38141650-6fcb-4e1f-9623-450cc0ed660b · inbound
Trends in AI Supercomputers Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0301b650-4d67-42df-9fa1-fd627987e3a5 · inbound
Taming the Titans: A Survey of Efficient LLM Inference Serving Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dc5a2f3-85a1-4c6f-a8c2-299c257b3ebc · inbound
Hetu v2: A General and Scalable Deep Learning System with Hierarchical and Heterogeneous Single Program Multiple Data Annotations Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b98d2e69-c40f-4c0f-a376-a80a2b2b4410 · inbound
Hardware-Efficient Attention for Fast Decoding Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b1adefe-1e6a-4e8c-8fe1-fb2f309c00e3 · inbound
Speeding up Model Loading with fastsafetensors Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2416c498-850c-4e9e-9e1a-2891e812dea0 · inbound
Technical Report of TeleChat2, TeleChat2.5 and T1 Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 434e3dfa-8e73-4bda-9894-afc73a21f4fb · inbound
Efficient and Scalable Agentic AI with Heterogeneous Systems Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98915eea-eaf6-4490-9f89-fa2b7b9db928 · inbound
SpikingBrain: Spiking Brain-inspired Large Models Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation fdb78c4e-ce1a-4b43-aee2-89eaf78d9c06 · inbound
Towards EnergyGPT: A Large Language Model Specialized for the Energy Sector Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5016bb78-fabd-40f9-90ce-84c305794c98 · inbound
Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 24b707cb-bacc-4ff0-948d-eaa90040a26f · inbound
Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 19c9e4ea-19e8-4fd7-82ab-82d278e94d87 · inbound
Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 328733c0-12be-44c1-8352-58b97e7d75af · inbound
Kimi K2.5: Visual Agentic Intelligence Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 177b40ef-d995-4946-a80f-645b9db7e617 · inbound
Attention Residuals Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c9817349-2284-48aa-8094-16363d79dabe · inbound
AEGIS: Scaling Long-Sequence Homomorphic Encrypted Transformer Inference via Hybrid Parallelism on Multi-GPU Systems Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6a0912e7-1778-41aa-9a95-863b705e390c · inbound
An Engineering Journey Training Large Language Models at Scale on Alps: The Apertus Experience Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b02ba0ef-36db-49ff-9907-d9d5d8aa0421 · inbound
Nautilus: An Auto-Scheduling Tensor Compiler for Efficient Tiled GPU Kernels Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b98fd996-021b-4bf4-bd1b-e32939204f3b · inbound
GPUOS: A GPU Operating System Primitive for Transparent Operation Fusion Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation fbf8fc09-41b9-4bdb-b3b7-64885625d174 · inbound
Efficient Training on Multiple Consumer GPUs with RoundPipe Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4459bbc7-997e-4ae8-bc91-8ddb8ad82895 · inbound
A Scalable Recipe on SuperMUC-NG Phase 2: Efficient Large-Scale Training of Language Models Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d8715644-38f9-4b6f-9f37-30e957cda438 · inbound
Navigating LLM Valley: From AdamW to Memory-Efficient and Matrix-Based Optimizers Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5e29785d-d1b4-451b-88e8-0ea4b476cde0 · inbound
MinT: Managed Infrastructure for Training and Serving Millions of LLMs Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 66d0774e-14e5-4333-ad20-0313dfe62708 · inbound
MinT: Managed Infrastructure for Training and Serving Millions of LLMs Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 98a325aa-8eea-4233-abb5-1b7f161c9428 · inbound
Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 74ea8c6e-298e-4381-a6e7-301a64b72879 · inbound
Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 79497a54-12f9-423d-aafb-370042858c82 · inbound
A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 12c7c42b-e7ee-4c9e-a26f-e5612c9b589e · inbound
Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 18dd1c7d-1d54-4fb3-b41d-d0ad1f22525a · inbound
Heterogeneous Parallelism for Multimodal Large Language Model Training Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 702f9e44-561d-495f-a5ef-ae30dc9bc281 · inbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5c6e3757-cf93-446d-827b-e7eeb1a5facb · inbound
Model Multiplicity for Adversarial Detection in Small Language Model Training on Edge Devices Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8ac83b28-42d8-4410-8f26-1dd54aa5ebbe · inbound
Piper: A Programmable Distributed Training System Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 132e3e92-7590-4d9e-87c2-8672827fad05 · inbound
The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 225
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation db29d1e9-7eb4-4784-8506-b0ee07621a42 · inbound
The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 211
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78b96fde-1b39-415a-b5e3-11a9ee3918da · inbound
PHOENIX: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e857b2bb-5386-442e-8cc5-6b2d1d92bfee · inbound
PHOENIX: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e919841b-30d3-4e66-a2d1-e594d06b6753 · inbound
Design-CP: Context Parallelism for Design of Protein Nanoparticles Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ed52c58-f009-466d-870c-16cd34446578 · inbound
GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 00d01956-3547-4166-88a1-c427c5c4440e · inbound
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73494bae-8a1d-4c9f-a9ac-3a3c3d90d1a0 · inbound
A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af076d3f-fb73-4ce8-9211-432cca962ff0 · inbound
MALT: Lightweight Curvature-Aware Muon via Diagonal Preconditioning Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e931b501-2f67-4399-aabb-423c49c75522 · inbound
Memory-Efficient Activation Checkpointing with Sliding Window and Hirschberg's Algorithm for 0/1 Knapsack Solving in PyTorch Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a8ef394-e7e8-4f03-adaf-c0e050621e9d · inbound
Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.