Pith. sign in

Paper Citation Record · LEDGER

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning

As of 22 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 0 inbound Pith citation observations for arXiv:2501.04266.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.04266 v2

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:42:38.233849Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

34 of 34 outbound references displayed

  • verified exact2
  • verified fuzzy12
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 556f303e-7734-4b11-b9d8-2844dd8417d2 · outbound

This paper cites Introducing the next generation of Claude — anthropic.com,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Introducing the next generation of Claude — anthropic.com,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.665918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:42:38.092503Z digest=sha256:b520daaf51193638e0863eec38e67ba33ee7fcdf13058e2a022d04d48f656fab

Observation d4e413ca-aabc-45e0-92b2-55fb4cbdfb9e · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Gemma: Open Models Based on Gemini Research and Technology

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.097420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.097420Z digest=sha256:3eb887377f93b26127cd066c7f45549abb8c0452b004925d5aa2e35cca55037b

Observation b15f156a-3aac-4470-a570-fb3a0d78245b · outbound

This paper cites Llama 3 model card,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Llama 3 model card,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.102113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.102113Z digest=sha256:498aba5b623e70e9819e1a9fb781319acf9ca350cca336e6ea7225823d4bd8a6

Observation 9dbe729f-675d-4788-ba68-023caeb7369a · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Evaluating Large Language Models Trained on Code

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.106390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.106390Z digest=sha256:a319722afddd12f3d93b0b6a9ed94a942bd7fe5e2f4403b6751a5d0770a7c9f4

Observation 0d7b9409-785c-44b4-912a-902f96bb8725 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.110991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.110991Z digest=sha256:2e260691039eab1babc0e73ea661b0460101cb974b827bbdaa3712c249f10d55

Observation 3b8c2180-fd1c-459e-9437-ea320f8ebbd7 · outbound

This paper cites Measuring massive multitask language understanding,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Measuring massive multitask language understanding,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.115345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.115345Z digest=sha256:55de6f5c6379fb956e5a0ac9857260eb6096263444e2413f396d569d840e0c2c

Observation bcbdff8a-f26d-4bc1-b441-fb0214fe0104 · outbound

This paper cites Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.123613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.123613Z digest=sha256:cba9e78e62383fe9800cb99f9bda5d53e4d23a3cc748eefe1b29b401ea8876aa

Observation b3fe7a0d-a675-4545-a842-7d5c137f3e71 · outbound

This paper cites The mvapich project: Transforming research into high-performance mpi library for hpc community,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning The mvapich project: Transforming research into high-performance mpi library for hpc community,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.640073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:42:38.127811Z digest=sha256:16e6b85fea94eb74958b75a0e0fc068d162cae8ad2b7bd1310e748f33c67a261

Observation 064a4e3b-3562-49de-9495-bfe45d86a64f · outbound

This paper cites Scalable Distributed DNN Training using TensorFlow and CUDA-Aware MPI: Characterization, Designs, and Performance Evaluation,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Scalable Distributed DNN Training using TensorFlow and CUDA-Aware MPI: Characterization, Designs, and Performance Evaluation,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.627707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:42:38.132220Z digest=sha256:5b99deafd62bcb0b71ee5b8366bbd81adeb029e98535001ccdca4666cf960ec5

Observation 954b19cf-6e28-4fac-b409-9ff30e210e7b · outbound

This paper cites NVIDIA Collective Communications Library (NCCL),.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning NVIDIA Collective Communications Library (NCCL),

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.614921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:42:38.135902Z digest=sha256:fbec0c96c15886916fff80b12fb36fa81775980bfe80f76c66c453c09393c74b

Observation 02bcaf6a-9e55-4078-a489-a98d61c9662c · outbound

This paper cites AMSP: Reducing Communication Overhead of ZeRO for Efficient LLM Training.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning AMSP: Reducing Communication Overhead of ZeRO for Efficient LLM Training

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:42:38.400265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:42:38.139962Z digest=sha256:d19ec75960bcbf7d7ef39e118074932472b99a869e8a01074f825bf47951c020

Observation 3da99280-6328-4b4c-a270-735321a48d59 · outbound

This paper cites ZeRO: Memory Optimizations Toward Training Trillion Parameter Models.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning ZeRO: Memory Optimizations Toward Training Trillion Parameter Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.144213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.144213Z digest=sha256:f8308e0bcd60969ec6cf04715cabc2c1ab050139ed868961e633e4d4f1427638

Observation a9c65893-895e-43d9-84ff-ba1a9ebdca58 · outbound

This paper cites Fairscale: A general purpose modular pytorch li- brary for high performance and large scale training,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Fairscale: A general purpose modular pytorch li- brary for high performance and large scale training,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.595578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:42:38.157012Z digest=sha256:30ba250d4ce832a0a5d0626583217f3f676add9ea48956baf1ab86f84797b93e

Observation 7d60d192-106f-499a-bdca-3fb29ecc7704 · outbound

This paper cites Megatron-LM: Ongoing research training transformer models at scale,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Megatron-LM: Ongoing research training transformer models at scale,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.583069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:42:38.160347Z digest=sha256:4376bd993b5f3f7759c278592959190cb2da054bcd9931d6aa0335fc585acc16

Observation 29ad6087-e169-4d2c-85c3-f1d09dcbbedc · outbound

This paper cites Frontier - HPE Cray EX235a, AMD Optimized 3rd Generation EPYC 64C 2GHz, AMD Instinct MI250X, Slingshot-11 | TOP500 — top500.org,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Frontier - HPE Cray EX235a, AMD Optimized 3rd Generation EPYC 64C 2GHz, AMD Instinct MI250X, Slingshot-11 | TOP500 — top500.org,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.571146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:42:38.164101Z digest=sha256:e406d9e5578c03a1d370b3cd7da84318a96dfe321a75a2195c4ca7ec012521e2

Observation f524f000-a4fe-4496-a184-36f228f07747 · outbound

This paper cites An in-depth analysis of the slingshot interconnect,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning An in-depth analysis of the slingshot interconnect,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.558817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:42:38.167716Z digest=sha256:7f0bd5043ca3d42852f1ce97ae90184100c879f5518eb8d86069cc02fb549433

Observation e6cee94d-9069-47d4-9335-4375fccf66f0 · outbound

This paper cites ZeRO++: Extremely Efficient Collective Communication for Giant Model Training.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning ZeRO++: Extremely Efficient Collective Communication for Giant Model Training

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.171907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.171907Z digest=sha256:d9d19b8ec2638e16f19f828ceb7df48a3938eff4a9107c40d3fc6bf45a974de0

Observation 18d3800b-6188-415e-b7f8-9f91ee29305c · outbound

This paper cites Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.175813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.175813Z digest=sha256:dd27fcff12ca32b5b3114e880fd7584ad9463c9f7e3e619d4fe93677e8bde85f

Observation 9b89cf4b-5e8f-44d7-9ef6-76c09e6094e6 · outbound

This paper cites Scaling single- image super-resolution training on modern hpc clusters: Early experi- ences,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Scaling single- image super-resolution training on modern hpc clusters: Early experi- ences,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.545770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:42:38.179854Z digest=sha256:5b8d0b412f25e37ddb02c36529825add32f9b765d9d1a595bb01a3b2599b3ac6

Observation 660dbb0d-91ae-41fb-bdd4-68fbb0370391 · outbound

This paper cites Adam: A method for stochastic optimization,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Adam: A method for stochastic optimization,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.183548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.183548Z digest=sha256:39109f71b93aaceda8d1164d06b8d78aac20d50f7978b05e1761f2c85aef670c

Observation a4623ac3-e50c-431e-8457-f0b7523e542d · outbound

This paper cites 8-bit Optimizers via Block-wise Quantization.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning 8-bit Optimizers via Block-wise Quantization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.187634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.187634Z digest=sha256:6a5fdf3f55b0f248ec3cd51fce567f216bf0bdae71b2f924d343337ca82401e1

Observation e42ab56d-3118-4ed5-8023-88c78fc161a3 · outbound

This paper cites Decoupled weight decay regularization,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Decoupled weight decay regularization,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.191520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.191520Z digest=sha256:c329d5ca0223858cd4f1b0f5f6ceba67ecc5e5f297309bbe897ea490bf512995

Observation b26f6e6b-5293-47ed-8b43-ed9f35407131 · outbound

This paper cites GPT-NeoX-20B: An open-source autoregressive language model,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning GPT-NeoX-20B: An open-source autoregressive language model,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.515169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:42:38.195131Z digest=sha256:ce67cf5657768a053f3d5452c7c31dc2217c1c9b20422ec4f8d5de063ce5ff47

Observation bfdced74-66cc-4749-82cc-914fa6cd931a · outbound

This paper cites Language models are few-shot learners,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Language models are few-shot learners,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.198731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.198731Z digest=sha256:89a1f285bb99e46b3e934e598aff2d98098685a73ccbc9b038a501ccf311ce59

Observation fa9efb2b-2a5e-4b36-a55e-8d563f5b3058 · outbound

This paper cites RWKV: Reinventing RNNs for the Transformer Era.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning RWKV: Reinventing RNNs for the Transformer Era

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.202557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.202557Z digest=sha256:9adb378b519a9f74af06225c054107184ce651041d3e2c3102c97c89facaeb97

Observation 71fdde9f-5416-476f-9f80-77c0ee2a3147 · outbound

This paper cites Pytorch fsdp: Experiences on scaling fully sharded data parallel,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Pytorch fsdp: Experiences on scaling fully sharded data parallel,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.206521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.206521Z digest=sha256:8eafaaf30a73ec4f0d2add959e2414c877f08f2405cdad3c7716c858c59c98af

Observation 960dfec8-e83b-4ee4-8161-11a0675d1929 · outbound

This paper cites Mics: Near-linear scaling for training gigantic model on public cloud,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Mics: Near-linear scaling for training gigantic model on public cloud,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.487222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:42:38.214329Z digest=sha256:fbfe5f173654f40f2716c8e4f44f60f3f10553bc6a24c233156520bb39fd0055

Observation 432033b2-76e4-4c6b-839f-cbbce1021da7 · outbound

This paper cites MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.218267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.218267Z digest=sha256:35007f41c81712d1c2337052c2197465dcfbddd3f8f7daa9a244336f50aa8b73

Observation 2b629f5e-1470-412e-a753-13dd96880d73 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.210400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.210400Z digest=sha256:45145ccf9e4982ddfc15bf9b1499755179e24ce19cae071f7f6728afab7f06f9

Observation 12f5fa30-71b4-4dd6-8ae2-27c76559ba6e · outbound

This paper cites Optimizing Distributed Training on Frontier for Large Language Models.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Optimizing Distributed Training on Frontier for Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.225511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.225511Z digest=sha256:6a8969dce16612e5aa9c85b76bf6147506d9b06ce12656f7292a4ae9a06bf49b

Observation 5e608e03-f496-4b9a-8945-d3c111500d6a · outbound

This paper cites Comparative Study of Large Language Model Architectures on Frontier.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Comparative Study of Large Language Model Architectures on Frontier

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:42:38.286343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:42:38.229469Z digest=sha256:d24c333f652cc40ef6f92305703c980965efc42d56fcdc291c99f634b8afefeb

Observation 75297a7e-1ce8-4d1a-8f66-b58553ba9c54 · outbound

This paper cites Accelerating large language model training with hybrid gpu-based compression,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Accelerating large language model training with hybrid gpu-based compression,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.472809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T21:42:38.222134Z digest=sha256:0d8c53e0c08f13fadc325a9ba41812b82f2bf03d89685475c3df93163b736945

Observation 70cee47f-4482-4411-b2cf-e42347847ace · outbound

This paper cites Hilfer fractional advection-diffusion equations with power-law initial condition; a Numerical study using variational iteration method.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Hilfer fractional advection-diffusion equations with power-law initial condition; a Numerical study using variational iteration method

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.233849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.233849Z digest=sha256:a47f966b2c53b2f8ce2b5f940ab8abd44ff6803a9ceaf775b19a037d26b2e3ad

Observation c3c524c6-630f-477b-b15e-2c946f85c94a · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Measuring Massive Multitask Language Understanding

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.119113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.119113Z digest=sha256:3be42b243248b19f198795b4d99b953f2e908432bb54f542aa632331679b23ec

Pith citing papers

No inbound Pith citation observations are available.