Pith. sign in

Paper Citation Record · LEDGER

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers

As of 10 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2502.08145.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.08145 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T10:20:02.783216Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact5
  • verified fuzzy33
  • unresolved10
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cfe60ddb-b8d2-4770-b176-54ab30e9e410 · outbound

This paper cites Super: Sub-graph parallelism for transformers,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Super: Sub-graph parallelism for transformers,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:05.277436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.624115Z digest=sha256:da32c341ef7de3d9933d54f3424a930782acd0254c019806eabf6fea9d347dd3

Observation e6e52360-d80b-4c1f-a0da-591463fe3c47 · outbound

This paper cites Scaling distributed deep learning work- loads beyond the memory capacity with karma,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Scaling distributed deep learning work- loads beyond the memory capacity with karma,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:05.268993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.628041Z digest=sha256:774838cbd7257a420ff7aaa3ba53c4685ab148f68d50aa2098f4973c080b9a0a

Observation 173818ab-a30e-46b0-8e04-3a0b664a992f · outbound

This paper cites Forge: Pre-training open foundation models for science,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Forge: Pre-training open foundation models for science,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:05.261411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.631258Z digest=sha256:b2fa0296a7aad4976ebc19455c4371b53bf5532554687590e22063ae1a078587

Observation 05f24982-92c5-4040-a46f-8d80de8a20a7 · outbound

This paper cites Optimizing distributed training on frontier for large language models,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Optimizing distributed training on frontier for large language models,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:05.251706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.634237Z digest=sha256:1194cf61ecb215cc831e381c4e817f613e696a72aef3d3c70d14641bd00c837d

Observation ae7568c9-1fd7-465d-8c55-ee1c7922a044 · outbound

This paper cites Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:05.242863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.638603Z digest=sha256:7dff6c789049fffd28a5451b250028687f3484073311f69b473cdd76376d84ee

Observation fb0670d6-8efc-4d2e-b247-9226f418704c · outbound

This paper cites Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:02.642128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:20:02.642128Z digest=sha256:00e920418c6e0539aa05b325ba85d1171c541d6e8a257e11d15160e8aba315f5

Observation e5e73c07-1c20-43a9-b9c3-b2cccfb97aa8 · outbound

This paper cites MegaScale: Scaling large language model training to more than 10,000 GPUs,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers MegaScale: Scaling large language model training to more than 10,000 GPUs,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:05.234090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.645980Z digest=sha256:3a949fabb09a4682684f5c136f893f8b76531a1d3ceb779001e32346d9481ed8

Observation 5fb685db-f7b6-4a83-b6f5-0fb09d67e5db · outbound

This paper cites Google cloud demonstrates the world’s largest distributed training job for large language models across 50000+ tpu v5e chips,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Google cloud demonstrates the world’s largest distributed training job for large language models across 50000+ tpu v5e chips,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:05.224647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.649415Z digest=sha256:13bd672e79370dc702bb5ff23ac557fc819ed307fb06d84910265ebbe59ce560

Observation d31439de-b84f-43b5-ab93-7aff0e022af2 · outbound

This paper cites AxoNN: An asynchronous, message-driven parallel framework for extreme-scale deep learning,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers AxoNN: An asynchronous, message-driven parallel framework for extreme-scale deep learning,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:05.215715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.652845Z digest=sha256:2021ed869a192ac6c2a442020948757476bb1b913007e87274c86e3b4e9de8cf

Observation 34d2b012-dcf0-4de7-9aac-58efaec82c63 · outbound

This paper cites Exploiting sparsity in pruned neural networks to optimize large model training,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Exploiting sparsity in pruned neural networks to optimize large model training,

Reference 10

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T10:20:04.956493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.655642Z digest=sha256:bee3468f8ac37aab71e34887bc6cfc717f3eb13010da14405446dd3c04e9e83b

Observation 92cc03c1-322f-49ff-9c7e-96cfae54dcd9 · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Zero: Memory optimizations toward training trillion parameter models,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:05.206332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.659372Z digest=sha256:a1490ec7d2eea095adcfeb73cc9153fbadb5236e90ede1325b128deec34d186f

Observation 15d514db-84c9-449c-8c4d-b2e5864ee6fd · outbound

This paper cites Pytorch fsdp: Experiences on scaling fully sharded data parallel,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Pytorch fsdp: Experiences on scaling fully sharded data parallel,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:05.197454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.662349Z digest=sha256:8c25e9d4f47ec6f91d4044b826c8a781aa3f5b0e081285c5a9d556a66dc1a508

Observation 3adfc4b9-0a73-4fa5-ad42-26269b8f8431 · outbound

This paper cites Megatron-lm: Training multi-billion parameter language models using model parallelism,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Megatron-lm: Training multi-billion parameter language models using model parallelism,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:05.179145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.668560Z digest=sha256:ff8558ec72bdd1829b47a775c8bc00cc23ab332ccee0ef433609b4a656529cb2

Observation a4b70986-2b49-46e6-8dc1-4fe17fee7aca · outbound

This paper cites GPipe: efficient training of giant neural networks using pipeline parallelism,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers GPipe: efficient training of giant neural networks using pipeline parallelism,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:05.170498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.671345Z digest=sha256:b2c9b69c654ce8b68eaa5ab794611455fc8dde557dbb20f1cd8b6825de211d74

Observation d1a34e15-d91d-4c5d-9bd6-52060981d60c · outbound

This paper cites Deepspeed: Extreme-scale model training for everyone,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Deepspeed: Extreme-scale model training for everyone,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:05.162555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.674012Z digest=sha256:024cd6680cf16bf3dddd222a1800a72c33d4e692797b3124b7346dfb967c69dc

Observation e8091f37-6dfe-4fa1-8251-e0e7012908a5 · outbound

This paper cites A hybrid tensor-expert-data parallelism approach to optimize mixture-of-experts training,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers A hybrid tensor-expert-data parallelism approach to optimize mixture-of-experts training,

Reference 17

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T10:20:04.609477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.677289Z digest=sha256:2453117c21101721d6bf0da020b88bbab233eb654c80fb3094cd158d8281f6a5

Observation 1dc7ce9b-2044-4557-8c17-4858138b72d0 · outbound

This paper cites GPT-NeoX: Large Scale Autoregressive Language Modeling in PyTorch,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers GPT-NeoX: Large Scale Autoregressive Language Modeling in PyTorch,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:05.152654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.680393Z digest=sha256:91f52ddd570e53eb6cb1cda10d0cd29e108541a8d77e6fa774fe1b7f7c072a1c

Observation 2a4c607d-2d59-4ad5-aba4-f0ec474863d6 · outbound

This paper cites Alpa: Automating Inter- and Intra-Operator Parallelism for Distributed Deep Learning.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Alpa: Automating Inter- and Intra-Operator Parallelism for Distributed Deep Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:02.683674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:20:02.683674Z digest=sha256:7026c47ece752f1ef8b6bf676afd49daed9aa33f709dc36308fc1e3b18fb3156

Observation 6f350aae-1cf5-48ed-b4fd-71628a22086c · outbound

This paper cites Colossal-AI: a unified deep learning system for large-scale parallel training,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Colossal-AI: a unified deep learning system for large-scale parallel training,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:05.144149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.686844Z digest=sha256:305cf4f32da0ef840e136a115424f0d38a557f711354e8843ddb59ad32f67c09

Observation 8854f878-52af-44af-9404-6ff66b5a0ad9 · outbound

This paper cites Llama 2: Open foundation and fine-tuned chat models,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Llama 2: Open foundation and fine-tuned chat models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:05.133794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.689764Z digest=sha256:49dc5131493bb413e5ae13a8dfaff1b6318b26fa957224c572572bfdc551f5c6

Observation 8b3f9544-2386-40d4-9509-223cf5ea0c6e · outbound

This paper cites Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:02.692885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:20:02.692885Z digest=sha256:58c58d52c704503b501ba5dc93516f26c44ccc8227b9af102a0503d15fdfb60f

Observation 9ba13cb0-3cdc-4355-a15a-12751caef3c7 · outbound

This paper cites LBANN: livermore big artificial neural network HPC toolkit,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers LBANN: livermore big artificial neural network HPC toolkit,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:05.125390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.696500Z digest=sha256:2105f6adfa72d604f6dad00524ee376b2b5b5bbf10f22d1fee3e378b54614d0a

Observation bfd7db8b-9d6c-4dd3-bd2e-e40fcaf5ce0d · outbound

This paper cites Nvidia selene supercomputer,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Nvidia selene supercomputer,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:05.117015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.699888Z digest=sha256:4dda1293722fa48e403bffc00532b95cc8440b1d50442ebb2adfe34dedadb1de

Observation d6c41d50-7547-4753-9b5c-04cfec9da0fd · outbound

This paper cites Frontier: Exploring exascale,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Frontier: Exploring exascale,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:05.108866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.703399Z digest=sha256:4ad0ec2b1b88d7e40c8b91a954a7f3b2c6d69286740618191c7bd6011445c401

Observation 02f325f9-583f-4c53-9893-f82b4e765e40 · outbound

This paper cites A three-dimensional approach to parallel matrix multiplication,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers A three-dimensional approach to parallel matrix multiplication,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:05.100532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.706230Z digest=sha256:ca73416f5831921c4af9965ad95a7b75b54d0ec37dbac76053c096e25ec8d99f

Observation 38b79e32-95e6-49cb-a25c-afe97b3a2065 · outbound

This paper cites ZeRO++: Extremely efficient collective communication for large model training,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers ZeRO++: Extremely efficient collective communication for large model training,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:05.188445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.709021Z digest=sha256:a110ac38d2551d64a04b1bb659d994dcfcc735735fe39980080bfb343fa408ff

Observation fb5c813e-084a-4de4-9796-71e62a58cd82 · outbound

This paper cites Improving the performance of collective operations in mpich,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Improving the performance of collective operations in mpich,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:05.091221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.712838Z digest=sha256:9037541c9117e1b40105544407e547bbb0f159c066d4ddd08aa9343713859638

Observation 961544f7-2ae5-40f2-877b-6060c90bdeb0 · outbound

This paper cites Optimization of collective reduction operations,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Optimization of collective reduction operations,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:05.082585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.715458Z digest=sha256:e5e9821357eb4fad6f4c9d0aef2eb91c8d7e062fd617a3719541a5821dd3abe1

Observation 027f16fe-75f9-4c37-b9be-d0606c3646e4 · outbound

This paper cites Improving communication performance in dense linear algebra via topology aware collectives,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Improving communication performance in dense linear algebra via topology aware collectives,

Reference 30

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T10:20:03.301128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.718427Z digest=sha256:24f199540e315a05ae605020a9c34619b7ef5eaf92526edc660265f99a746a55

Observation 2856382c-b8dc-419c-9ece-b2bb727d5dbb · outbound

This paper cites Mapping applications with collectives over sub-communicators on torus networks,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Mapping applications with collectives over sub-communicators on torus networks,

Reference 31

Resolution
malformed identifier
doi_truncated, observed 2026-08-08T10:20:02.820258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.721918Z digest=sha256:c58ef5612dca41c0263efe979a96446a402bee801a2810e598204fa9072845de

Observation ffae8d80-2e29-432e-8d87-6220294e365e · outbound

This paper cites RAHTM: Routing- algorithm aware hierarchical task mapping,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers RAHTM: Routing- algorithm aware hierarchical task mapping,

Reference 32

Resolution
verified exact
doi, observed 2026-08-08T10:20:02.809304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.725877Z digest=sha256:519e2a388a6597718ee59defc63b229254a5f412a04c8abf2e3611a71ffdd407

Observation 18f394f7-eb24-445c-80cb-9b33769fb30e · outbound

This paper cites Optimizing the performance of parallel applications on a 5D torus via task mapping,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Optimizing the performance of parallel applications on a 5D torus via task mapping,

Reference 33

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T10:20:03.048400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.728858Z digest=sha256:96b4ebc3b3bf8a5074c8648da7f51e0bbcf0c0a0010d63d99020a80671e91161

Observation fa8248a8-893e-42fb-8106-6fa3d1da566c · outbound

This paper cites Supervised learning based algorithm selection for deep neural networks,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Supervised learning based algorithm selection for deep neural networks,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:05.073494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.732776Z digest=sha256:2757942c9828de8fe5f97122002be2ef220da702eefee17fca16332f384e094b

Observation bcbddde1-a435-4540-98b6-c815a313d1fd · outbound

This paper cites Language Models are Few-Shot Learners.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Language Models are Few-Shot Learners

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:02.735880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:20:02.735880Z digest=sha256:41430ae9d0c5b6c5a0f2fa679732aca52923f83ef58ec65a7ba46bc3e5129bd0

Observation 25f95a0b-c2c5-408b-85ae-6bb5b9478acc · outbound

This paper cites Attention Is All You Need.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Attention Is All You Need

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:02.739697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:20:02.739697Z digest=sha256:ccbb619eb99c14fb2bb6d9761ad659a9b784ef3e121be36882a1090105582d8f

Observation b6d68b7c-b16a-48b5-9784-ea7c3494d9d0 · outbound

This paper cites Bigscience large open-science open-access multilingual language model,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Bigscience large open-science open-access multilingual language model,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:05.065419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.742736Z digest=sha256:9492180b62cd93b599f9cd8f7741149fb1597ad3d0d1894eb7efcfc18f91d01d

Observation 9a5e2c51-0a27-4a59-85c6-11fea84bdb77 · outbound

This paper cites Language models are unsupervised multitask learners,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Language models are unsupervised multitask learners,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:05.056115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.746066Z digest=sha256:6048a3726d86fece59a2cc96e2c9a7ad61527d19aeb859f0448e0ec0e1e32663

Observation 3a6163ec-71d9-4a64-9421-e53764fd4d5f · outbound

This paper cites Training Deep Nets with Sublinear Memory Cost.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Training Deep Nets with Sublinear Memory Cost

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:02.748714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:20:02.748714Z digest=sha256:a852c2355dda1c2e44a656e34835b385da9916ea6d1c073a3fa613b5688b59eb

Observation 933cbc24-0067-4915-9828-d982de5f4ef6 · outbound

This paper cites A Study of BFLOAT16 for Deep Learning Training.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers A Study of BFLOAT16 for Deep Learning Training

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:02.751774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:20:02.751774Z digest=sha256:dadfefb7b5da2442a54a270e7a5971cdfc0bec5b9a2c862003fbcfbbfcbdb8b8

Observation 43ef57a5-c4a9-4925-8a90-b4a201bfa18f · outbound

This paper cites an unresolved cited work.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-08T10:20:05.046128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.754450Z digest=sha256:2b9b3a178564648f869b907b7be79ad37cf791ecd8b9ca8ce506600f5f2a2cf9

Observation 7920df23-0bb7-4aa8-a738-cf3dca58ebf1 · outbound

This paper cites Interactive investigation of traffic congestion on fat-tree networks using TreeScope,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Interactive investigation of traffic congestion on fat-tree networks using TreeScope,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:05.036954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.757875Z digest=sha256:5cd69f6ba49a5f75d98586be4e60aca341eed51509e418396e2d6de08d41cee2

Observation 0f67b059-0295-489a-8c68-5dbbec1579ef · outbound

This paper cites Quantifying I/O and communication traffic interference on dragonfly networks equipped with burst buffers,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Quantifying I/O and communication traffic interference on dragonfly networks equipped with burst buffers,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:05.028284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.761481Z digest=sha256:99c9aba3b2a64928cfb4ff00bfe1e66fdc49836617cb2a640ac1c2414fce2366

Observation 26feff98-4dfa-4171-9546-48f55f5f623d · outbound

This paper cites Quantifying memorization across neural language models,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Quantifying memorization across neural language models,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:05.018885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.764626Z digest=sha256:533ae97d6b9a6df9c2128080b2fbf52d773d0a406712fd2ac03a3087556de178

Observation 648bc8e9-bf90-4fe9-b9dd-27e200a2238b · outbound

This paper cites The times sues openai and microsoft over ai use of copyrighted work,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers The times sues openai and microsoft over ai use of copyrighted work,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:05.008971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.768307Z digest=sha256:408f0e916b098a78a7947e8c78113bc6f3bfea78c4db97e700b152a671379844

Observation 9efce775-2041-4e4a-ab0f-cbf5d21b4bf6 · outbound

This paper cites Extracting training data from large language models,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Extracting training data from large language models,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:02.771544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:20:02.771544Z digest=sha256:47b5562f9a5ccee501c1864ba386f4c188c84523b96917a3c3d07ab983291bcf

Observation 9d69f755-c00b-4b91-97d0-6c7080ca91fa · outbound

This paper cites Pythia: A suite for analyzing large language models across training and scaling,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Pythia: A suite for analyzing large language models across training and scaling,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:04.994313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.774029Z digest=sha256:17706d734519675d67a7702d795fcfdcdb3ab07a2b7d34ee12edb1c680f2c458

Observation 178cdb6d-caa0-4152-ae75-80a4004fa1dc · outbound

This paper cites Tinyllama: An open-source small language model,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Tinyllama: An open-source small language model,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:04.985533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.777191Z digest=sha256:68f47f0fc015eca3db823a275272e8e5467d0fa4b059b1797fc86b52c62015e3

Observation df901eb7-3427-494b-baae-7e033835fb2a · outbound

This paper cites The llama 3 herd of models,.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers The llama 3 herd of models,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:20:04.976637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T10:20:02.780659Z digest=sha256:bf56e876a3b58d94be630d85f9289a6e36cee1414a3dd382777c212fbe3f4366

Observation 2145c077-2e84-45e2-9e97-4ed4002bfff4 · outbound

This paper cites Be like a Goldfish, Don't Memorize! Mitigating Memorization in Generative LLMs.

Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers Be like a Goldfish, Don't Memorize! Mitigating Memorization in Generative LLMs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T10:20:02.783216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:20:02.783216Z digest=sha256:11ecdc963bf8177e3e0bdb0665bb85c9fd5a4f79c33c58dac9d1bbf86893f656

Pith citing papers

No inbound Pith citation observations are available.