Pith. sign in

Paper Citation Record · LEDGER

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models

As of 13 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2411.10003.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.10003 v2

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:10:00.382168Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e36d548d-3952-47a7-a449-edec124c2c0a · outbound

This paper cites Gshard: Scaling giant models with conditional computation and automatic sharding,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Gshard: Scaling giant models with conditional computation and automatic sharding,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.199938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.199938Z digest=sha256:56b4acde9a9902fc91566af761eacb7d394770fee8bc024549202ade00a1e162

Observation b452d13e-0844-444c-9281-3036ffbb069a · outbound

This paper cites Glam: Efficient scaling of language models with mixture-of-experts,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Glam: Efficient scaling of language models with mixture-of-experts,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.203536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.203536Z digest=sha256:0891ad062e4a761f140c6a4ae2e7e345e88a1e090ae0bdb8d2dfe85b512dd6dc

Observation 7f798d6e-bf4f-47ca-b80d-96dc7b51556f · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.206774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.206774Z digest=sha256:96d5e179891203a10ac42c7f8b07c21b90b1e1218f2428f3cac3b7a664609e00

Observation cc6d02f9-b1e9-4827-a2fe-338b2737372e · outbound

This paper cites Deepspeed-moe: Advancing mixture-of- experts inference and training to power next-generation ai scale,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Deepspeed-moe: Advancing mixture-of- experts inference and training to power next-generation ai scale,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.210378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.210378Z digest=sha256:0b89a022a6fa73ddf64dc52345169f2496257131a5c20d6e78ffed7268f4b16a

Observation 05b77c6d-d58f-4c68-99c6-487f56dee8dd · outbound

This paper cites Tutel: Adaptive mixture-of-experts at scale,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Tutel: Adaptive mixture-of-experts at scale,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.213587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.213587Z digest=sha256:79318a5c5d21444d2d4f34d00afe32b97640409b050c361ba60b9627fd24065c

Observation 9dfd5072-c6ac-4d8d-b664-59294ae5c7ce · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.216540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.216540Z digest=sha256:d145bce3b1dca3df27e3c0c8d79800bc7005f51c40cf56496fa3d775b5de14fe

Observation c9efb5b9-d7e9-4890-bca6-8868a99fa61e · outbound

This paper cites Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.219726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.219726Z digest=sha256:bad45681909490a58eca52cde3266be1c779c21ea0880b177a4a8186e34040b8

Observation 09976f2d-8bef-4494-8040-2d022e3473b2 · outbound

This paper cites Faster- moe: modeling and optimizing training of large-scale dynamic pre- trained models,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Faster- moe: modeling and optimizing training of large-scale dynamic pre- trained models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.788777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:10:00.222415Z digest=sha256:4d9ff91a4b403d3da8cf97c64515694ab2a1b16cc956d00e2b60b74e7e257745

Observation c36279c3-9609-4d50-a18a-153a4be382b1 · outbound

This paper cites Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.224968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.224968Z digest=sha256:3ff8b3fe1ebb60bed443f23058779cb7bc103593e875f20756ea1221ac9d019e

Observation 29dd0267-7124-4cae-8baa-246793263d54 · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Zero: Memory optimizations toward training trillion parameter models,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.227609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.227609Z digest=sha256:2459e2947245ba4e137ea711c21aed98fb5023e40ecc29d35e365697ab2b6d15

Observation 49b1da81-cc16-4eb7-ae87-f6c63a864304 · outbound

This paper cites Scaling Laws for Neural Language Models.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Scaling Laws for Neural Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.231213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.231213Z digest=sha256:bdf3ced5374dd202907abd966c5b32b4f13c43d7ae700173be64440e63dde485

Observation c9d435d4-ea4a-4a82-bbef-b908c81fa615 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.234422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.234422Z digest=sha256:82b03685453b8152b7f4ddd4bd2264617cc856a9c9aa05a4d73ae13f364d8931

Observation 2e8315ff-cb5d-4199-b7e1-652a144dcd75 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.237219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.237219Z digest=sha256:1c94dc23079533ceaacd828bae1163eec9c1b25628a0dc48f84c25c4d28c749a

Observation 4440c7bc-d19d-4091-a2cf-65cfecebc704 · outbound

This paper cites Xlnet: Generalized autoregressive pretraining for language understanding,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Xlnet: Generalized autoregressive pretraining for language understanding,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.240658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.240658Z digest=sha256:298d992db5709e75ab730252c21c90e7222153231e65b756cbb6efdfb874ceff

Observation 34821e7a-2354-4cba-853e-7974fa80fe15 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.243213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.243213Z digest=sha256:fe60d4912766673ac8968277d848febb75208bd086571f737785ef3930266e1a

Observation d145997b-c7d3-446b-80e6-fb8425bac0f5 · outbound

This paper cites Language models are unsupervised multitask learners,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Language models are unsupervised multitask learners,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.246182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.246182Z digest=sha256:8877f5c6134e443a135fbcc965b43f7bc3870350316fe8361f96bc93262c659f

Observation 21a68326-a51d-4cf2-9ef0-cb58780bbe26 · outbound

This paper cites Language mod- els are few-shot learners,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Language mod- els are few-shot learners,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.248941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.248941Z digest=sha256:9cc3a3bf1005a697c6cb1ac405c3811bf25c31a9f65cb30154764edd2f3a1824

Observation aff407bc-51ed-4917-b305-4a238d19ef69 · outbound

This paper cites Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.252688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.252688Z digest=sha256:16b6174573e6151cd5a924b9a83479c9f39924845368459d46ad9af367659c78

Observation 3809ba1a-448b-47f8-b717-3d5b47fde7ba · outbound

This paper cites Mixture of A Million Experts.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Mixture of A Million Experts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.257132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.257132Z digest=sha256:7dc63ea1f14c85fcd7a95b8899ffd64f231086c805e5072bc7dd7dcca76e09ac

Observation 52602f3f-2653-4d8f-b858-a60f96c09ede · outbound

This paper cites Taming Sparsely Activated Transformer with Stochastic Experts.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Taming Sparsely Activated Transformer with Stochastic Experts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.260529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.260529Z digest=sha256:bf9ac2f750131598976762013445aa6a71adfbb6750c0715f5b9cff90bc78082

Observation fc4f260d-5b54-4dd9-a9aa-07fe5e8455b5 · outbound

This paper cites Scaling laws for fine-grained mixture of experts,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Scaling laws for fine-grained mixture of experts,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.752028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:10:00.263578Z digest=sha256:bcf333d8e5652bc117694f16ef3475e867745622534c8e8aff232abc825b767a

Observation e9b2b9de-b84d-42fa-82f9-8122320186ae · outbound

This paper cites Sparse Upcycling: Training Mixture-of-Experts from Dense Checkpoints.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Sparse Upcycling: Training Mixture-of-Experts from Dense Checkpoints

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.266892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.266892Z digest=sha256:5abdf44b75c1e27833fc83cd8b9759066bf64b314b196dd29e3ce65787cd9b1a

Observation 290384aa-3190-4961-b0ae-38e50cd62f5c · outbound

This paper cites Go wider instead of deeper,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Go wider instead of deeper,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.270242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.270242Z digest=sha256:f0562e03d74e61465a6ded867c44f4c4cd98d8e4eaa478df045fd9615938526f

Observation cc186b9a-7c60-4240-9bf8-99c451518807 · outbound

This paper cites One Student Knows All Experts Know: From Sparse to Dense.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models One Student Knows All Experts Know: From Sparse to Dense

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.272918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.272918Z digest=sha256:ec0a1cf429b31ef96635a13cc9fecef19575f3693c647e5a593f6fe7c8adc516

Observation 38ee513d-0e03-42ba-b6f2-90914a5cf819 · outbound

This paper cites ST-MoE: Designing Stable and Transferable Sparse Expert Models.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models ST-MoE: Designing Stable and Transferable Sparse Expert Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.275906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.275906Z digest=sha256:850d147f9a7ff1fe7422f275a28814de9b92c2c99a5c7bb040823daa506d3cde

Observation 2d53ee6d-218c-4ed4-aa53-ae7d224fd501 · outbound

This paper cites GPT-4 Technical Report.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models GPT-4 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.279105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.279105Z digest=sha256:610ef05971ca318a70dc9f1e7d82d065b159f396bc2ea58e3108c7e8153f929e

Observation 699f02f3-d1f9-4b75-8533-38a0486d714c · outbound

This paper cites Stabilization of planar collective motion: All-to-all communication,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Stabilization of planar collective motion: All-to-all communication,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.736945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:10:00.282279Z digest=sha256:4fc62c96bb720fe9f7bbbcd999f462b055b8b8f3c0e6d4c099e90aea471a321c

Observation 8c2d3890-f7a6-4d17-bd70-7f01b8af434f · outbound

This paper cites Optimization of all-to-all communication on the blue gene/l supercomputer,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Optimization of all-to-all communication on the blue gene/l supercomputer,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.727086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:10:00.285105Z digest=sha256:c68b009de1b8ac93246907c3d4c590ae1a24364ccf51d004303d640d676b8060

Observation f37a21b1-5643-4dd9-9d6c-4b0286378d15 · outbound

This paper cites The hierarchical factor algorithm for all-to-all communication,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models The hierarchical factor algorithm for all-to-all communication,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.718422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:10:00.288038Z digest=sha256:99f2a2e16a0f89dba8e7047db837354976889b59e6e03b5ccec71f291f9e6bb3

Observation af6da1f4-89f5-4d0c-8aec-55ac85df5188 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.290731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.290731Z digest=sha256:a9d5427b59bce7c935dedc069c24a1c40f9eacbc8ec36be417c2111d346e965e

Observation 7ebe3cf2-a672-4597-8f35-550841df79cc · outbound

This paper cites HetuMoE: An Efficient Trillion-scale Mixture-of-Expert Distributed Training System.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models HetuMoE: An Efficient Trillion-scale Mixture-of-Expert Distributed Training System

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.293880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.293880Z digest=sha256:834dddd656249bb94b4c8f9e8a781ba00bdcc79e6f3a1125086182a5f375782a

Observation 826989a0-251e-4486-b7c8-e460c0678026 · outbound

This paper cites FastMoE: A Fast Mixture-of-Expert Training System.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models FastMoE: A Fast Mixture-of-Expert Training System

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.296977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.296977Z digest=sha256:57bf0d543089ec024791539226115e731cb0c81372d3b4d169a2dcc0e93bf55d

Observation 80806519-2428-4747-949b-66856a7bfbb3 · outbound

This paper cites Moesys: A distributed and efficient mixture-of-experts training and inference system for internet services,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Moesys: A distributed and efficient mixture-of-experts training and inference system for internet services,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.300195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.300195Z digest=sha256:474300974ce56931292b803158e3675ac3ff428011e0eb492482dc1defd1b48c

Observation 9f26843c-db03-4bab-92db-f7371fc18d86 · outbound

This paper cites Accelerating distributed moe training and inference with lina,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Accelerating distributed moe training and inference with lina,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.704177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:10:00.302914Z digest=sha256:d5a6cfa15291011dbea06295b7425254c572e5c8f090e531b7b2114de8b9b01f

Observation dd018d99-45b1-47b4-9e39-d869372a3e2e · outbound

This paper cites Janus: A unified distributed training framework for sparse mixture-of-experts models,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Janus: A unified distributed training framework for sparse mixture-of-experts models,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.695266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:10:00.305921Z digest=sha256:88e56508acc3517738a3ae331175f9fec6ef83c1533fd1bd57a0a1890f49c8a5

Observation 7e2bfbb1-5918-4be1-88c8-a3c798fc9ed9 · outbound

This paper cites Amp: Automatically finding model parallel strategies with heterogeneity awareness,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Amp: Automatically finding model parallel strategies with heterogeneity awareness,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.686137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:10:00.308883Z digest=sha256:0bdce8c70bd939b9feadba533a6ac9173f800026542d11581ddf4c411d331d3f

Observation 9f5c84a0-ef36-470f-be87-de44ef55257a · outbound

This paper cites Merak: An efficient distributed dnn training framework with automated 3d parallelism for giant foundation models,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Merak: An efficient distributed dnn training framework with automated 3d parallelism for giant foundation models,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.311581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.311581Z digest=sha256:6bcb9a529eeaaaaba1bb83fdb9b6c39c6d7194fefa6c820458db6b932e3139d1

Observation 46c22993-8ac0-47e4-a2ca-434d3d1db28e · outbound

This paper cites Hippie: A data-paralleled pipeline approach to improve memory-efficiency and scalability for large dnn training,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Hippie: A data-paralleled pipeline approach to improve memory-efficiency and scalability for large dnn training,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.672088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:10:00.314245Z digest=sha256:f5c36355e28d649d7bb9eeaea4db5b9716373e33deedb0048b90e0b576ab3d7d

Observation 360af553-4124-49e6-bedf-1f840cc44f51 · outbound

This paper cites Colossal-ai: A unified deep learning system for large-scale parallel training,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Colossal-ai: A unified deep learning system for large-scale parallel training,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.317161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.317161Z digest=sha256:5aafbeda18e9413d5bdb7c60a27c9d11916c824e4ca2db1c6b32fe6b0d9836b0

Observation 6e0242cb-1274-4526-880d-b762bbc4d5af · outbound

This paper cites Piper: Mul- tidimensional planner for dnn parallelization,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Piper: Mul- tidimensional planner for dnn parallelization,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.658455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:10:00.320208Z digest=sha256:3f40eb784933156eb0be969e59aa3c1f5acc2e4880c44c8af1e8c7221e4f3dab

Observation b951c0a5-908f-4ecc-bc98-899fa3472acf · outbound

This paper cites Parallel intelligent computing: development and challenges,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Parallel intelligent computing: development and challenges,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.649130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:10:00.322933Z digest=sha256:b2a0c8b3d95e402631ba0d26eeb123c28086f047b1053736848cadb842c5ac2e

Observation 8d3a6b4a-c378-4a98-b349-19dc4cfa6be5 · outbound

This paper cites Zero- infinity: Breaking the gpu memory wall for extreme scale deep learning,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Zero- infinity: Breaking the gpu memory wall for extreme scale deep learning,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.639818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:10:00.325797Z digest=sha256:c86003451e1bd4446c3e3d69a31baffaf729b536dd4429f8633f1cef0e58365f

Observation 72044f93-4f03-47a2-b08f-2c859365b88a · outbound

This paper cites Zero-offload: Democratizing billion-scale model training,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Zero-offload: Democratizing billion-scale model training,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.630265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:10:00.329512Z digest=sha256:d120e46511b1a3d135b264d7a2e5b8ed2e0aefceb2dc0a22a0ae2b36a123fe03

Observation 10ce9216-144c-4931-bc69-a483136a8b34 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.332429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.332429Z digest=sha256:30466b0d03241636214a05db47987b369ee81eac22bc4cd87529af3cc84c889a

Observation 51bd6a73-b9cd-4b03-8802-1d55e0763222 · outbound

This paper cites MiCS: Near-linear Scaling for Training Gigantic Model on Public Cloud.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models MiCS: Near-linear Scaling for Training Gigantic Model on Public Cloud

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.335691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.335691Z digest=sha256:72cc15071cc2a5cfc193983c9c947553ec89d25a69509b2417787f9716970a9b

Observation c05e997e-1f81-4404-a78f-22131273c31c · outbound

This paper cites Maximizing Parallelism in Distributed Training for Huge Neural Networks.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Maximizing Parallelism in Distributed Training for Huge Neural Networks

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.338766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.338766Z digest=sha256:0d5313d8d2819d5201935cb8ecd3a37af564b556f8de65561f739caf0b6a5ff3

Observation 9b97b545-503f-4084-9ef2-f3196c6a566f · outbound

This paper cites Autopipe: A fast pipeline parallelism approach with balanced partitioning and micro- batch slicing,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Autopipe: A fast pipeline parallelism approach with balanced partitioning and micro- batch slicing,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.620566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:10:00.342063Z digest=sha256:856e8b8cc0e7a04ac96c3eb9f8b9fbfad2c54f1647a3d2a069007a4346103e63

Observation 24de2ed7-f199-4900-a52d-3a1c7e66e0e0 · outbound

This paper cites Hph: Hybrid parallelism on heterogeneous clusters for accelerating large-scale dnns training,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Hph: Hybrid parallelism on heterogeneous clusters for accelerating large-scale dnns training,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.611602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:10:00.345381Z digest=sha256:f59311b76876eb72e0205140936403d78e5767493dcf2ba44ddae19b6e3cb22c

Observation 85eae24b-9be0-41e4-9ced-6facda989775 · outbound

This paper cites Sequence Parallelism: Long Sequence Training from System Perspective.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Sequence Parallelism: Long Sequence Training from System Perspective

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.348204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.348204Z digest=sha256:1dba35a20132110138e723c72125199e118b0468974969e57c619cee76513983

Observation ae5eb7df-23fc-48f6-86c1-912b5b2efea5 · outbound

This paper cites Reducing activation recomputation in large transformer models,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Reducing activation recomputation in large transformer models,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.351017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.351017Z digest=sha256:5238acf69e2074d6ad710b858ae62513ed85f7631305899adc35e1afdf50dd0c

Observation 459f212d-2063-40f8-a3d1-cb5d50a95469 · outbound

This paper cites DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.354500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.354500Z digest=sha256:a27b32fe837a42754f9357f4dd7ca06ef01da8c5db7b98d5b259b4a646a04eb7

Observation c2d30e42-31d6-44f3-ace7-11f1858ea94a · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.357780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.357780Z digest=sha256:d6c267c1598b6114f8aedfaa5d85deee08d79c8d433a3aebaa49547d6957825e

Observation 0b245745-c6a7-491f-bf00-0ca0fe2d9f0c · outbound

This paper cites Bagualu: targeting brain scale pretrained models with over 37 million cores,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Bagualu: targeting brain scale pretrained models with over 37 million cores,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.361017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.361017Z digest=sha256:bdcc768f33c189dd2dbf5bf654891715a4fd8ad3fb271229fe58e207f3770c9f

Observation 90b3f15f-2ac1-461d-a45f-6634f43a9305 · outbound

This paper cites Parm: Efficient training of large sparsely-activated models with dedicated schedules,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Parm: Efficient training of large sparsely-activated models with dedicated schedules,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.593618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:10:00.363917Z digest=sha256:39e345efae1bc590003ac737e75dea7a7e726e9883eb388a099950ecc1f9f1ea

Observation 170fcd03-b062-4ae9-8564-7850af247d9e · outbound

This paper cites A hybrid tensor-expert-data parallelism approach to optimize mixture- of-experts training,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models A hybrid tensor-expert-data parallelism approach to optimize mixture- of-experts training,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.367178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.367178Z digest=sha256:b45c2af2ada7cd89a172314da3f916c0ef817c1e8a3f3b07a8bff967a8805784

Observation 8b1933bb-90f3-4d74-878b-0bbc965268e7 · outbound

This paper cites Mg-wfbp: Efficient data communication for distributed synchronous sgd algorithms,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Mg-wfbp: Efficient data communication for distributed synchronous sgd algorithms,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.579108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:10:00.370350Z digest=sha256:fb5a9f4912f8440872ad58b5eae1ff242c628cec5793e634f8474eec0b14b5af

Observation 330a99b6-790c-430b-ace4-199aeb2e85b3 · outbound

This paper cites PipeTransformer: Automated Elastic Pipelining for Distributed Training of Transformers.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models PipeTransformer: Automated Elastic Pipelining for Distributed Training of Transformers

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.372977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.372977Z digest=sha256:f30317111d52dc069e2a225b25cd3cc47bdc1bd999a16ddb64c991e0f6bd90bf

Observation fe168bc5-30cb-4bf0-a885-1718896f5399 · outbound

This paper cites A multidimensional communication scheduling method for hybrid parallel dnn training,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models A multidimensional communication scheduling method for hybrid parallel dnn training,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.570645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:10:00.376113Z digest=sha256:effbd30381f35ff98e6d288aae318ba9cfee9c5496b86ba270f8801b9b8ae1dd

Observation b1c76ec9-0a12-47b8-96bd-a305aff7edc3 · outbound

This paper cites Schemoe: An extensible mixture-of-experts distributed training system with tasks scheduling,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Schemoe: An extensible mixture-of-experts distributed training system with tasks scheduling,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.379017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.379017Z digest=sha256:2eb6033dd23673d045f2fb6c0588cad29f8a25950346ee0b4c5d9b5ad68dae9b

Observation 7b64b832-e524-4522-b65d-dd22fe32f106 · outbound

This paper cites Pipemoe: Accelerating mixture- of-experts through adaptive pipelining,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Pipemoe: Accelerating mixture- of-experts through adaptive pipelining,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.382168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.382168Z digest=sha256:e5393f1ca843bfb5a8514e89d9f51f839ab1e7c6bfacf8157b556754d63d7a3a

Pith citing papers

No inbound Pith citation observations are available.