Pith. sign in

Paper Citation Record · LEDGER

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism

As of 15 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2506.22175.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22175 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:15:28.228455Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact3
  • verified fuzzy20
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 80e302e1-70b3-48e4-9fb2-efa3d16c2011 · outbound

This paper cites On the optimization of deep networks: Implicit acceleration by overparameterization,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism On the optimization of deep networks: Implicit acceleration by overparameterization,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:32.788377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:25.501953Z digest=sha256:f462f6b849acd79327b9f43eb2ba157614af07831de5c252f6cc2f470382b73f

Observation da7e2b1f-e19f-4471-8b74-013bcc6e3fac · outbound

This paper cites Exploring the limits of weakly supervised pretraining,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Exploring the limits of weakly supervised pretraining,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:32.565005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:25.611239Z digest=sha256:e5566614b2442c1cb9204006dda3524404063b38edeb6fd72250fa801632da12

Observation c0832550-adb5-4c73-999e-1b1107862a00 · outbound

This paper cites Antman: Dynamic scaling on gpu clusters for deep learning,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Antman: Dynamic scaling on gpu clusters for deep learning,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:32.311875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:25.755050Z digest=sha256:2643ea69b21dc75497d43707acf0991bc6eaf520bdeb2ace011576b0caf10f60

Observation b4e4428f-2f87-4a99-a327-3399924419b9 · outbound

This paper cites Whale: Efficient giant model training over heterogeneous gpus,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Whale: Efficient giant model training over heterogeneous gpus,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:32.114853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:25.901684Z digest=sha256:4a61496f9eb4f2b6ed6f4ddde297d57b39cb995f54a2d68ce0cd639432d29e76

Observation 4a45c5aa-fed0-4ad2-903a-aa8ea74b4ac7 · outbound

This paper cites Axonn: An asynchronous, message-driven parallel framework for extreme-scale deep learning,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Axonn: An asynchronous, message-driven parallel framework for extreme-scale deep learning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:31.870218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:26.037612Z digest=sha256:9ceef13f41e1e52609445588271da77a7b805360ac5638ac613ab5f2f6190642

Observation 202fd7a0-346f-4d78-b43d-6d7a2e0321a8 · outbound

This paper cites An efficient and non-intrusive gpu schedul- ing framework for deep learning training systems,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism An efficient and non-intrusive gpu schedul- ing framework for deep learning training systems,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:31.651693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:26.141093Z digest=sha256:bfcec1172e87fe6fb47c338ff8235ab13a90591003a05b4eca2578f830cc6c96

Observation f4fe10b9-0f1b-4450-9e60-b23b2571da6f · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Bert: Pre-training of deep bidirectional transformers for language understanding,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:31.543830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:26.269023Z digest=sha256:62b00ae6f390cbdcadb05cf78bd50b27500ad3cd938fe201b8fe7b5fd56d917e

Observation ace6e3a0-ef11-45f4-a95e-78859d2aefd2 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:26.398427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:26.398427Z digest=sha256:d29f02516d646d2ef933da4c5132cac3556de8959703d5d688942de5ab1435ff

Observation 5ea1faec-ca1e-4b18-8d4d-900f39b7ebea · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:31.337050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:26.502541Z digest=sha256:7acd6c529e09bcf48b5949b868ebd628431edfc2e1a57d08299d747c63221261

Observation c0368665-aa1a-4d71-a103-10d3923d409e · outbound

This paper cites Language mod- els are few-shot learners,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Language mod- els are few-shot learners,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:26.600395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:26.600395Z digest=sha256:eb85cc55345999b37ac0aa3a4df619ee9d51c4728a770c1f3df414fdebe2bdff

Observation 9e9724b0-cf24-495c-a834-667601713bd7 · outbound

This paper cites BLISS: Robust Sequence-to-Sequence Learning via Self-Supervised Input Representation.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism BLISS: Robust Sequence-to-Sequence Learning via Self-Supervised Input Representation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:15:28.679030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:26.715460Z digest=sha256:cf1e03c866b184c0157c324f0b285029dd2bbe43a49dca50617f102a009354ab

Observation 9129e399-e0bb-4626-8d95-36dbff9f54fe · outbound

This paper cites E2S2: Encoding-Enhanced Sequence-to-Sequence Pretraining for Language Understanding and Generation.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism E2S2: Encoding-Enhanced Sequence-to-Sequence Pretraining for Language Understanding and Generation

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:15:28.537261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:26.817024Z digest=sha256:35813e8681cc2f030ec513b804fadfadddd01bda2b414b6fb1af33758f6ccf7b

Observation a588efc2-b4fd-4dd4-8bdc-8dc75e06d433 · outbound

This paper cites Unsu- pervised cross-lingual representation learning at scale,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Unsu- pervised cross-lingual representation learning at scale,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:31.199126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:26.871985Z digest=sha256:de8a3328e89e6ac5e786774c4810b446e32f9a6ad7ac66de209dda292fde30c0

Observation c51cf98b-e571-425f-bbf9-cb073dfb0325 · outbound

This paper cites Toward Efficient Language Model Pretraining and Downstream Adaptation via Self-Evolution: A Case Study on SuperGLUE.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Toward Efficient Language Model Pretraining and Downstream Adaptation via Self-Evolution: A Case Study on SuperGLUE

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:26.922616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:26.922616Z digest=sha256:f775484f9a48f52b1a4f9e8e80981888552aa8d5cf507de9ecdc4c23b2243a34

Observation 45b1a33d-616a-487a-bccb-fb8ccd773284 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:26.970238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:26.970238Z digest=sha256:686cbe286c3dbcd26e2bf049f94e4a0935aeb9933c84dbba58087a4ca38df64f

Observation 1e365e47-038f-4936-851c-975894fcde0d · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:27.043950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:27.043950Z digest=sha256:5f63b674cd4f4c23410e3cd465602257c2b3a57078020a9bb9004264cfa64627

Observation 25121131-c2c5-4c52-98d2-7cfd717c1894 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:30.997680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:27.110599Z digest=sha256:2ead37c34aaf2df13740f45de5a02b2264300e946162d232017412141898b44b

Observation 52d7527b-358b-4360-af2c-cbe2b11d13a6 · outbound

This paper cites PAD-Net: An Efficient Framework for Dynamic Networks.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism PAD-Net: An Efficient Framework for Dynamic Networks

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:15:28.401357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:27.198838Z digest=sha256:cfcc21b9beebfaf52f9247155badcc393358d399ef664c4a49d960cdb176ce57

Observation d1ab2883-4be9-4ebd-8bcc-18f2bfd914b5 · outbound

This paper cites Base layers: Simplifying training of large, sparse models,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Base layers: Simplifying training of large, sparse models,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:30.739187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:27.260718Z digest=sha256:21cac81a978579d06d0eedf8dad01ca1bc7ac51c46aec43b5a388460b74a4b12

Observation 81121fdc-711e-4918-8173-af010a61e654 · outbound

This paper cites Gating dropout: Communication-efficient regularization for sparsely activated transform- ers,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Gating dropout: Communication-efficient regularization for sparsely activated transform- ers,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:30.504450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:27.323226Z digest=sha256:4cf95000ec68dd23210f18073787fe5ee6d29e99886b58825d2190eb8b199535

Observation f5762c01-f99e-42bd-903b-cb895bac91df · outbound

This paper cites Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation AI scale,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation AI scale,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:30.256375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:27.397140Z digest=sha256:f25a3a811b55b87a015d8fc7ef5f78aeaed3babcb1d96445f79e8a03c1d94cb5

Observation 10ad3246-ce10-4f75-8cf0-b201c7eae5b4 · outbound

This paper cites Fastermoe: modeling and optimizing training of large-scale dynamic pre-trained models,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Fastermoe: modeling and optimizing training of large-scale dynamic pre-trained models,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:30.040462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:27.451658Z digest=sha256:fd998e2a3d41313f9195e89e4dcb022c94382d143b7047d1bdb0cda84f9238ec

Observation da501c63-0f1f-4977-a420-cf8acc594f93 · outbound

This paper cites Scalable distributed dl training: Batching communication and computation,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Scalable distributed dl training: Batching communication and computation,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:29.822345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:27.495731Z digest=sha256:63a0a026ce5227f40a3eaff8503484f4f4cdcdc16c193325dbf5b7fd43079fc4

Observation 77bc7860-4971-447f-867d-c7965f20e622 · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Zero: Memory optimizations toward training trillion parameter models,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:29.593956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:27.557041Z digest=sha256:dd37070b5c1e8c393036f03f0cf72e1bb0b27d5406623adacd1da3b325cdbc85

Observation 89d2c24a-c7ce-4c6e-b494-0c6e08c1f0db · outbound

This paper cites Scalable and Efficient MoE Training for Multitask Multilingual Models.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Scalable and Efficient MoE Training for Multitask Multilingual Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:27.618354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:27.618354Z digest=sha256:72ab404c035fce70921838169a7c62ef3fe3ad0942853e90c2ec7e11558f4db4

Observation e3157efc-d1e2-4f18-928f-81e7d59f0558 · outbound

This paper cites Training Deep Nets with Sublinear Memory Cost.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Training Deep Nets with Sublinear Memory Cost

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:27.666690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:27.666690Z digest=sha256:2b06d22475e854b8e58034ba4a91ceca3bdc0fd05ae6f5ef83f4d016b199fd3d

Observation 61131dc7-3ae8-4b15-b543-4815bf7e6cda · outbound

This paper cites vdnn: Virtualized deep neural networks for scalable, memory-efficient neural network design,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism vdnn: Virtualized deep neural networks for scalable, memory-efficient neural network design,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:29.248516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:27.733735Z digest=sha256:0af2be6395da8a84b110d29a5cc7c858dabfbb56dd0bca3262d2ce2e76f9bf34

Observation 954749e8-cf15-4840-8bcd-07b8fddcad2b · outbound

This paper cites Buddy compression: Enabling larger memory for deep learning and hpc workloads on gpus,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Buddy compression: Enabling larger memory for deep learning and hpc workloads on gpus,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:29.046713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:27.805975Z digest=sha256:241e82ad9da40a3cfe9b4cc4a7222637e169612acc6baccf09c59b635dec90b5

Observation 0637980d-29df-4b1c-b823-c91db9e0624f · outbound

This paper cites Efficient large-scale language model training on gpu clusters using megatron-lm,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Efficient large-scale language model training on gpu clusters using megatron-lm,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:28.945997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:27.873864Z digest=sha256:d0b6f1bc0859201162b4e61fb81fe54b72647dd7a627126600281aa648fddc81

Observation 5f77f9e4-68da-46a4-9a5b-fe15d0db598c · outbound

This paper cites Adam: A method for stochastic optimization,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Adam: A method for stochastic optimization,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:27.918794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:27.918794Z digest=sha256:329130aad6478fb1f32a0b572e4132ba6bf9317db961e353a2b5192da3c99bcb

Observation 1df6a646-9d37-49d9-94b6-61dbb4a70cfb · outbound

This paper cites Gpipe: Efficient training of giant neu- ral networks using pipeline parallelism,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Gpipe: Efficient training of giant neu- ral networks using pipeline parallelism,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:27.975062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:27.975062Z digest=sha256:1216e10417ca66294b216ff09c61da849b5c0b6ca3e2e07d4d3d50d1bcb07ce2

Observation cc6a1c04-8cc7-41c1-9dfd-bd462064a2e0 · outbound

This paper cites Tutel: Adaptive Mixture-of-Experts at Scale.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Tutel: Adaptive Mixture-of-Experts at Scale

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:28.050477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:28.050477Z digest=sha256:cb5cc33ba8770cc39ffd3f33e8a63ca5f76dbbe2328239ed6ed8ec222b134528

Observation 74773f93-26a2-4e1a-bccc-250a63346df8 · outbound

This paper cites Mesh-tensorflow: Deep learning for supercomputers,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Mesh-tensorflow: Deep learning for supercomputers,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:28.801078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:15:28.110425Z digest=sha256:0e0fa06e4cedb13c594d1c4d277e7cfd7dbb381ce02145aab5ccd26aadf4360e

Observation 0cc0be77-5179-45f1-b922-db552ce61723 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:28.158421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:28.158421Z digest=sha256:5963d60901f5691b7e0ba0a7bf253fe7fe79d5ae3d5e491c460f3c0c523f12b6

Observation fcc388e8-10f4-48cc-a89d-78c841b56b78 · outbound

This paper cites PipeDream: Fast and Efficient Pipeline Parallel DNN Training.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism PipeDream: Fast and Efficient Pipeline Parallel DNN Training

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:28.228455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:28.228455Z digest=sha256:81ddd84d187ab00ae5ad35429347bc2ffd1a5f119ebc95d6a23f0abd306ed48d

Pith citing papers

No inbound Pith citation observations are available.