Pith. sign in

Paper Citation Record · LEDGER

FastMoE: A Fast Mixture-of-Expert Training System

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2103.13262.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2103.13262 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:52:07.475849Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T10:47:02.152647Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c7019b44-4439-473e-9fe7-b5d9f57b046b · inbound

A Survey of Large Language Models cites this paper.

A Survey of Large Language Models FastMoE: A Fast Mixture-of-Expert Training System

Reference 211

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:46:39.997102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T22:46:39.268353Z digest=sha256:fce109ae409ee717cc36cd45d34b1baf1290191523457521e3b8dea1e206efac

Observation e463f93a-f277-4bcc-bbdd-44c1eb3066f8 · inbound

A Comprehensive Overview of Large Language Models cites this paper.

A Comprehensive Overview of Large Language Models FastMoE: A Fast Mixture-of-Expert Training System

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:28:39.310778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T20:28:38.900026Z digest=sha256:52be1d5df036e6576a632c7cad1a4d6c810cdb6ccda4824a536e21ee814aa842

Observation 3aed77df-0255-4bb4-a6b4-d1179ddc6545 · inbound

DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models cites this paper.

DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models FastMoE: A Fast Mixture-of-Expert Training System

Reference 172

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:07:22.306886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T01:07:22.166595Z digest=sha256:93f32b3da074303b96871da72d2f482793ffa8bdf41512e7bb9ea0e1c3ab4ba5

Observation 25c84bbd-c5fb-4cb7-a2a6-361beccaed4f · inbound

TabICL: A Tabular Foundation Model for In-Context Learning on Large Data cites this paper.

TabICL: A Tabular Foundation Model for In-Context Learning on Large Data FastMoE: A Fast Mixture-of-Expert Training System

Reference 218

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:35:02.311552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T13:35:02.018244Z digest=sha256:85b79ae84ef3af9796646fbf3d2eac5e24a29cbe78f7a69433d087748c6a979b

Observation 6a27268a-ffa9-45a5-abaf-bc11ad728dae · inbound

MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing cites this paper.

MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing FastMoE: A Fast Mixture-of-Expert Training System

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T14:52:07.475849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:52:07.475849Z digest=sha256:bf8225df0bd36d7b9682acbadc1bc8c90d21cd6c93ef2361229c43105627ba7d

Observation caec5c3e-4025-450a-a09e-6bd5c6ccc09e · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models FastMoE: A Fast Mixture-of-Expert Training System

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:36:23.955606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:c687c4510a2bc49e4daa39d9b670151cfbe12a2f54d39e1006091fbb1289f1c0

Observation d73357d5-c507-4370-9956-4e1199b6fa67 · inbound

Two Is Better Than One: Rotations Scale LoRAs cites this paper.

Two Is Better Than One: Rotations Scale LoRAs FastMoE: A Fast Mixture-of-Expert Training System

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:57:43.646076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:57:43.646076Z digest=sha256:55ecdb5e0b4259ce231b6fcd1767b2d581a95acbc9ae86a370dad79ce44bdc17

Observation 22486253-f1f6-422c-88ad-715fb65d80bc · inbound

MoE-GPS: Guidlines for Prediction Strategy for Dynamic Expert Duplication in MoE Load Balancing cites this paper.

MoE-GPS: Guidlines for Prediction Strategy for Dynamic Expert Duplication in MoE Load Balancing FastMoE: A Fast Mixture-of-Expert Training System

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:39.404771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:39.404771Z digest=sha256:743a1b07cf647e0e926bc7291b68a00ecb7c67a1b02bd9eecf351a10002e8cf1

Observation d4340967-aa22-4a69-aad8-fd6b8cf1e4c5 · inbound

HarMoEny: Efficient Multi-GPU Inference of MoE Models cites this paper.

HarMoEny: Efficient Multi-GPU Inference of MoE Models FastMoE: A Fast Mixture-of-Expert Training System

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.858076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.858076Z digest=sha256:ff04f25537173f7155eef31517ac1a0803abd5a143d604102c0748c46472cee4

Observation 8e5341b7-536e-46e5-8b90-21ddb068f966 · inbound

Towards Accurate and Efficient 3D Object Detection for Autonomous Driving: A Mixture of Experts Computing System on Edge cites this paper.

Towards Accurate and Efficient 3D Object Detection for Autonomous Driving: A Mixture of Experts Computing System on Edge FastMoE: A Fast Mixture-of-Expert Training System

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.072590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:29.072590Z digest=sha256:e4cd215427f0f3a449777b94b02b690448f047ee9cae5fdfac3ac678b4873303

Observation 8f82d564-9f3c-467a-876f-ff097377227b · inbound

BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity cites this paper.

BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity FastMoE: A Fast Mixture-of-Expert Training System

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:00.189812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:20:00.189812Z digest=sha256:fea53752c033b516a7823f7004b936a8d64344094c62ccbeccdb85ab97cce182

Observation 90764af7-1ac7-4b06-8082-f40da840398e · inbound

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning cites this paper.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning FastMoE: A Fast Mixture-of-Expert Training System

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:41.941220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:41.941220Z digest=sha256:cf2a0f1f74659e260862e8aec27040ee232567c093c3dfa3184dae0d1cba8093

Observation 79ae7997-4a84-4a1b-a01a-b522e7a2786d · inbound

AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMs cites this paper.

AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMs FastMoE: A Fast Mixture-of-Expert Training System

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:31:17.988693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-09T19:46:13.015064Z digest=sha256:9b8b01e33fb7d929e7ae84bcd391b1fdfb840cd200ce2e7a0410e67ca2a6f1d0

Observation 8365c891-d4f8-42c2-a69e-9252d4686c31 · inbound

AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMs cites this paper.

AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMs FastMoE: A Fast Mixture-of-Expert Training System

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:21:30.567351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T05:17:09.793360Z digest=sha256:3429f883a124d45a25c83aa0503bd79b763d415c163db5f8d60f460706f9a7c6

Observation c7825e9c-b3cf-4d82-9b09-622930c946f0 · inbound

MoE-Prefill: Zero Redundancy Overheads in MoE Prefill Serving cites this paper.

MoE-Prefill: Zero Redundancy Overheads in MoE Prefill Serving FastMoE: A Fast Mixture-of-Expert Training System

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T05:50:26.994223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T19:29:28.916831Z digest=sha256:a063006f5e3df70f056c595185573f14aa890508d618389cb8a781464278cd00

Observation d14e1f5f-ceca-4087-8593-731afe4b1243 · inbound

MoE-Prefill: Zero Redundancy Overheads in MoE Prefill Serving cites this paper.

MoE-Prefill: Zero Redundancy Overheads in MoE Prefill Serving FastMoE: A Fast Mixture-of-Expert Training System

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T17:47:41.970798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T17:43:56.767714Z digest=sha256:9610cdaa0acab80619def4c774854fda4a2ae9d8006bd0d3f2cb4bd65bae673d

Observation bb418645-f69b-465f-aabc-c9edaeedf98b · inbound

Hierarchical Mixture-of-Experts with Two-Stage Optimization cites this paper.

Hierarchical Mixture-of-Experts with Two-Stage Optimization FastMoE: A Fast Mixture-of-Expert Training System

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:01:15.400906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T01:58:04.218139Z digest=sha256:03c811e43776fdc9107f5fc0c1ce9aee4a9e1978a88705a314c8f720e7a4f07c

Observation 39caa9a9-2175-4036-9aec-4f53d3543cf3 · inbound

Fast MoE Inference via Predictive Prefetching and Expert Replication cites this paper.

Fast MoE Inference via Predictive Prefetching and Expert Replication FastMoE: A Fast Mixture-of-Expert Training System

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:27:06.992817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T02:25:59.268868Z digest=sha256:5b26a214ad1e6f090d7fcd46ee88ed8a15c6fbe13e9a2e0137bfa07f734742de

Observation 18fcc67d-21a3-4bb4-978c-03ae259594ac · inbound

EMO: Frustratingly Easy Progressive Training of Extendable MoE cites this paper.

EMO: Frustratingly Easy Progressive Training of Extendable MoE FastMoE: A Fast Mixture-of-Expert Training System

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:07:53.589798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T20:06:33.816250Z digest=sha256:75d298e3c28f7b1f2be3dce0a3cd175f7272cbf7eff31d57ecb6196e54f8005d

Observation 609fc325-8fc2-44b9-b51e-d6daba8d4007 · inbound

EMO: Frustratingly Easy Progressive Training of Extendable MoE cites this paper.

EMO: Frustratingly Easy Progressive Training of Extendable MoE FastMoE: A Fast Mixture-of-Expert Training System

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:35:04.303635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T05:32:10.642355Z digest=sha256:e38059c6b89b81120ef4ecd9e9c6f81810b26a44773c93e4451db6ee0ba9382d

Observation f1bb5ced-086e-4759-8a61-6d9a87a84c41 · inbound

ViBE: Co-Optimizing Workload Skew and Hardware Variability for MoE Serving cites this paper.

ViBE: Co-Optimizing Workload Skew and Hardware Variability for MoE Serving FastMoE: A Fast Mixture-of-Expert Training System

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:46:13.242761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T18:07:52.038640Z digest=sha256:cfa948c8937acb71ca5deac098e0d38677cb24f8ba0da2975345eb1e03e2d1a1

Observation d4629689-ac79-4e57-82d9-f7b0836d5666 · inbound

Language-Assisted Super-Resolution from Real-World Low-Resolution Patches cites this paper.

Language-Assisted Super-Resolution from Real-World Low-Resolution Patches FastMoE: A Fast Mixture-of-Expert Training System

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:05:41.498921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-01T05:49:20.689248Z digest=sha256:37aac35fd3a470509cfd88995d0650320170274c7e94abd5a9c3d10dcceb0dde

Observation 28101062-bbed-4856-a13e-525275d988c6 · inbound

Language-Assisted Super-Resolution from Real-World Low-Resolution Patches cites this paper.

Language-Assisted Super-Resolution from Real-World Low-Resolution Patches FastMoE: A Fast Mixture-of-Expert Training System

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T22:18:59.533997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-03T22:14:30.734906Z digest=sha256:8d64d20bea505a1695779702141d485fa7a2bba51b57724a9f8a407ff4bd1921

Observation d503ead8-7bd4-46ff-b381-6b01c11a7451 · inbound

BrainFIBRE: A Foundation Model via Information Decomposition for Brain Microstructure cites this paper.

BrainFIBRE: A Foundation Model via Information Decomposition for Brain Microstructure FastMoE: A Fast Mixture-of-Expert Training System

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T15:07:04.023990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-02T15:00:39.357350Z digest=sha256:14ead32484e85af811179c59026e2a19ed89862f57887501613e689f5c36615d

Observation 999c85a5-de8c-4ad3-bcdc-2ad0e354f571 · inbound

BrainFIBRE: A Foundation Model via Information Decomposition for Brain Microstructure cites this paper.

BrainFIBRE: A Foundation Model via Information Decomposition for Brain Microstructure FastMoE: A Fast Mixture-of-Expert Training System

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T09:20:43.421695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:20:43.421695Z digest=sha256:f7916993dfd313f4d1dfcd9405719149416bef57e9eb5bd1b91659d66cd0d181

Observation 6152210c-84fd-4605-a866-bb8536a07119 · inbound

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving cites this paper.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving FastMoE: A Fast Mixture-of-Expert Training System

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:b510d6cca5a231c1fb9c2fa6e5bdfd227f6cb6487cab2e7785ca712770c29c6f

Observation 0dd51e7f-3b5e-46f6-a27b-d4845e2ef976 · inbound

On the Design of Mixture-of-Experts for Dynamic Gaussian Splatting cites this paper.

On the Design of Mixture-of-Experts for Dynamic Gaussian Splatting FastMoE: A Fast Mixture-of-Expert Training System

Reference 63

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T10:47:02.153873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T10:37:15.252017Z digest=sha256:7b439c9b7dba49a06564f16c32ab7d7f1b245d3a3d7a809bf1c77d6795129714

Observation 454c2b12-edc9-4083-bbb7-c81d2a7e2a93 · inbound

On the Design of Mixture-of-Experts for Dynamic Gaussian Splatting cites this paper.

On the Design of Mixture-of-Experts for Dynamic Gaussian Splatting FastMoE: A Fast Mixture-of-Expert Training System

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-14T15:37:19.276654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:37:19.276654Z digest=sha256:be0b1ffb7503ef82d2c6c5ee8d9d2cc6b204a7e3f0e9d9aab37fb84da06b2526

Observation e515c608-a980-4977-be24-17d305f175e5 · inbound

Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts cites this paper.

Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts FastMoE: A Fast Mixture-of-Expert Training System

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T00:38:32.753301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:38:32.753301Z digest=sha256:a0c3ab17e944ab9f02a836f3d1b64c72cbdbd62c2ed0816f0ebeef841ae253f9