Pith. sign in

Paper Citation Record · LEDGER

MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2211.15841.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2211.15841 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:46:02.060413Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T15:24:49.924239Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 405b044b-1f8b-4e02-ab51-dfb4f02a968f · inbound

Mixtral of Experts cites this paper.

Mixtral of Experts MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:13:53.846793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T04:09:15.921778Z digest=sha256:9c648d60078961e7fb639e0fd662d7e23b0081caee269144f8b6f250205d95b1

Observation 107736f1-d318-43c3-ba93-bbc2de294613 · inbound

Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models cites this paper.

Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-18T02:48:44.988800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T02:48:44.900467Z digest=sha256:3311fd26a0f003f31fa0d7bf1a187444794fd068c6705c2ca601aed842ce9a00

Observation 4e0de21e-2c55-46f3-928f-df9205b76c01 · inbound

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism cites this paper.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T11:46:02.060413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:46:02.060413Z digest=sha256:6d90cea1fa1fb957e3f3c9e8c01ae59879deeaea6cb11c8ee8934ce4a6a0de7f

Observation d227385f-692c-4923-ac5c-c4a5c7d59f4d · inbound

Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient cites this paper.

Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:23.496812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:06:23.496812Z digest=sha256:db426d79d2703a539e81184037848bf1c674717f629df458bbd188bd37680f53

Observation 99384180-b8a9-4ab8-a0b7-1a9605377599 · inbound

Training Sparse Mixture Of Experts Text Embedding Models cites this paper.

Training Sparse Mixture Of Experts Text Embedding Models MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T11:20:23.608575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:20:23.608575Z digest=sha256:457932f43fd43ca6c29f5bb3afe87a94a3b97ed1403371898f471680c6b8a947

Observation 55dafae2-588d-4e18-b4ee-64b9a1a0224f · inbound

Scaling Fine-Grained MoE Beyond 50B Parameters: Empirical Evaluation and Practical Insights cites this paper.

Scaling Fine-Grained MoE Beyond 50B Parameters: Empirical Evaluation and Practical Insights MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:58.768850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:58.768850Z digest=sha256:493e449f46c77474b137954c04dea9b1dfbf6afce0f4a867231809913b22128d

Observation afe6b634-b2a1-4f08-ba14-063b67a85fe5 · inbound

Apple Intelligence Foundation Language Models: Tech Report 2025 cites this paper.

Apple Intelligence Foundation Language Models: Tech Report 2025 MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:26:58.583406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:26:58.583406Z digest=sha256:30918504c0dec69fc36ead924c493ff2b4c54c98a1bb2cbc161fb7964efb964c

Observation 9eb5650d-b3eb-45be-924d-73e7de6ca05d · inbound

Maximum Score Routing For Mixture-of-Experts cites this paper.

Maximum Score Routing For Mixture-of-Experts MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T19:23:15.344590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T19:23:15.344590Z digest=sha256:ba5c8123f9ecf3dac5b4e68a6d8d0e3f9178f9046434c231831774cc02f9f8c6

Observation 479a75b0-ca0a-40cb-aac4-948b76d48d76 · inbound

When Does Sparsity Mitigate the Curse of Depth in LLMs cites this paper.

When Does Sparsity Mitigate the Curse of Depth in LLMs MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T20:29:33.439034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T20:29:33.439034Z digest=sha256:c7e4ab1f81b0a2f495831ceb07c496c154d8fecf7754b04fcf26497ec8cb8b1a

Observation c71d9f5f-e52f-4b4b-924c-2810a0ffde9c · inbound

Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns cites this paper.

Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:36:10.790930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T08:29:17.149710Z digest=sha256:92c568376b0c1c6dc4067a5cc31b5683c3d3ef4738469134fba4cff8b90482c6

Observation 52548489-bca6-4a80-a3ec-91609d656bed · inbound

TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM Training cites this paper.

TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM Training MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:06:20.725388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T01:36:41.804171Z digest=sha256:28a815363bedda7e79d17bdcc15409f2b87c2845bf0daf4ecd4a79d571010846

Observation 4171cd97-1054-4062-be7b-654dd935ea7f · inbound

RaMP: Runtime-Aware Megakernel Polymorphism for Mixture-of-Experts cites this paper.

RaMP: Runtime-Aware Megakernel Polymorphism for Mixture-of-Experts MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 10

Resolution
malformed identifier
arxiv_id, observed 2026-05-11T23:36:31.194142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T16:39:35.128879Z digest=sha256:a8c05226ae06c4067d8b681bb3663ebc3a68578418206435b31fa5850a9c131b

Observation 9575abc1-549e-4c21-b324-4d33f24d60a2 · inbound

Surviving Partial Rank Failures in Wide Expert-Parallel MoE Inference cites this paper.

Surviving Partial Rank Failures in Wide Expert-Parallel MoE Inference MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:11:24.699038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:30:58.729425Z digest=sha256:8ec8eeee678d4464948064a5f6e56b9bff4085fd8d93d7c277d121ab8e4780ef

Observation 1235cce1-ca82-4a06-910f-d8a7456eb1cb · inbound

Safety-Oriented Routing Analysis of Mixtral MoE Under Benign and Harmful Prompts cites this paper.

Safety-Oriented Routing Analysis of Mixtral MoE Under Benign and Harmful Prompts MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:24:49.925986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T15:21:34.127680Z digest=sha256:c67e977f631262b5fc53c7e5a2f4ad45ec79a4f374d26051a86fb52bc1361aca

Observation ba00ad91-0f43-49e1-944c-510d54818d01 · inbound

Communication-Aware Placement and Pruning for Efficient Mixture-of-Experts Inference cites this paper.

Communication-Aware Placement and Pruning for Efficient Mixture-of-Experts Inference MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-11T08:35:22.347459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:35:22.347459Z digest=sha256:809be4c410b61c20bd328e463f9e852de029296666648ada7f2472e33d31bbff

Observation fb0b9cf3-8307-4248-9806-c49d1a1b584e · inbound

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix cites this paper.

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T17:32:37.115161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:32:37.115161Z digest=sha256:6e550ede12429a2a598a6c5ce0d3b4b701be483d7f6ed27f8505e315eda159fe

Observation ce1b3d73-d720-4578-998b-fb5c930107d3 · inbound

Route-Block Membership Selects Packed-AWQ Arithmetic: A Controlled Single-Fixture Mechanism Study cites this paper.

Route-Block Membership Selects Packed-AWQ Arithmetic: A Controlled Single-Fixture Mechanism Study MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T00:15:48.390607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:15:48.390607Z digest=sha256:e5ddf965ee560045693f297329a4eaa7537ea80f56188fda19eb8c74fb2b9265

Observation abfb2628-53e8-4d81-a6a7-fcac483bb81d · inbound

Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts cites this paper.

Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T00:38:32.637398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:38:32.637398Z digest=sha256:ff0bae6378c27cd5cf1fb12b35fbe6daf5e10c511950e0ebc69027ede5ba92a3