Pith. sign in

REVIEW 7 cited by

DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2201.05596 v2 pith:7E6MSA6G submitted 2022-01-14 cs.LG cs.AIcs.DC

classification cs.LGcs.AIcs.DC
keywords modelsmodelinferencetrainingdensecomparedcostdeepspeed-moe
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As the training of giant dense models hits the boundary on the availability and capability of the hardware resources today, Mixture-of-Experts (MoE) models become one of the most promising model architectures due to their significant training cost reduction compared to a quality-equivalent dense model. Its training cost saving is demonstrated from encoder-decoder models (prior works) to a 5x saving for auto-aggressive language models (this work along with parallel explorations). However, due to the much larger model size and unique architecture, how to provide fast MoE model inference remains challenging and unsolved, limiting its practical usage. To tackle this, we present DeepSpeed-MoE, an end-to-end MoE training and inference solution as part of the DeepSpeed library, including novel MoE architecture designs and model compression techniques that reduce MoE model size by up to 3.7x, and a highly optimized inference system that provides 7.3x better latency and cost compared to existing MoE inference solutions. DeepSpeed-MoE offers an unprecedented scale and efficiency to serve massive MoE models with up to 4.5x faster and 9x cheaper inference compared to quality-equivalent dense models. We hope our innovations and systems help open a promising path to new directions in the large model landscape, a shift from dense to sparse MoE models, where training and deploying higher-quality models with fewer resources becomes more widely possible.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 55 citations worldwide. Full citation record

  1. Communication-Aware Placement and Pruning for Efficient Mixture-of-Experts Inference

    cs.DC 2026-07 conditional novelty 6.0 of 10

    Communication-aware expert placement plus device-level pruning yields 1.23–1.86× MoE inference throughput and better accuracy at equal speedup than load-balance or sequential baselines.

  2. Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference

    cs.DC 2025-08 conditional novelty 6.0 of 10

    HeteroScale coordinates scaling of prefill and decode pools using decode TPS as a single robust signal, reporting a 26.6 percentage point GPU utilization gain in production.

  3. Lilith: Developmental Modular LLMs with Chemical Signaling

    q-bio.NC 2025-07 reject novelty 6.0 of 10

    A conceptual framework in which untrained modular LLMs are developed through simulated life and token-based chemical signaling, with the goal of enabling empirical study of consciousness emergence via Integrated Infor...

  4. THOR-MoE: Hierarchical Task-Guided and Context-Responsive Routing for Neural Machine Translation

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A hierarchical routing method that combines predicted task labels with context-aware token routing improves BLEU and reduces activated experts in translation MoE models.

  5. MoE-Beyond: Learning-Based Expert Activation Prediction on Edge Devices

    cs.LG 2025-08 unverdicted novelty 4.0 of 10

    An abstract-only MoE paper claiming 97.5% activation prediction accuracy and a 17% to 72% cache hit rate gain, whose full text is a different paper on functional equations, making the results unverifiable.

  6. ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing

    cs.MM 2025-06 reject novelty 4.0 of 10

    ViFusion combines dynamic tensor fusion with hierarchical AllReduce to speed up distributed video feature indexing, but the 8-22x throughput claim is an overstatement of bandwidth gains over a self-defined baseline.

  7. Hecto: Modular Sparse Experts for Adaptive and Interpretable Reasoning

    cs.AI 2025-06 reject novelty 2.0 of 10

    A lightweight heterogeneous MoE with a GRU and an FFNN expert trails homogeneous baselines, and its claimed reasoning-type specialization is confounded by unequal expert inputs.

Pith tools