Pith. sign in

REVIEW 2 cited by

Towards Understanding Mixture of Experts in Deep Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.02813 v1 pith:LTIA47SD submitted 2022-08-04 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords layerlearningproblemdeepexpertsmodelunderstandingclassification
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The Mixture-of-Experts (MoE) layer, a sparsely-activated model controlled by a router, has achieved great success in deep learning. However, the understanding of such architecture remains elusive. In this paper, we formally study how the MoE layer improves the performance of neural network learning and why the mixture model will not collapse into a single model. Our empirical results suggest that the cluster structure of the underlying problem and the non-linearity of the expert are pivotal to the success of MoE. To further understand this, we consider a challenging classification problem with intrinsic cluster structures, which is hard to learn using a single expert. Yet with the MoE layer, by choosing the experts as two-layer nonlinear convolutional neural networks (CNNs), we show that the problem can be learned successfully. Furthermore, our theory shows that the router can learn the cluster-center features, which helps divide the input complex problem into simpler linear classification sub-problems that individual experts can conquer. To our knowledge, this is the first result towards formally understanding the mechanism of the MoE layer for deep learning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Atmos-Bench: 3D Atmospheric Structures for Climate Insight

    cs.CV 2025-07 reject novelty 5.0 of 10

    Atmos-Bench introduces a synthetic 3D benchmark for satellite LiDAR backscatter recovery, and FourCastX, a frequency-MoE inpainting model, reports substantially higher PSNR/SSIM than six baselines on it.

  2. NeuroMoE: A Transformer-Based Mixture-of-Experts Framework for Multi-Modal Neurological Disorder Classification

    eess.IV 2025-06 reject novelty 4.0 of 10

    NeuroMoE fuses anatomical, diffusion, and functional MRI with clinical and serum biomarkers via a gated mixture-of-experts, reporting 82.47% accuracy for PD/iRBD/HC classification on a proprietary 113-subject dataset.

Pith tools