REVIEW 6 cited by
A Review of Sparse Expert Models in Deep Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Sparse expert models are a thirty-year old concept re-emerging as a popular architecture in deep learning. This class of architecture encompasses Mixture-of-Experts, Switch Transformers, Routing Networks, BASE layers, and others, all with the unifying idea that each example is acted on by a subset of the parameters. By doing so, the degree of sparsity decouples the parameter count from the compute per example allowing for extremely large, but efficient models. The resulting models have demonstrated significant improvements across diverse domains such as natural language processing, computer vision, and speech recognition. We review the concept of sparse expert models, provide a basic description of the common algorithms, contextualize the advances in the deep learning era, and conclude by highlighting areas for future work.
Forward citations
Cited by 6 Pith papers
-
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging
Soup-of-Experts pretrains a shared parameter bank and many expert vectors, plus a router, so a small specialist language model can be instantiated instantly from any domain-weight mixture without retraining.
-
EmoStyle: Affective Conditioning of Style-Specialist Experts for Emotional Image Generation
EmoStyle injects LLM-inferred valence-arousal and emotion labels into Z-Image via AdaLN-style residual modulation over style-bucket LoRA experts, plus VLM candidate ranking, and ranked first on AffectiveArt Track 1.
-
Mixture of LoRA Experts with Multi-Modal and Multi-Granularity LLM Generative Error Correction for Accented Speech Recognition
A three-stage pipeline of mono-accent LoRA experts, hierarchical routing, and phoneme-plus-word LLM error correction cuts accented-English WER from 6.34% to 2.07% on a combined 9-accent test set.
-
Heterogeneous Ranking in Industrial-Scale Recommender Systems: A Case Study
Heterogeneity-conditioned gating and expert modulation improved multi-task ranking in Google Discover, narrowing the articles-vs-videos ranking gap and lifting feed engagement in A/B tests.
-
(GG) MoE vs. MLP on Tabular Data
A Gumbel-Softmax-gated mixture of experts with numerical embeddings matches MLP accuracy on 38 tabular datasets while using roughly 10x fewer parameters, but its edge over MLP is not statistically significant.
-
Position: AI Scaling: From Up to Down and Out
AI scaling is reframed as three paradigms: Scaling Up, Scaling Down, and Scaling Out, with future gains predicted to come from down and out.
Discussion (0). Continue with ORCID to comment.