Pith. sign in

REVIEW 5 cited by

Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.12370 v3 pith:X3GJ3PWS submitted 2025-01-21 cs.LG cs.AI

classification cs.LGcs.AI
keywords scalingparameterssparsitycapacitymodelperformancecomputeexample
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Scaling the capacity of language models has consistently proven to be a reliable approach for improving performance and unlocking new capabilities. Capacity can be primarily defined by two dimensions: the number of model parameters and the compute per example. While scaling typically involves increasing both, the precise interplay between these factors and their combined contribution to overall capacity remains not fully understood. We explore this relationship in the context of sparse Mixture-of-Experts (MoEs), which allow scaling the number of parameters without proportionally increasing the FLOPs per example. We investigate how varying the sparsity level, i.e., the fraction of inactive parameters, impacts model's performance during pretraining and downstream few-shot evaluation. We find that under different constraints (e.g., parameter size and total training compute), there is an optimal level of sparsity that improves both training efficiency and model performance. These results provide a better understanding of the impact of sparsity in scaling laws for MoEs and complement existing works in this area, offering insights for designing more efficient architectures.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When Does Sparsity Mitigate the Curse of Depth in LLMs

    cs.CL 2026-03 conditional novelty 6.0 of 10

    Implicit and explicit sparsity reduce residual-stream variance and improve layer effectiveness metrics, enabling a depth-scaling recipe with about 4.6 points higher downstream accuracy.

  2. Predictable Scale: Part II, Farseer: A Refined Scaling Law in Large Language Models

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A newly fitted scaling law with model-size-dependent data exponents predicts LLM loss more accurately than Chinchilla, including at a held-out 25.1B model.

  3. Towards Foundational Models for Dynamical System Reconstruction: Hierarchical Meta-Learning via Mixture of Experts

    cs.LG 2025-02 conditional novelty 6.0 of 10

    MixER uses K-means clustering and least squares to route each dynamical system to a specialist expert, improving reconstruction across loosely related ODE families in low-data regimes while underperforming on closely ...

  4. Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient

    cs.LG 2025-02 conditional novelty 6.0 of 10

    A joint scaling law fitted to over 280 models shows that, under fixed memory or total-parameter budgets, MoE models can achieve lower loss than dense models when trained on more tokens.

  5. Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Soup-of-Experts pretrains a shared parameter bank and many expert vectors, plus a router, so a small specialist language model can be instantiated instantly from any domain-weight mixture without retraining.

Pith tools