Pith. sign in

REVIEW 3 major objections 6 minor 1 cited by

Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective

T0 review · 3 major / 6 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read Sigmoid self-attention needs fewer samples than softmax attention.

desk verdict The sigmoid-gating MoE analysis is a real technical contribution, but the headline claim about self-attention sample complexity does not follow from it. read the letter →

arxiv 2502.00281 v2 pith:WWV7XGQO submitted 2025-02-01 cs.LG cs.AI

classification cs.LGcs.AI MSC 62F1268T07
keywords sigmoidself-attentionsoftmaxsamplecomplexitymixtureofexpertsquadraticgatingconvergenceratesVoronoilossTransformertheory
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to prove that replacing the row-wise softmax in Transformer self-attention with an element-wise sigmoid makes the mechanism more sample-efficient: fewer training examples are needed to reach a given approximation accuracy. It does this by rewriting one row of the attention matrix as a mixture-of-experts model whose gate is quadratic, then analyzing how fast least-squares estimation recovers the expert parameters in that model. In the dense regime, where gating parameters are nonzero, the paper derives a polynomial sample complexity of $O(\epsilon^{-2})$ for sigmoid attention, compared with $O(\epsilon^{-4})$ or exponential rates for softmax attention from the comparison baseline. In the sparse regime, the two mechanisms have the same sample complexity. The paper concludes that sigmoid self-attention is at least as data-efficient as softmax, and strictly better in the regime the authors argue is common in practice.

What carries the argument

The load-bearing object is the representation of one row of the attention matrix as a mixture of experts with quadratic affinity scores: $[\mathrm{SigmoidAttn}(X)]_{i,:} = \sum_j \sigma(x_i B x_j^\top)\, x_j W_V$, with $B = W_Q W_K^\top / \sqrt{d_k}$. This rewrites learning attention weights as estimating the parameters of a sigmoid-gating mixture-of-experts regression. The proofs are carried by Voronoi loss functions, which measure parameter discrepancies cell by cell, together with strong and weak identifiability conditions expressed as linear independence of partial derivatives; those conditions determine whether Taylor-expanded parameter differences can be separated. The sigmoid's element-wise, unnormalized structure removes the softmax normalization constraint and, in the dense regime, lets first-order Taylor terms dominate, yielding the fast $O_P((\log n/n)^{1/2})$ expert rate.

What would settle it

Fit a sigmoid-gated mixture of experts with polynomial experts to synthetic data from the same model in the dense over-specified regime and plot the Voronoi loss against $n$; the paper predicts decay of order $n^{-1/2}$, so observing decay of order $n^{-1/4}$ or slower, the softmax baseline rate, would refute the central claim.

Watch

Extended reading notes

Core claim

The central claim is a concrete sample-complexity separation between sigmoid and softmax attention. Using the equivalence that each row of the attention output is a mixture of experts, with gate $\sigma(x_i B x_j^\top)$ for sigmoid and value rows $x_j W_V$ as experts, the authors analyze the sigmoid-gating mixture-of-experts regression model. They show that in the dense regime, weakly identifiable experts such as ReLU, GELU, and polynomial experts are estimated at rate $O_P((\log n/n)^{1/2})$, so only $O(\epsilon^{-2})$ samples are needed for approximation error $\epsilon$. The comparable softmax analysis from the baseline they compare against yields $O(\epsilon^{-4})$ for strongly identifiable experts and exponential $O(\exp(\epsilon^{-1/\tau}))$ for polynomial experts. The paper therefore claims sigmoid attention is more sample-efficient than softmax attention in the dense regime and equally efficient in the sparse regime.

Load-bearing premise

The analysis assumes the sample complexity of estimating a sigmoid-gated mixture-of-experts regression transfers directly to the sample complexity of learning self-attention parameters in a Transformer, even though in attention the experts are random input tokens shared across rows and the fitted parameters are $W_Q$, $W_K$, and $W_V$.

Editorial extensions

If this is right

  • In the dense regime, sigmoid self-attention reaches the same approximation error as softmax with quadratically fewer samples, $O(\epsilon^{-2})$ versus $O(\epsilon^{-4})$.
  • Polynomial experts, which are exponentially hard under softmax attention, become polynomially easy under sigmoid attention in the dense regime.
  • Under the sparse regime, sigmoid attention is not worse than softmax: both require $O(\epsilon^{-4})$ samples for strongly identifiable experts.
  • The same separation holds under partially quadratic affinity scores, where linear experts also move from exponential to polynomial sample complexity.
  • The result provides a statistical justification for the empirical success of sigmoid attention: removing token competition also removes a statistical bottleneck in the dense regime.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the transfer assumption holds, the paper implies sigmoid attention is preferable in small-sample settings such as few-shot learning or low-resource modeling, where sample efficiency matters more than raw capacity.
  • The mixture-of-experts representation suggests a testable architectural prediction: attention heads whose value projections are well approximated by low-degree polynomials should benefit most from switching to sigmoid gating.
  • A rigorous extension to multi-head attention via hierarchical mixtures of experts, which the authors flag as future work, would likely preserve the dense-regime separation if the hierarchy inherits weak identifiability.
  • Because input-dependent gating weights are typical in trained models, the paper's argument predicts that practical gains from sigmoid attention should be widespread rather than confined to specially constructed cases.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper claims to prove that sigmoid self-attention is more sample-efficient than softmax self-attention. It does so by representing each row of a self-attention matrix as a mixture of experts (MoE) with quadratic affinity scores, then analyzing the sample complexity of estimating the parameters of a sigmoid-gating MoE regression model (Eq. 5). The authors derive regression-function convergence rates under sparse and dense regimes for the gating parameters, translate these into parameter and expert convergence rates via Voronoi losses, and compare them with rates imported from a prior softmax-MoE analysis [1]. In the dense regime they claim a polynomial O(epsilon^-2) sample complexity for sigmoid-gating experts, versus O(epsilon^-4) or exponential rates for softmax, concluding that sigmoid self-attention is more sample-efficient. The paper also presents numerical experiments on the MoE regression models.

Significance. If the central claim were established, this would be a notable theoretical result giving a rigorous statistical justification for the empirically observed advantages of sigmoid self-attention. The technical machinery developed for sigmoid-gating MoE convergence—bracketing-entropy regression bounds, Voronoi-loss lower bounds, weak identifiability conditions, and a minimax lower bound—contains interesting components that could be of independent value. However, the paper does not provide a theorem connecting the MoE regression rates to the estimation of actual self-attention parameters, and the dense-regime analysis is carried out against a proxy target rather than the ground-truth experts. As a result, the headline claim about sigmoid versus softmax self-attention is not supported by the results as they stand.

major comments (3)
  1. [Section 2 and Section 4] No formal result connects the sample complexity of the MoE regression model in Eq. (5) to the sample complexity of learning self-attention parameters. In Eq. (5) the unknown quantities are fixed expert and gating parameters estimated from i.i.d. pairs (X_i, Y_i), whereas in the self-attention representation of Section 2 the 'experts' are the input token projections x_j W_V, which are random and shared across rows, and the learnable parameters are W_Q, W_K, W_V. The algebraic rewriting of one attention row as an MoE does not map one estimation problem onto the other, and no theorem in the paper states that the rates for the regression model transfer to attention parameter estimation. Therefore the abstract claim that 'sigmoid self-attention has lower sample complexity than softmax self-attention' does not follow from the MoE analysis.
  2. [Section 4.2, Theorem 3, Corollary 1, and Proposition 3] The dense-regime rate is for an over-parameterized proxy, not for the true experts. Corollary 1 states that inf_{G in M_N(Theta)\M_{N*}(Theta)} ||f_{\hat G_n} - f_G||_{L2(mu)} = O_P(sqrt(log n/n)), and Theorem 3 bounds L3(\hat G_n, \bar G) where \bar G is the minimizer of ||f_G - f_{G*}|| over that same excluded set. The text then concludes that 'it takes those experts only a polynomial number of samples O(epsilon^-2) to achieve an approximation error of epsilon'. This conclusion is not established because the error is measured against \bar G, not against the ground-truth expert parameters G*. In fact, Proposition 3 in Appendix B.5 shows that a dense over-specified sum of sigmoid gates cannot converge to the true single-sigmoid gating function, so there is no reason to expect \bar eta_i = eta*_i or \bar f_G = f_{G*}. The softmax rates imported from [1] are for estimation of the ground-truth experts, so the comparison in Table 1 compares distances to two different targets and does not support the claimed sample-efficiency advantage.
  3. [Section 4.1.2, Theorem 2] The exponential sample-complexity claim for polynomial experts is not supported by Theorem 2. That theorem gives inf_{\hat G_n} sup_{G} E[L_{2,r}(\hat G_n, G)] \gtrsim n^{-1/2} for every r \ge 1. This is a minimax lower bound on the r-th-power Voronoi loss, and it implies at most that the corresponding parameter discrepancies cannot be estimated faster than a polynomial rate of order n^{-1/(2r)} for fixed r. The subsequent text claims that the parameter convergence rates are 'slower than any polynomial rates O_P(n^{-1/2r}) for any r \ge 1, potentially as slow as O_P(1/log^tau(n))'. This is internally inconsistent because n^{-1/2r} is itself a polynomial rate, and the 'potentially as slow as n^{-1/log^tau(n)}' assertion is not a consequence of any proved statement. Consequently, the 'exponential number of data O(exp(epsilon^{-1/tau}))' for sigmoid-gating polynomial experts in the sparse regime, as listed in Table 1, is not proven, and the comparison with the softmax exponential rate from [1] is not established.
minor comments (6)
  1. [Section 5, Setup] The phrase 'ynthetic data' should be 'synthetic data'.
  2. [Section 6, Conclusion] The sentence 'Our results show that sigmoid self-attention has a higher sample complexity than the softmax version in the more common dense regime' contradicts the abstract, the title, and Table 1; it should read 'lower sample complexity'.
  3. [Appendix A, Related Works] The sentence 'Le et al. Furthermore, Akbarian et al. [1]...' contains an incomplete citation 'Le et al.' with no reference or title; please complete it or remove it.
  4. [Section 4.1.1, Theorem 1] The statement 'If the expert function ... then the lower bound ... holds true ... then L1( bGn, G∗) = OP(...)' has a double-'then' construction that should be rephrased for clarity.
  5. [Appendix B.1, Step 4] The covering numbers |\Delta_\tau| and |\Omega_\tau| are deterministic quantities but are written with OP(...); they should use O(...) notation.
  6. [Section 5 and Figure 1] The Voronoi loss L3(\hat G_n, G) is plotted for the sigmoid model fitted to data generated from a softmax-gating MoE, but the target measure G for the sigmoid fit is never defined; please specify how G is chosen in that setting.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: the sigmoid MoE rates are independently derived, though the dense-regime sample-complexity claim targets a proxy rather than the true experts.

full rationale

No circular step is established. The sigmoid-gating MoE rates in Sections 3 and 4 are derived from a least-squares estimator defined in equation (6) via bracketing-entropy bounds (Appendix B.1) and Voronoi-loss lower bounds (Theorems 1-3), so they do not presuppose the softmax rates or the attention claim. The Section 2 attention-to-MoE identity is an algebraic rewriting of one attention row, not a fit. The softmax baseline is imported from reference [1], whose authors overlap with the present paper; this is load-bearing self-citation, but [1] is a separate stated-assumption analysis of softmax quadratic gating, so it does not make the sigmoid derivation circular under the hard rules. The main concern is a target mismatch, not circularity: in the dense regime, Theorem 3 gives inf_G L3(hat G_n, G) = O_P(sqrt(log n/n)) for G defined as the argmin over over-specified mixtures, and Section 4.2 then concludes that 'it takes those experts only a polynomial number of samples O(epsilon^-2) to achieve an approximation error of epsilon' without showing that the proxy parameters equal the true experts or that the proxy regression function equals f_{G*}; this is a transfer gap that affects the correctness of the comparison, but it is not an equation-level reduction of the prediction to its inputs. The Section 6 single-head limitation is a scope statement, not a circular step. The score is set to 2 to reflect the low-level self-citation burden while noting that no circular derivation was found.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The MoE analysis relies on standard statistical machinery and identifiability assumptions. The main unsupported input is the transfer of MoE sample complexity to self-attention, plus the assertion that the dense regime is the practically relevant one.

assumptions (5)
  • domain assumption The data follows the regression model Yi = f_G*(Xi) + eps_i with i.i.d. Gaussian noise and bounded input domain (Eq. 4).
    Section 3, Eq. (4). This setup is standard for MoE convergence analysis but is not checked against actual self-attention data distributions.
  • domain assumption Strong or weak identifiability conditions (Definitions 1-4) hold for the expert functions, including ReLU/GELU networks and polynomial experts.
    Section 4.1.1 and 4.2. These conditions are asserted with examples but not derived from attention mechanisms; they are tailored to make the Taylor expansion arguments work.
  • standard math Standard empirical process theory results (bracketing entropy, Le Cam's lemma, Fatou's lemma) apply as used in the proofs.
    Appendix B uses these tools to derive the OP(sqrt(log n/n)) regression rates. These are standard and likely valid.
  • ad hoc to paper The self-attention row representation as an MoE with quadratic affinity scores (Section 2) preserves the statistical estimation problem of attention.
    This is the key transfer assumption: the paper rewrites one row of attention as a weighted sum of value vectors, then treats it as an MoE regression model with fixed expert parameters. No proof is given that the sample complexity of estimating W_Q, W_K, W_V equals that of estimating the MoE parameters.
  • ad hoc to paper The dense regime is more common in practice than the sparse regime.
    Section 4.2 states this without empirical evidence and uses it to present the dense-regime improvement as the main practical conclusion.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective." pith.science (2026). https://pith.science/paper/WWV7XGQO

@misc{pith2026250200281,
  author       = {Pith},
  title        = {Pith review of: Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WWV7XGQO}},
  note         = {Machine review of arXiv:2502.00281}
}
read the original abstract

At the core of the popular Transformer architecture is the self-attention mechanism, which dynamically assigns softmax weights to each input token so that the model can focus on the most salient information. However, the softmax structure slows down the attention computation due to its row-wise nature, and it inherently introduces competition among tokens: as the weight assigned to one token increases, the weights of others decrease. This competitive dynamic may narrow the focus of self-attention to a limited set of features, potentially overlooking other informative characteristics. Recent experimental studies have shown that using the element-wise sigmoid function helps eliminate token competition and reduce the computational overhead. Despite these promising empirical results, a rigorous comparison between sigmoid and softmax self-attention mechanisms remains absent in the literature. This paper closes this gap by theoretically demonstrating that sigmoid self-attention is more sample-efficient than its softmax counterpart. Toward that goal, we represent the self-attention matrix as a mixture of experts and show that ``experts'' in sigmoid self-attention require significantly less data to achieve the same approximation error as those in softmax self-attention.

Figures

Figures reproduced from arXiv: 2502.00281 by the authors.

Figure 1
Figure 1. Log-log plots of the convergence rates of Voronoi losses for softmax and sigmoid quadratic [PITH_FULL_IMAGE:figures/full_fig_p011_1.png] view at source ↗
Figure 2
Figure 2. Log-log plots of empirical convergence rates of Voronoi losses for softmax and sigmoid [PITH_FULL_IMAGE:figures/full_fig_p044_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SAGE: Shape-Adapting Gated Experts for Adaptive Histopathology Image Segmentation

    eess.IV 2025-11 conditional novelty 5.0 of 10

    A CNN-Transformer U-Net with hierarchical expert routing and a shape-adapting hub reports state-of-the-art colon histopathology segmentation Dice of 95.57% on EBHI, 95.16% on DigestPath, and 94.17% on GlaS.

Reference graph

Works this paper leans on

51 extracted references · 36 canonical work pages · cited by 1 Pith paper

  1. [1]

    Akbarian, H

    P. Akbarian, H. Nguyen, X. Han, and N. Ho. Quadratic gating functions in mixture of experts: A statistical insight.arXiv preprint arXiv:2410.11222, 2024. (Cited on pages 2, 5, 9, 10, 12, and 33.)

  2. [2]

    Brown, B

    T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al. Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020. (Cited on page 1.)

  3. [3]

    Carion, F

    N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko. End-to-end object detection with transformers. InEuropean conference on computer vision, pages 213–229. Springer, 2020. (Cited on page 1.)

  4. [4]

    L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch. Decision transformer: Reinforcement learning via sequence modeling.Advances in neural information processing systems, 34:15084–15097, 2021. (Cited on page 1.)

  5. [5]

    Z. Chen, Y. Deng, Y. Wu, Q. Gu, and Y. Li. Towards understanding the mixture-of-experts layer in deep learning. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems, volume 35, pages 23049–23062. Curran Associates, Inc., 2022.(Cited on page 3.)

  6. [6]

    Z. Chen, Y. Shen, M. Ding, Z. Chen, H. Zhao, E. G. Learned-Miller, and C. Gan. Mod-squad: Designing mixtures of experts as modular multi-task learners. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11828–11837, 2023.(Cited on page 3.) 45

  7. [7]

    Z. Chi, L. Dong, S. Huang, D. Dai, S. Ma, B. Patra, S. Singhal, P. Bajaj, X. Song, X.-L. Mao, H. Huang, and F. Wei. On the representation collapse of sparse mixture of experts. In A. H. Oh, A. Agarwal, D. Belgrave, and K. Cho, editors,Advances in Neural Information Processing Systems, 2022. (Cited on page 3.)

  8. [8]

    Csordás, K

    R. Csordás, K. Irie, and J. Schmidhuber. Approximating two-layer feedforward networks for efficient transformers. In H. Bouamor, J. Pino, and K. Bali, editors,Findings of the Association for Computational Linguistics: EMNLP 2023, pages 674–692, Singapore, Dec. 2023. Association for Computational Linguistics. (Cited on page 3.)

Show all 51 references
  1. [9]

    Csordás, P

    R. Csordás, P. Pi´kekos, K. Irie, and J. Schmidhuber. Switchhead: Accelerating transformers with mixture-of-experts attention. arXiv preprint arXiv:2312.07987, 2023. (Cited on pages 2 and 12.)

  2. [10]

    T. Dao. Flashattention-2: Faster attention with better parallelism and work partitioning. In The Twelfth International Conference on Learning Representations, 2024. (Cited on page 2.)

  3. [11]

    T. Dao, D. Y. Fu, S. Ermon, A. Rudra, and C. Re. Flashattention: Fast and memory-efficient exact attention with IO-awareness. InAdvances in Neural Information Processing Systems,

  4. [12]

    DeepSeek-AI, A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, D. Dai, D. Guo, D. Yang, D. Chen, D. Ji, E. Li, F. Lin, F. Dai, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Bao, H. Xu, H. Wang, H. Zhang, H. Ding, H. Xin, H. Gao, H. Li, H. Qu, ...

  5. [13]

    Devlin, M.-W

    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. InNorth American Chapter of the Association for Computational Linguistics, 2019. (Cited on page 1.)

  6. [14]

    Dosovitskiy, L

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby. An image is worth 16x16 46 words: Transformers for image recognition at scale.ArXiv, abs/2010.11929, 2020. (Cited on page 1.)

  7. [15]

    N. Du, Y. Huang, A. M. Dai, S. Tong, D. Lepikhin, Y. Xu, M. Krikun, Y. Zhou, A. W. Yu, O. Firat, B. Zoph, L. Fedus, M. Bosma, Z. Zhou, T. Wang, Y. E. Wang, K. Webster, M. Pellat, K. Robinson, K. S. Meier-Hellstern, T. Duke, L. Dixon, K. Zhang, Q. V. Le, Y. Wu, Z. Chen, and C. ...

  8. [16]

    Faria and G

    S. Faria and G. Soromenho. Fitting mixtures of linear regressions. Journal of Statistical Computation and Simulation, 80(2):201–225, 2010. (Cited on page 3.)

  9. [17]

    Fedus, B

    W. Fedus, B. Zoph, and N. Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research, 23:1–39, 2022. (Cited on page 3.)

  10. [18]

    N. Gaur, B. Farris, P. Haghani, I. Leal, P. J. Moreno, M. Prasad, B. Ramabhadran, and Y. Zhu. Mixture of informed experts for multilingual speech recognition. InICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 6234–6238....

  11. [19]

    X. Gu, T. Pang, C. Du, Q. Liu, F. Zhang, C. Du, Y. Wang, and M. Lin. When attention sink emerges in language models: An empirical view.arXiv preprint arXiv:2410.10781, 2024. (Cited on page 1.)

  12. [20]

    Hazimeh, Z

    H. Hazimeh, Z. Zhao, A. Chowdhery, M. Sathiamoorthy, Y. Chen, R. Mazumder, L. Hong, and E. Chi. Dselect-k: Differentiable selection in the mixture of experts with applications to multi-task learning. Advances in Neural Information Processing Systems, 34:29335–29347, 2021. (Cit...

  13. [21]

    Ho, C.-Y

    N. Ho, C.-Y. Yang, and M. I. Jordan. Convergence rates for Gaussian mixtures of experts. Journal of Machine Learning Research, 23(323):1–81, 2022. (Cited on page 13.)

  14. [22]

    E. S. Hu, K. Ahn, Q. Liu, H. Xu, M. Tomar, A. Langford, D. Jayaraman, A. Lamb, and J. Langford. Learning to achieve goals with belief state transformers.ArXiv, abs/2410.23506,

  15. [23]

    R. A. Jacobs, M. I. Jordan, S. J. Nowlan, and G. E. Hinton. Adaptive mixtures of local experts. Neural Computation, 3, 1991. (Cited on pages 2 and 3.)

  16. [24]

    A. Jain, V. P. Singh, and S. P. Rath. A multi-accent acoustic model using mixture of experts for speech recognition. InInterspeech, pages 779–783, 2019.(Cited on page 3.)

  17. [25]

    A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. de las Casas, E. B. Hanna, F. Bressand, G. Lengyel, G. Bour, G. Lample, L. R. Lavaud, L. Saulnier, M.-A. Lachaux, P. Stock, S. Subramanian, S. Yang, S. Antoniak, T. L. Scao, T. Gervet, T...

  18. [26]

    P. Jin, B. Zhu, L. Yuan, and S. Yan. Moh: Multi-head attention as mixture-of-head attention. arXiv preprint arXiv:2410.11842, 2024. (Cited on page 12.)

  19. [27]

    M. I. Jordan and R. A. Jacobs. Hierarchical mixtures of experts and the EM algorithm.Neural Computation, 6:181–214, 1994. (Cited on pages 3 and 12.)

  20. [28]

    C. Kim, J. Park, J. Shin, H. Lee, P. Abbeel, and K. Lee. Preference transformer: Modeling human preferences using transformers for rl.arXiv preprint arXiv:2303.00957, 2023. (Cited on page 1.)

  21. [29]

    Kwon and S.-W

    Y. Kwon and S.-W. Chung. Mole: Mixture of language experts for multi-lingual automatic speech recognition. InICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5. IEEE, 2023.(Cited on page 3.)

  22. [30]

    B. Lindsay. Mixture models: Theory, geometry and applications. In NSF-CBMS Regional Conference Series in Probability and Statistics. IMS, Hayward, CA., 1995.(Cited on page 3.)

  23. [31]

    Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF international conference on computer vision, pages 10012–10022, 2021.(Cited on page 1.)

  24. [32]

    J. Ma, Z. Zhao, X. Yi, J. Chen, L. Hong, and E. H. Chi. Modeling task relationships in multi- task learning with multi-gate mixture-of-experts. InProceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining, pages 1930–1939, 2018.(Cited on page 3.)

  25. [33]

    Manole and N

    T. Manole and N. Ho. Refined convergence rates for maximum likelihood estimation under finite mixture models. InProceedings of the 39th International Conference on Machine Learning, volume 162 ofProceedings of Machine Learning Research, pages 14979–15006. PMLR, 17–23 Jul 2022....

  26. [34]

    E. F. Mendes and W. Jiang. Convergence rates for mixture-of-experts.arXiv preprint arxiv 1110.2058, 2011. (Cited on page 13.)

  27. [35]

    Muennighoff, L

    N. Muennighoff, L. Soldaini, D. Groeneveld, K. Lo, J. D. Morrison, S. Min, W. Shi, P. Walsh, O. Tafjord, N. Lambert, Y. Gu, S. Arora, A. Bhagia, D. Schwenk, D. Wadden, A. Wettig, B. Hui, T. Dettmers, D. Kiela, A. Farhadi, N. A. Smith, P. W. Koh, A. Singh, and H. Hajishirzi. Ol...

  28. [36]

    Nguyen, P

    H. Nguyen, P. Akbarian, F. Yan, and N. Ho. Statistical perspective of top-k sparse softmax gating mixture of experts. InInternational Conference on Learning Representations, 2024. (Cited on page 13.)

  29. [37]

    Nguyen, X

    H. Nguyen, X. Han, C. W. Harris, S. Saria, and N. Ho. On expert estimation in hierarchical mixture of experts: Beyond softmax gating functions.arXiv preprint arXiv:2410.02935, 2024. (Cited on page 12.)

  30. [38]

    Nguyen, N

    H. Nguyen, N. Ho, and A. Rinaldo. On least square estimation in softmax gating mixture of experts. In Proceedings of the ICML, 2024. (Cited on pages 6 and 12.) 48

  31. [39]

    Nguyen, T

    H. Nguyen, T. Nguyen, and N. Ho. Demystifying softmax gating function in Gaussian mixture of experts. InAdvances in Neural Information Processing Systems, 2023. (Cited on page 13.)

  32. [40]

    Puigcerver, C

    J. Puigcerver, C. Riquelme, B. Mustafa, and N. Houlsby. From sparse to soft mixtures of experts. In The Twelfth International Conference on Learning Representations, 2024. (Cited on page 3.)

  33. [41]

    Radford, J

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021.(Cited on page 1.)

  34. [42]

    Raffel, N

    C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67, 2020. (Cited on page 1.)

  35. [43]

    Ramapuram, F

    J. Ramapuram, F. Danieli, E. Dhekane, F. Weers, D. Busbridge, P. Ablin, T. Likhomanenko, J. Digani, Z. Gu, A. Shidani, and R. Webb. Theory, analysis, and best practices for sigmoid self-attention. arXiv preprint arXiv:2409.04431, 2024. (Cited on pages 1, 2, and 12.)

  36. [44]

    Shazeer, A

    N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean. Outra- geously large neural networks: The sparsely-gated mixture-of-experts layer. InIn International Conference on Learning Representations, 2017. (Cited on page 3.)

  37. [45]

    Touvron, T

    H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al. Llama: Open and efficient foundation language models.arXiv preprint arXiv:2302.13971, 2023. (Cited on page 1.)

  38. [46]

    van de Geer.Empirical Processes in M-estimation

    S. van de Geer.Empirical Processes in M-estimation. Cambridge University Press, 2000.(Cited on pages 5 and 13.)

  39. [47]

    Vaswani, N

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin. Attention is all you need. InAdvances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017.(Cited on pages 1, 3, and 12.)

  40. [48]

    X. Wu, S. Huang, W. Wang, and F. Wei. Multi-head mixture-of-experts.arXiv preprint arXiv:2404.15045, 2024. (Cited on pages 2 and 12.)

  41. [49]

    F. Yan, H. Nguyen, D. Le, P. Akbarian, and N. Ho. Understanding expert structures on minimax parameter estimation in contaminated mixture of experts. InProceedings of The 28th International Conference on Artificial Intelligence and Statistics, 2025. (Cited on page 13.)

  42. [50]

    Z. You, S. Feng, D. Su, and D. Yu. Speechmoe: Scaling to large acoustic models with dynamic routing mixture of experts.arXiv preprint arXiv:2105.03036, 2021. (Cited on page 3.)

  43. [51]

    Y. Zhou, N. Du, Y. Huang, D. Peng, C. Lan, D. Huang, S. Shakeri, D. So, A. Dai, Y. Lu, Z. Chen, Q. Le, C. Cui, J. Laudon, and J. Dean. Brainformers: Trading simplicity for efficiency. In International Conference on Machine Learning, pages 42531–42542. PMLR, 2023.(Cited on page 3.) 49

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.