Pith. sign in

REVIEW 2 cited by

SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.07987 v3 pith:7RDKHMXP submitted 2023-12-13 cs.LG cs.CLcs.NE

classification cs.LGcs.CLcs.NE
keywords switchheadcomputeattentionbaselineperformancetransformerachievingfeedforward
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite many recent works on Mixture of Experts (MoEs) for resource-efficient Transformer language models, existing methods mostly focus on MoEs for feedforward layers. Previous attempts at extending MoE to the self-attention layer fail to match the performance of the parameter-matched baseline. Our novel SwitchHead is an effective MoE method for the attention layer that successfully reduces both the compute and memory requirements, achieving wall-clock speedup, while matching the language modeling performance of the baseline Transformer. Our novel MoE mechanism allows SwitchHead to compute up to 8 times fewer attention matrices than the standard Transformer. SwitchHead can also be combined with MoE feedforward layers, resulting in fully-MoE "SwitchAll" Transformers. For our 262M parameter model trained on C4, SwitchHead matches the perplexity of standard models with only 44% compute and 27% memory usage. Zero-shot experiments on downstream tasks confirm the performance of SwitchHead, e.g., achieving more than 3.5% absolute improvements on BliMP compared to the baseline with an equal compute resource.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A shared per-token controller jointly routes attention resolution, FFN experts, and KV bit-width and is claimed to Pareto-dominate independently tuned MoD+MoE+KV-quant at matched cost while protecting rare-token accuracy.

  2. Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Routing Mamba applies mixture-of-experts to Mamba projection layers with one shared router, reporting perplexity parity with dense Mamba at roughly half the active parameters on 20B-token pretraining.

Pith tools