Pith. sign in

REVIEW 5 cited by

Mega: Moving Average Equipped Gated Attention

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.10655 v3 pith:LFH4ZRFL submitted 2022-09-21 cs.LG

classification cs.LG
keywords attentionmegaincludingmechanismmodelingsequenceaveragebias
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The design choices in the Transformer attention mechanism, including weak inductive bias and quadratic computational complexity, have limited its application for modeling long sequences. In this paper, we introduce Mega, a simple, theoretically grounded, single-head gated attention mechanism equipped with (exponential) moving average to incorporate inductive bias of position-aware local dependencies into the position-agnostic attention mechanism. We further propose a variant of Mega that offers linear time and space complexity yet yields only minimal quality loss, by efficiently splitting the whole sequence into multiple chunks with fixed length. Extensive experiments on a wide range of sequence modeling benchmarks, including the Long Range Arena, neural machine translation, auto-regressive language modeling, and image and speech classification, show that Mega achieves significant improvements over other sequence models, including variants of Transformers and recent state space models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 36 citations worldwide. Full citation record

  1. Partial Ring Scan: Revisiting Scan Order in Vision State Space Models

    cs.CV 2026-02 conditional novelty 6.0 of 10

    Ring-based scanning with selective channel routing improves accuracy, speed, and rotation robustness of vision state-space models.

  2. Cortical-SSM: A Deep State Space Model for Motor Imagery Decoding from EEG Signals

    cs.CV 2025-10 conditional novelty 6.0 of 10

    Cortical-SSM, a dual state-space architecture with wavelet-based frequency features, reports state-of-the-art motor-imagery decoding accuracy on OpenBMI, Stieger2021, and a clinical ECoG-ALS dataset.

  3. Elucidating the Design Space of Decay in Linear Attention

    cs.CL 2025-09 conditional novelty 5.0 of 10

    A controlled study of decay in linear attention finds median decay near 0.8 works best, vector decay generally beats scalar decay, and RoPE/TPE give little benefit for models with sub-unity decay.

  4. Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA)

    cs.LG 2025-07 reject novelty 5.0 of 10

    WERSA is a linear-complexity attention mechanism combining Haar wavelets with random feature projections, reporting small accuracy gains over baselines but resting on a flawed softmax approximation proof.

  5. Block-based Symmetric Pruning and Fusion for Efficient Vision Transformers

    cs.CV 2025-07 reject novelty 4.0 of 10

    A block-based attention pruning and fusion method for ViTs that reports large accuracy gains at reduced FLOPs, but the gain is mostly from the chunk-attention backbone and the core symmetry claim is false.

Pith tools