Pith. sign in

REVIEW 6 cited by

Hyena Hierarchy: Towards Larger Convolutional Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.10866 v3 pith:N6L5BWGB submitted 2023-02-21 cs.LG cs.CL

classification cs.LGcs.CL
keywords attentionhyenalengthsequencetransformerslanguagemethodsmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in deep learning have relied heavily on the use of large Transformers due to their ability to learn at scale. However, the core building block of Transformers, the attention operator, exhibits quadratic cost in sequence length, limiting the amount of context accessible. Existing subquadratic methods based on low-rank and sparse approximations need to be combined with dense attention layers to match Transformers, indicating a gap in capability. In this work, we propose Hyena, a subquadratic drop-in replacement for attention constructed by interleaving implicitly parametrized long convolutions and data-controlled gating. In recall and reasoning tasks on sequences of thousands to hundreds of thousands of tokens, Hyena improves accuracy by more than 50 points over operators relying on state-spaces and other implicit and explicit methods, matching attention-based models. We set a new state-of-the-art for dense-attention-free architectures on language modeling in standard datasets (WikiText103 and The Pile), reaching Transformer quality with a 20% reduction in training compute required at sequence length 2K. Hyena operators are twice as fast as highly optimized attention at sequence length 8K, and 100x faster at sequence length 64K.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 75 citations worldwide. Full citation record

  1. RoVE: Rotary Value Embeddings Attention for Relative Position-dependent Value Pathways

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    RoVE rotates value embeddings simultaneously with keys in attention to make values position-dependent, reframing RoPE as attentive convolution and reporting gains on long-context tasks in 124M and 354M GPT-2 models.

  2. Structured Recurrent Mixers for Massively Parallelized Sequence Generation

    cs.CL 2026-05 unverdicted novelty 6.0 of 10

    Structured Recurrent Mixers provide a dual parallel-recurrent representation for sequence models, claiming superior training efficiency, information capacity, and inference throughput over linear complexity alternatives.

  3. Customizing the Inductive Biases of Softmax Attention using Structured Matrices

    cs.LG 2025-09 conditional novelty 6.0 of 10

    Structured-matrix scoring functions, BTT and MLR, let attention escape the low-rank bottleneck and add a distance-dependent compute bias, improving accuracy for fixed compute on regression, language modeling, and forecasting.

  4. CellPainTR: Generalizable Representation Learning for Cross-Dataset Cell Painting Analysis

    cs.CV 2025-09 conditional novelty 6.0 of 10

    CellPainTR applies Hyena-based Transformers with source context tokens to Cell Painting profiles, achieving better batch integration and OOD generalization than ComBat and Harmony.

  5. Mamba for Wireless Communications and Networking: Principles and Opportunities

    cs.NI 2025-08 conditional novelty 4.0 of 10

    The paper argues Mamba can improve efficiency and performance in wireless tasks, backed by two small case studies with mixed results.

  6. Directional Non-Commutative Monoidal Structures for Compositional Embeddings in Machine Learning

    cs.LG 2025-05 conditional novelty 4.0 of 10

    A D-axis composition algebra on (vector, matrix-power) tuples provides associative per-axis operators and an interchange law when axis matrices commute, recovering RoPE, affine embedding composition, and SSM-style rec...

Pith tools