Pith. sign in

REVIEW 5 cited by

One-layer transformers fail to solve the induction heads task

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.14332 v1 pith:SYSI26BT submitted 2024-08-26 cs.LG stat.ML

classification cs.LGstat.ML
keywords headsinductionone-layersizesolvetasktransformerargument
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A simple communication complexity argument proves that no one-layer transformer can solve the induction heads task unless its size is exponentially larger than the size sufficient for a two-layer transformer.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Attention-based representations for multi-task computation

    cs.LG 2026-08 accept novelty 7.0 of 10

    For min/max readout, two attention heads beat one head by an exponential resource gap, and for n-bit parity and symmetric Boolean functions, heads times polynomial degree must reach the threshold degree, with matching...

  2. Indexing: the Beginning and the End

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Causal-complexity bounds show RNNs, SSMs, and masked linear attention need ω(1) layers for right-hand indexing, while a one-layer softmax transformer solves it; when the index is first, a one-layer RNN suffices.

  3. Understanding Input Selectivity in Mamba: Impact on Approximation Power, Memorization, and Associative Recall Capacity

    cs.LG 2025-06 conditional novelty 7.0 of 10

    Mamba's S6 layer can represent Haar wavelets and solve associative recall tasks with explicit size bounds, though its memory still decays exponentially unless input-dependent time steps counteract it.

  4. Eigenvalues as a Metric for Memory Dynamics in Sequence Models

    cs.LG 2025-10 conditional novelty 6.0 of 10

    Eigenvalue spectra of attention and SSM dynamics show consistent signatures of memory retention and selective forgetting that align with task requirements.

  5. Fast attention mechanisms: a tale of parallelism

    cs.LG 2025-09 conditional novelty 6.0 of 10

    ANNA, a hashing-based sub-quadratic attention mechanism, provably preserves standard attention's MPC expressiveness and can simulate low-rank attention, while being simulable by MPC with near-linear machines.

Pith tools