REVIEW 5 cited by
One-layer transformers fail to solve the induction heads task
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
A simple communication complexity argument proves that no one-layer transformer can solve the induction heads task unless its size is exponentially larger than the size sufficient for a two-layer transformer.
Forward citations
Cited by 5 Pith papers
-
Attention-based representations for multi-task computation
For min/max readout, two attention heads beat one head by an exponential resource gap, and for n-bit parity and symmetric Boolean functions, heads times polynomial degree must reach the threshold degree, with matching...
-
Indexing: the Beginning and the End
Causal-complexity bounds show RNNs, SSMs, and masked linear attention need ω(1) layers for right-hand indexing, while a one-layer softmax transformer solves it; when the index is first, a one-layer RNN suffices.
-
Understanding Input Selectivity in Mamba: Impact on Approximation Power, Memorization, and Associative Recall Capacity
Mamba's S6 layer can represent Haar wavelets and solve associative recall tasks with explicit size bounds, though its memory still decays exponentially unless input-dependent time steps counteract it.
-
Eigenvalues as a Metric for Memory Dynamics in Sequence Models
Eigenvalue spectra of attention and SSM dynamics show consistent signatures of memory retention and selective forgetting that align with task requirements.
-
Fast attention mechanisms: a tale of parallelism
ANNA, a hashing-based sub-quadratic attention mechanism, provably preserves standard attention's MPC expressiveness and can simulate low-rank attention, while being simulable by MPC with near-linear machines.
Discussion (0). Continue with ORCID to comment.