Pith. sign in

REVIEW 3 cited by

On the Computational Power of Transformers and its Implications in Sequence Modeling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.09286 v3 pith:WAYJNVES submitted 2020-06-16 cs.LG cs.CLstat.ML

classification cs.LGcs.CLstat.ML
keywords transformerspositionalpoweranalyzecomputationalimplicationsmodelingparticular
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Transformers are being used extensively across several sequence modeling tasks. Significant research effort has been devoted to experimentally probe the inner workings of Transformers. However, our conceptual and theoretical understanding of their power and inherent limitations is still nascent. In particular, the roles of various components in Transformers such as positional encodings, attention heads, residual connections, and feedforward networks, are not clear. In this paper, we take a step towards answering these questions. We analyze the computational power as captured by Turing-completeness. We first provide an alternate and simpler proof to show that vanilla Transformers are Turing-complete and then we prove that Transformers with only positional masking and without any positional encoding are also Turing-complete. We further analyze the necessity of each component for the Turing-completeness of the network; interestingly, we find that a particular type of residual connection is necessary. We demonstrate the practical implications of our results via experiments on machine translation and synthetic tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Transformers versus the EM Algorithm in Multi-class Clustering

    stat.ML 2025-02 conditional novelty 6.0 of 10

    A pretrained transformer can approximate Lloyd's EM algorithm for multi-class Gaussian clustering and can achieve the minimax optimal clustering error with enough pretraining data.

  2. A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization

    cs.LG 2025-06 conditional novelty 5.0 of 10

    Single-layer linear self-attention can represent, train on, and length-generalize pairwise interaction functions under data-versatility and exact-realizability assumptions, and the paper introduces higher-order HyperA...

  3. Eliciting Fine-Tuned Transformer Capabilities via Inference-Time Techniques

    cs.LG 2025-06 reject novelty 2.0 of 10

    The paper claims that in-context learning with finite example sets can approximate supervised fine-tuning in transformers, but the proof assumes the very approximation it sets out to establish.

Pith tools