Pith. sign in

REVIEW 1 cited by

The Case for Translation-Invariant Self-Attention in Transformer-Based Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.01950 v1 pith:3LXIAZ7S submitted 2021-06-03 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords embeddingslanguagemodelspositionself-attentionexistinginvariancepositional
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Mechanisms for encoding positional information are central for transformer-based language models. In this paper, we analyze the position embeddings of existing language models, finding strong evidence of translation invariance, both for the embeddings themselves and for their effect on self-attention. The degree of translation invariance increases during training and correlates positively with model performance. Our findings lead us to propose translation-invariant self-attention (TISA), which accounts for the relative position between tokens in an interpretable fashion without needing conventional position embeddings. Our proposal has several theoretical advantages over existing position-representation approaches. Experiments show that it improves on regular ALBERT on GLUE tasks, while only adding orders of magnitude less positional parameters.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. KoopAGRU: A Koopman-based Anomaly Detection in Time-Series using Gated Recurrent Units

    cs.LG 2025-01 conditional novelty 5.0 of 10

    KoopAGRU, a GRU-based Koopman model with FFT time-variant/invariant decomposition, reports an average F1 of 90.88% on five anomaly detection benchmarks, exceeding cited baselines.

Pith tools