Pith. sign in

REVIEW 7 cited by

Transformer Language Models without Positional Encodings Still Learn Positional Information

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.16634 v2 pith:OJVQRWQ7 submitted 2022-03-30 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords positionalcausalmodelsabsoluteencodingexplicitinformationlanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Causal transformer language models (LMs), such as GPT-3, typically require some form of positional encoding, such as positional embeddings. However, we show that LMs without any explicit positional encoding are still competitive with standard models, and that this phenomenon is robust across different datasets, model sizes, and sequence lengths. Probing experiments reveal that such models acquire an implicit notion of absolute positions throughout the network, effectively compensating for the missing information. We conjecture that causal attention enables the model to infer the number of predecessors that each token can attend to, thereby approximating its absolute position. Our findings indicate that causal LMs might derive positional awareness not only from the explicit positioning mechanism, but also from the effects of the causal mask.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CEHR-XGPT: A Scalable Multi-Task Foundation Model for Electronic Health Records

    cs.LG 2025-09 conditional novelty 6.0 of 10

    CEHR-XGPT unifies feature representation, zero-shot prediction, and synthetic data generation in a single GPT-2 style EHR model using artificial time tokens with time-decomposition and time-to-event losses.

  2. Multispin Physics of AI Tipping Points and Hallucinations

    cs.AI 2025-08 conditional novelty 6.0 of 10

    A closed-form formula predicts the iteration at which a simplified attention head tips from good to bad output, determined only by token embedding dot products.

  3. SeqPE: Transformer with Sequential Position Encoding

    cs.LG 2025-06 reject novelty 6.0 of 10

    SeqPE encodes each position as a symbolic digit sequence through a small Transformer, and with contrastive plus distillation losses it reports improved extrapolation in language, QA, and image classification.

  4. Unlocking the Power of Diffusion Models in Sequential Recommendation: A Simple and Effective Approach

    cs.IR 2025-05 conditional novelty 6.0 of 10

    ADRec applies token-level, per-token diffusion with causal attention to sequential recommendation, reducing embedding collapse and outperforming ten baselines on six datasets.

  5. AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally

    cs.AI 2026-06 conditional novelty 5.0 of 10

    Head-wise learnable rotary frequencies and length-dependent attention scaling (AdaRoPE) beat uniform RoPE and YaRN schedules in pretraining and 8k-to-64k context extension up to 8B scale.

  6. BALM-TSF: Balanced Multimodal Alignment for LLM-Based Time Series Forecasting

    cs.AI 2025-08 conditional novelty 5.0 of 10

    BALM-TSF combines a statistical-prompt text branch with a patch-based time series branch, using scaling plus contrastive alignment to balance the two modalities, improving long-term and few-shot forecasting on five of...

  7. HoPE: Hyperbolic Rotary Positional Encoding for Stable Long-Range Dependency Modeling in Large Language Models

    cs.CL 2025-09 reject novelty 4.0 of 10

    HoPE replaces RoPE's sine/cosine rotations with hyperbolic functions plus an exponential damping term to enforce monotonic attention decay, but the claimed consistent superiority and the 'RoPE as special case' theorem...

Pith tools