REVIEW 7 cited by
Transformer Language Models without Positional Encodings Still Learn Positional Information
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Causal transformer language models (LMs), such as GPT-3, typically require some form of positional encoding, such as positional embeddings. However, we show that LMs without any explicit positional encoding are still competitive with standard models, and that this phenomenon is robust across different datasets, model sizes, and sequence lengths. Probing experiments reveal that such models acquire an implicit notion of absolute positions throughout the network, effectively compensating for the missing information. We conjecture that causal attention enables the model to infer the number of predecessors that each token can attend to, thereby approximating its absolute position. Our findings indicate that causal LMs might derive positional awareness not only from the explicit positioning mechanism, but also from the effects of the causal mask.
Forward citations
Cited by 7 Pith papers
-
CEHR-XGPT: A Scalable Multi-Task Foundation Model for Electronic Health Records
CEHR-XGPT unifies feature representation, zero-shot prediction, and synthetic data generation in a single GPT-2 style EHR model using artificial time tokens with time-decomposition and time-to-event losses.
-
Multispin Physics of AI Tipping Points and Hallucinations
A closed-form formula predicts the iteration at which a simplified attention head tips from good to bad output, determined only by token embedding dot products.
-
SeqPE: Transformer with Sequential Position Encoding
SeqPE encodes each position as a symbolic digit sequence through a small Transformer, and with contrastive plus distillation losses it reports improved extrapolation in language, QA, and image classification.
-
Unlocking the Power of Diffusion Models in Sequential Recommendation: A Simple and Effective Approach
ADRec applies token-level, per-token diffusion with causal attention to sequential recommendation, reducing embedding collapse and outperforming ten baselines on six datasets.
-
AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally
Head-wise learnable rotary frequencies and length-dependent attention scaling (AdaRoPE) beat uniform RoPE and YaRN schedules in pretraining and 8k-to-64k context extension up to 8B scale.
-
BALM-TSF: Balanced Multimodal Alignment for LLM-Based Time Series Forecasting
BALM-TSF combines a statistical-prompt text branch with a patch-based time series branch, using scaling plus contrastive alignment to balance the two modalities, improving long-term and few-shot forecasting on five of...
-
HoPE: Hyperbolic Rotary Positional Encoding for Stable Long-Range Dependency Modeling in Large Language Models
HoPE replaces RoPE's sine/cosine rotations with hyperbolic functions plus an exponential damping term to enforce monotonic attention decay, but the claimed consistent superiority and the 'RoPE as special case' theorem...
Discussion (0). Continue with ORCID to comment.