Scaling attention logits by s log(context size) prevents attention from flattening in long contexts and improves length generalization and key-information retrieval in a 162M-parameter language model.
Redpajama: An open source recipe to reproduce llama training dataset, April 2023
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Scalable-Softmax Is Superior for Attention
Scaling attention logits by s log(context size) prevents attention from flattening in long contexts and improves length generalization and key-information retrieval in a 162M-parameter language model.