REVIEW 5 major objections 5 minor 69 references
Softplus Attention with Re-weighting Boosts Length Extrapolation in Large Language Models
T0 review · 5 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read By splitting attention into a Softplus normalisation stage and a sharpening re-weight stage, the paper shows a 124M-parameter GPT-2 can keep nearly flat validation loss at 16 times its 1K training length.
desk verdict A legitimate two-stage attention idea undermined by a tuned sharpening exponent and an unablated NTK scaling; the 16x extrapolation claim is as yet unsupported. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is a decomposition of Softmax into a positivity map followed by an $\ell^1$-normalisation, generalised to $\phi(x)/\lVert \phi(x)\rVert_1$ and rebuilt as two explicit stages. In the normalisation stage, LSSA, the paper uses $\phi(x)=\mathrm{Softplus}(x)$ and multiplies the score matrix by $\log d \log N$, where $N$ is a matrix whose row $i$ counts the tokens attended to up to position $i$, a scale factor derived from the idea that row-wise entropy should stay invariant as sequence length changes. In the sharpening stage, the normalised weights are passed through $A \leftarrow \mathrm{ReLU}_p(A \otimes N - O)$ and re-normalised, where $N$ again counts attended tokens, $O$ is an offset matrix, and $\mathrm{ReLU}_p$ masks non-positive entries and raises the rest to power $p$. The limit $\lim_{p\to\infty} (x_m - x_l) = 1$ for any $x_l < x_m$ is what guarantees the maximum weight tends to 1 and all smaller weights to 0, turning a smooth attention row into a sharp one and thereby preventing the attention smoothing that the paper identifies as the cause of poor length extrapolation.
What would settle it
Train the same 124M-parameter GPT-2 with LSSAR and with Softmax under identical conditions, switch off the test-time position-embedding stretch, and fix $p=15$ before inspecting any length sweep. If LSSAR's validation loss at 8K or 16K tokens then rises as steeply as Softmax's, the re-weighting mechanism is not the cause of the extrapolation; if Softmax also flattens when the stretch is present, the two-stage attention adds nothing beyond the shared test-time support.
Extended reading notes
Core claim
The paper's central discovery is that Softmax's contribution to a language model is its $\ell^1$-normalisation, not its positivity: inverting or re-centring the attention scores barely changes validation loss, while restoring $\ell^1$-normalisation recovers almost all performance. On that basis, the paper defines a general form $\phi(x)/\lVert \phi(x)\rVert_1$, instantiates it with $\phi=\mathrm{Softplus}$, and scales the score matrix by $\log d \log N$ so that each row's entropy stays invariant as the number of attended tokens grows. The resulting normalisation stage, LSSA, is then followed by a sharpening stage, $A \leftarrow \mathrm{ReLU}_p(A \otimes N - O)$ with a final $\ell^1$-normalisation, whose power $p$ drives the largest attention weight toward 1 and all others toward 0. The paper claims this two-stage mechanism, LSSAR, maintains a nearly constant validation loss at 16x the training token length and outperforms Softmax and Softmax-free baselines on long-context retrieval and downstream benchmarks.
Load-bearing premise
The claim stands on the assumption that the flat 16x-length loss comes from the new attention re-weighting rather than from the test-time position-embedding stretch (Dynamic NTK Scaling) applied to all models, or from choosing the sharpening strength $p=15$ after seeing the results.
Editorial extensions
If this is right
- Trained on 1024-token contexts, LSSAR holds its validation loss nearly flat out to 16,384 tokens, so a small model can be deployed on much longer documents without retraining.
- Passkey retrieval stays non-zero at 8K tokens under LSSAR, while Softmax attention drops to 0% beyond 1.5K tokens, showing the sharpening stage prevents the model from losing a single critical token in long contexts.
- The re-weighting operation improves extrapolation for other normalised attention variants at $p=3$, indicating the sharpening stage can be added to existing attention designs rather than requiring a full architecture change.
- Because the improvement is inside the attention computation, LSSAR combines with RoPE and Dynamic NTK Scaling, and the authors argue it should transfer to larger RoPE-based models by analogy with deeper networks.
- Softplus's bounded derivative keeps training stable at large $p$ values, where Softmax-based re-weighting suffers gradient explosion, so the sharpening strength can be increased without numerical failure.
Reading between the lines
- Editorial extension: the sharpening stage appears transferable, since the paper applies it to Softmax, Sigmoid, and LSSA attention; a natural untested step is fine-tuning an already-pretrained model with re-weighting instead of training from scratch.
- Editorial extension: the abstract's symbolic-regression claim about recovering Newton's gravitational law is not developed in the provided experimental section, so it should be read as a secondary, unverified assertion rather than part of the length-extrapolation evidence.
- Editorial extension: the paper's Discussion supports larger models only by analogy between large $p$ and increased depth, so scaling to billions of parameters remains an extrapolation rather than a demonstrated result.
- Editorial extension: a principled choice of $p$ from the desired attention entropy or the score distribution's tail could remove the empirical sweep and is not attempted; a length-dependent $p$ schedule may push the flat-loss region beyond 16x.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a two-stage redesign of self-attention. The first stage (LSSA) decomposes Softmax into a positive nonlinearity plus l1-normalization, replaces the exponential with Softplus, and applies a scale factor log d log N. The second stage (LSSAR) sharpens the normalized attention distribution through a power-based re-weighting operation. The authors train GPT-2-small (124M) on 1024-token sequences and report that LSSAR maintains a nearly constant validation loss up to 16x the training length, outperforms several Softmax-free attention variants on length extrapolation, succeeds on passkey retrieval, and improves downstream benchmarks. The abstract also claims, without any supporting experiments in the body, that the method recovers Newton's gravitational law via symbolic regression.
Significance. If the extrapolation claims are correct, the two-stage decomposition offers a useful design principle for attention and a potential alternative to positional-interpolation fixes. The paper contributes a clean decomposition experiment (Table 2), a broad comparison against ReLU- and Sigmoid-based attentions under shared settings (Table 4), and a release of code, all of which are valuable. However, the central headline result is currently contingent on two unestablished choices: Dynamic NTK Scaling is applied to every model at inference and is never ablated, and the sharpening exponent p=15 is selected from the very extrapolation lengths the paper aims to predict. The significance is therefore conditional on additional experiments that separate the attention mechanism from these test-time and tuning effects.
major comments (5)
- [Experiments, first paragraph; Introduction paragraph 2] Dynamic NTK Scaling is incorporated into every model at inference and is never ablated. Dynamic NTK Scaling is itself a post-hoc RoPE-based positional embedding adjustment, the same class of technique the Introduction says LSSAR renders unnecessary. Consequently, the claim that LSSAR 'fundamentally improves length extrapolation' through its two-stage attention mechanism is not supported: the gains could arise from the interaction between the re-weighting and the unablated NTK scaling. Please provide an ablation of LSSAR (and, for completeness, Softmax) with and without Dynamic NTK Scaling, and report the extrapolation losses in all four configurations.
- [Ablation Study for Re-weighting Mechanism; Table 4; Fig. 2] The parameter p=15 is selected from the extrapolation lengths themselves. The ablation sweeps p over 1-15, 50, and 100 using validation loss at 1K-16K, and the paper then reports LSSAR(p=15) as the headline result. Table 4 shows that at p=3, which has the best training-length loss (1K: 3.1782 vs 3.1905 for p=15), extrapolation fails: 4K loss is 5.4056 and 8K loss is 6.3007, worse than both unsharpened LSSA (5.9403 at 8K) and Softmax (6.2823 at 8K). Equation (6) only establishes behavior as p approaches infinity and offers no principle for choosing p=15. The near-constant-loss claim is therefore a result of tuning p on the evaluation lengths. Please fix p a priori using a validation split that does not overlap the extrapolation lengths, or derive a principled selection criterion for p.
- [Abstract] The abstract states that 'symbolic regression experiments demonstrate that our method enables models to recover Newton's gravitational law from orbital trajectory sequences.' No symbolic regression experiments, orbital trajectory data, or Newton's-law results appear anywhere in the manuscript. This is a claimed contribution that is entirely missing from the body. Either add the experiments and their results, or remove the claim from the abstract.
- [Experiments; Tables 2-4; Fig. 2] No confidence intervals or repeated seeds are reported. All loss values appear to come from single runs. The central 'nearly constant' validation loss of LSSAR(p=15) (3.1905, 3.1930, 3.2291, 3.3171 at 1K, 2K, 4K, 8K) may be within ordinary run-to-run variability for a 124M-parameter model. Please report means and standard deviations over at least 3 seeds and, if possible, a paired comparison that tests whether the loss increase from 1K to 8K is statistically distinguishable from that of the baselines.
- [Length Scaled Softplus Attention, Eq. (4)] The scale factor log d log N is introduced with the statement that it 'ensures entropy invariance' across sequence lengths, but no derivation or citation for this specific combination is provided. As written, the factor is an additional free parameter of LSSA. Please either supply a derivation showing why log d log N preserves entropy invariance, or explicitly present it as an empirical design choice and ablate its contribution to the extrapolation results.
minor comments (5)
- [Attention Re-weighting Mechanism, Eq. (6)] Equation (6) appears to misstate the quantity being computed: the left-hand side 'xm - xl' is the original score difference, but the limit expression is the difference between the re-weighted values after the power transformation. The intended statement is that the re-weighted distance approaches 1 as p goes to infinity, not that the original score difference equals 1. Please correct the notation.
- [Downstream Evaluation, Table 5] The SummScreen ROUGE-1 scores of 1.682 and 6.309 are far below typical ROUGE-1 values for summarization benchmarks. Please clarify the metric configuration or verify that the reported numbers are not the result of an evaluation error.
- [Comparison with State-of-the-Art Softmax-Free Attention Functions] The text repeatedly refers to 'state-of-the-art Softmax-free alternatives,' but the comparison set contains only three ReLU-based and two Sigmoid-based attention variants, not the full range of recent Softmax-free attention methods. Consider softening the characterization or expanding the baseline list.
- [Fig. 2] The x-axis labels '1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 50100' are cramped and difficult to read. Use a log scale or split the axis to show the p=50 and p=100 points clearly.
- [Ablation Study for Re-weighting Mechanism] The text says 'optimal results observed around p = 15,' but the sweep includes p=50 and p=100 with noticeably higher loss, so 'around p=15' should be clarified to mean, for example, p between 10 and 20, rather than implying a plateau that extends to p=100.
Circularity Check
The near-constant extrapolation loss is reported after p=15 is selected from the same extrapolation lengths in Fig. 2; at p=3 LSSAR fails to extrapolate, so the central claim is a tuned result rather than an independent prediction.
-
fitted input called prediction
[Ablation Study for Re-weighting Mechanism (Fig. 2 and Table 4)]
"increasing p values generally improve the performance of LSSA across different sequence lengths, with optimal results observed around p = 15. The validation loss remains relatively stable as the sequence length increases, underscoring LSSA’s strong scalability. ... We compared LSSAR (p = 15) against the standard Softmax attention baseline."
The central claim that LSSAR maintains nearly constant validation loss at 16x the training length is not an independent prediction: the sharpening exponent p is a hyperparameter in Eq. (5), and p=15 is chosen as optimal on the very sequence lengths (1K to 16K) used to report the near-constant loss. The paper's own Table 4 shows that with p=3, LSSAR loses extrapolation stability (loss 5.4056 at 4K and 6.3007 at 8K), so the flat-loss behavior is a property of the tuned configuration, not of the two-stage attention mechanism per se. Eq. (6) only characterizes p approaching infinity and provides no principled reason to select p=15. Reporting the curve after optimizing p on that same curve converts a tuned design into the appearance of a predicted extrapolation result.
full rationale
The comparison against Softmax, Sigmoid, and ReLU baselines under shared training settings is genuine evidence and not tautological, and there is no load-bearing self-citation chain: the only self-citation (Gao and Pavel 2017) supports background material on softmax properties. The circular component is confined to the headline extrapolation result: p is optimized on the extrapolation lengths in Fig. 2, and the near-constant-loss claim is then reported for that same p=15 configuration. Since p=3, which is better at the training length, fails to extrapolate, the result is best described as a tuned design rather than a prediction from first principles. Additionally, all models use Dynamic NTK Scaling at inference, a post-hoc RoPE extrapolation fix that is never ablated, so the paper's inference that the attention mechanism alone 'fundamentally' improves extrapolation is not isolated from this positional-embedding adjustment. The passkey and downstream results at p=15 are not fitted to those particular tasks and retain some independent content, but the central length-extrapolation claim itself partially reduces to a hyperparameter selection on the reported curve, yielding a circularity score of 6 rather than a full 8 or 10.
Assumptions & free parameters
free parameters (2)
- Re-weighting exponent p =
15 (main results); 3, 50, 100 also tested
- Length scale factor log d log N =
log 64 * log N (with d=64)
assumptions (4)
- domain assumption The crucial component of softmax is l1-normalisation, not positivity or non-negativity.
- domain assumption Attention smoothing is the cause of length extrapolation failure and sharpening fixes it.
- domain assumption Dynamic NTK Scaling is neutral or orthogonal to the attention comparison.
- ad hoc to paper The log d log N scale factor preserves entropy invariance across sequence lengths.
Cite this review
Pith. "Pith review of Softplus Attention with Re-weighting Boosts Length Extrapolation in Large Language Models." pith.science (2026). https://pith.science/paper/4FWCNJDN
@misc{pith2026250113428,
author = {Pith},
title = {Pith review of: Softplus Attention with Re-weighting Boosts Length Extrapolation in Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/4FWCNJDN}},
note = {Machine review of arXiv:2501.13428}
}
abstract
Large language models have achieved remarkable success in recent years, primarily due to self-attention. However, traditional Softmax attention suffers from numerical instability and reduced performance as the number of inference tokens increases. This work addresses these issues by proposing a new design principle for attention, viewing it as a two-stage process. The first stage (normalisation) refines standard attention by replacing Softmax with the more numerically stable Softplus followed by $l_{1}$-normalisation. Furthermore, we introduce a dynamic scale factor based on invariance entropy. We show that this novel attention mechanism outperforms conventional Softmax attention, and state-of-the-art Softmax-free alternatives. Our second proposal is to introduce a second processing stage (sharpening) which consists of a re-weighting mechanism that amplifies significant attentional weights while diminishing weaker ones. This enables the model to concentrate more effectively on relevant tokens, mitigating the attention sink phenomenon, and fundamentally improving length extrapolation. This novel, two-stage, replacement for self-attention is shown to ensure numerical stability and dramatically improve length extrapolation, maintaining a nearly constant validation loss at 16$\times$ the training length while achieving superior results on challenging long-context retrieval tasks and downstream benchmarks. Furthermore, symbolic regression experiments demonstrate that our method enables models to recover Newton's gravitational law from orbital trajectory sequences, providing evidence that appropriate attention mechanisms are crucial for foundation models to develop genuine physical world models. Our code is available at https://github.com/iminfine/freeattn.
Figures
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023. GPT -4 technical report. arXiv:2303.08774
arXiv 2023
-
[4]
Arora, S.; Eyuboglu, S.; Zhang, M.; Timalsina, A.; Alberti, S.; Zinsley, D.; Zou, J.; Rudra, A.; and R \'e , C. 2024. Simple linear attention language models balance the recall-throughput tradeoff. arXiv:2402.18668
arXiv 2024
- [5]
-
[6]
Ba, J.; Kiros, J. R.; and Hinton, G. E. 2016. Layer normalization. In Proceedings of the 33rd International Conference on Machine Learning (ICML), 198--206. PMLR
work page 2016
-
[7]
Bai, Y.; Chen, F.; Wang, H.; Xiong, C.; and Mei, S. 2024. Transformers as statisticians: Provable in-context learning with in-context algorithm selection. Advances in neural information processing systems, 36
work page 2024
-
[8]
Bisk, Y.; Zellers, R.; Gao, J.; Choi, Y.; et al. 2020. Piqa: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI conference on artificial intelligence, volume 34, 7432--7439
2020
Show all 69 references
-
[9]
by parts
bloc97. 2023 a . Add NTK -Aware interpolation "by parts" correction. GitHub Pull Request
2023
-
[10]
bloc97. 2023 b . NTK -Aware Scaled RoPE allows LLaMA models to have extended (8k+) context size without any fine-tuning and minimal perplexity degradation. Reddit post
2023
-
[11]
Chen, M.; Chu, Z.; Wiseman, S.; and Gimpel, K. 2021. SummScreen: A dataset for abstractive screenplay summarization. arXiv:2104.07091
2021 arXiv
-
[12]
Chen, S.; Wong, S.; Chen, L.; and Tian, Y. 2023. Extending Context Window of Large Language Models via Positional Interpolation
2023
-
[13]
J.; and Rudnicky, A
Chi, T.-C.; Fan, T.-H.; Ramadge, P. J.; and Rudnicky, A. 2022. KERPLE: Kernelized Relative Positional Embedding for Length Extrapolation. In Advances in Neural Information Processing Systems, volume 35, 8386--8399
2022
-
[14]
Chiang, D.; and Cholak, P. 2022. Overcoming a Theoretical Limitation of Self-Attention. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7654--7664
2022
-
[15]
Clark, P.; Cowhey, I.; Etzioni, O.; Khot, T.; Sabharwal, A.; Schoenick, C.; and Tafjord, O. 2018. Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge. arXiv:1803.05457
2018 arXiv
-
[16]
DeepSeek-AI; Guo, D.; Yang, D.; Zhang, H.; Song, J.; Zhang, R.; Xu, R.; Zhu, Q.; Ma, S.; Wang, P.; Bi, X.; Zhang, X.; Yu, X.; Wu, Y.; Wu, Z. F.; Gou, Z.; Shao, Z.; Li, Z.; Gao, Z.; Liu, A.; Xue, B.; Wang, B.; Wu, B.; Feng, B.; Lu, C.; Zhao, C.; Deng, C.; Zhang, C.; Ruan, C.; D...
2025 arXiv
-
[17]
P.; Caron, M.; Geirhos, R.; Alabdul mohsin, I.; Jenatton, R.; Beyer, L.; Tschannen, M.; Arnab, A.; Wang, X.; Ruiz, C
Dehghani, M.; Djolonga, J.; Mustafa, B.; Padlewski, P.; Heek, J.; Gilmer, J.; Steiner, A. P.; Caron, M.; Geirhos, R.; Alabdul mohsin, I.; Jenatton, R.; Beyer, L.; Tschannen, M.; Arnab, A.; Wang, X.; Ruiz, C. R.; Minderer, M.; Puigcerver, J.; Evci, U.; Kumar, M.; van Steenkiste...
2023
-
[18]
Dubey, A.; Jauhri, A.; Pandey, A.; Kadian, A.; Al-Dahle, A.; Letman, A.; Mathur, A.; Schelten, A.; Yang, A.; Fan, A.; et al. 2024. The Llama 3 herd of models. arXiv:2407.21783
2024 arXiv
-
[19]
emozilla. 2023. Dynamically Scaled RoPE further increases performance of long context LLaMA with zero fine-tuning. Reddit post
2023
-
[20]
Fu, H.; Guo, T.; Bai, Y.; and Mei, S. 2024. What can a single attention layer learn? a study through the random features lens. Advances in Neural Information Processing Systems, 36
2024
-
[21]
Gao, B.; and Pavel, L. 2017. On the properties of the softmax function with application in game theory and reinforcement learning. arXiv:1704.00805
2017 arXiv
-
[22]
Gao, L.; Tow, J.; Abbasi, B.; Biderman, S.; Black, S.; DiPofi, A.; Foster, C.; Golding, L.; Hsu, J.; Le Noac'h, A.; Li, H.; McDonell, K.; Muennighoff, N.; Ociepa, C.; Phang, J.; Reynolds, L.; Schoelkopf, H.; Skowron, A.; Sutawika, L.; Tang, E.; Thite, A.; Wang, B.; Wang, K.; a...
2024
-
[23]
Golovneva, O.; Wang, T.; Weston, J.; and Sukhbaatar, S. 2024. Contextual Position Encoding: Learning to Count What's Important. arXiv
2024
-
[24]
Han, D.; Pan, X.; Han, Y.; Song, S.; and Huang, G. 2023. Flatten transformer: Vision transformer using focused linear attention. In Proceedings of the IEEE/CVF international conference on computer vision, 5961--5971
2023
-
[25]
He, Z.; Feng, G.; Luo, S.; Yang, K.; He, D.; Xu, J.; Zhang, Z.; Yang, H.; and Wang, L. 2024. Two Stones Hit One Bird: Bilevel Positional Encoding for Better Length Extrapolation. arXiv
2024
-
[26]
Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2020. Measuring massive multitask language understanding
2020
-
[27]
Hendrycks, D.; and Gimpel, K. 2016. Gaussian Error Linear Units ( GELUs ). arXiv:1606.08415
2016 arXiv
-
[28]
R.; Pawar, S
Henry, A.; Dachapally, P. R.; Pawar, S. S.; and Chen, Y. 2020. Query-key normalization for transformers. In Findings of the Association for Computational Linguistics: EMNLP 2020, 4246--4253. Association for Computational Linguistics
2020
-
[29]
G.; Zhu, M.; Chen, B.; Kalenichenko, D.; Wang, W.; Weyand, T.; Andreetto, M.; and Adam, H
Howard, A. G.; Zhu, M.; Chen, B.; Kalenichenko, D.; Wang, W.; Weyand, T.; Andreetto, M.; and Adam, H. 2017. MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVP...
2017
-
[30]
Hron, J.; Bahri, Y.; Sohl-Dickstein, J.; and Novak, R. 2020. Infinite attention: NNGP and NTK for deep attention networks. In International Conference on Machine Learning, 4376--4386. PMLR
2020
-
[31]
Hua, W.; Dai, Z.; Liu, H.; and Le, Q. 2022. Transformer quality in linear time. In International conference on machine learning, 9099--9117. PMLR
2022
-
[32]
Huang, Z.; Liang, D.; Xu, P.; and Xiang, B. 2020. Improve Transformer Models with Better Relative Position Embeddings. In Findings of the Association for Computational Linguistics: EMNLP 2020, 3327--3335. Online: Association for Computational Linguistics
2020
-
[33]
Jang, E.; Gu, S.; and Poole, B. 2016. Categorical reparameterization with gumbel-softmax. arXiv:1611.01144
2016 arXiv
-
[34]
kaiokendev. 2023. Things I'm learning while training superhot. Accessed: [Insert Access Date]
2023
-
[35]
Katharopoulos, A.; Vyas, A.; Pappas, N.; and Fleuret, F. 2020. Transformers are RNN : Fast autoregressive transformers with linear attention. In International conference on machine learning, 5156--5165. PMLR
2020
-
[36]
N.; Das, P.; and Reddy, S
Kazemnejad, A.; Padhi, I.; Ramamurthy, K. N.; Das, P.; and Reddy, S. 2023. The Impact of Positional Encoding on Length Generalization in Transformers. arXiv:2305.19466
2023 arXiv
-
[37]
Kiyono, S.; Kobayashi, S.; Suzuki, J.; and Inui, K. 2021. SHAPE: Shifted Absolute Position Embedding for Transformers. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 3309--3321. Online and Punta Cana, Dominican Republic: Association ...
2021
-
[38]
Lai, Z.; Lim, L.-H.; and Liu, Y. 2024. Attention is a smoothed cubic spline. arXiv:2408.09624
2024 arXiv
-
[39]
Li, S.; You, C.; Guruganesh, G.; Ainslie, J.; Ontanon, S.; Zaheer, M.; Sanghai, S.; Yang, Y.; Kumar, S.; and Bhojanapalli, S. 2023. Functional Interpolation for Relative Positions Improves Long Context Transformers
2023
-
[40]
Li, Z.; Bhojanapalli, S.; Zaheer, M.; Reddi, S.; and Kumar, S. 2022. Robust training of neural networks using scale invariant architectures. In International Conference on Machine Learning, 12656--12684. PMLR
2022
-
[41]
Likhomanenko, T.; Xu, Q.; Synnaeve, G.; Collobert, R.; and Rogozhnikov, A. 2021. CAPE : Encoding Relative Positions with Continuous Augmented Positional Embeddings. In Advances in Neural Information Processing Systems, volume 34, 16079--16092. Curran Associates, Inc
2021
-
[42]
Liu, A.; Feng, B.; Xue, B.; Wang, B.; Wu, B.; Lu, C.; Zhao, C.; Deng, C.; Zhang, C.; Ruan, C.; et al. 2024. Deepseek-v3 technical report. arXiv:2412.19437
2024 arXiv
-
[43]
Liu, X.; Yan, H.; Zhang, S.; An, C.; Qiu, X.; and Lin, D. 2023. Scaling Laws of RoPE-based Extrapolation. arXiv
2023
-
[44]
Liu, Z.; Hu, H.; Lin, Y.; Yao, Z.; Xie, Z.; Wei, Y.; Ning, J.; Cao, Y.; Zhang, Z.; Dong, L.; Wei, F.; and Guo, B. 2022. Swin Transformer V2: Scaling Up Capacity and Resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 11999--...
2022
-
[45]
Lu, J.; Yao, J.; Zhang, J.; Zhu, X.; Xu, H.; Gao, W.; Xu, C.; Xiang, T.; and Zhang, L. 2021. Soft: Softmax-free transformer with linear complexity. Advances in Neural Information Processing Systems, 34: 21297--21309
2021
-
[46]
Misra, D. 2019. Mish: A Self Regularized Non-Monotonic Neural Activation Function. arXiv:1908.08681
2019 arXiv
-
[47]
Mohtashami, A.; and Jaggi, M. 2023. Random-access infinite context length for transformers. Advances in Neural Information Processing Systems, 36: 54567--54585
2023
-
[48]
B.; Lozhkov, A.; Mitchell, M.; Raffel, C.; Werra, L
Penedo, G.; Kydlíček, H.; allal, L. B.; Lozhkov, A.; Mitchell, M.; Raffel, C.; Werra, L. V.; and Wolf, T. 2024. The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale. arXiv:2406.17557
2024 arXiv
-
[49]
Qi, X.; Ye, J.; He, Y.; Li, C.-G.; Zi, B.; Dai, X.; Zou, Q.; and Xiao, R. 2024. Stable-Transformer: Towards a Stable Transformer Training. Accessed: 2024-11-15
2024
-
[50]
Qin, Z.; Sun, W.; Deng, H.; Li, D.; Wei, Y.; Lv, B.; Yan, J.; Kong, L.; and Zhong, Y. 2022. cosFormer : Rethinking Softmax In Attention. In International Conference on Learning Representations
2022
-
[51]
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; and Sutskever, I. 2019. Language Models are Unsupervised Multitask Learners
2019
-
[52]
Ramapuram, J.; Danieli, F.; Dhekane, E.; Weers, F.; Busbridge, D.; Ablin, P.; Likhomanenko, T.; Digani, J.; Gu, Z.; Shidani, A.; et al. 2024. Theory, Analysis, and Best Practices for Sigmoid Self-Attention. arXiv:2409.04431
2024 arXiv
-
[53]
E.; Hinton, G
Rumelhart, D. E.; Hinton, G. E.; and Williams, R. J. 1986. Learning Internal Representations by Error Propagation. Nature, 323(6088): 533--536
1986
-
[54]
Shah, J.; Bikshandi, G.; Zhang, Y.; Thakkar, V.; Ramani, P.; and Dao, T. 2024. Flashattention-3: Fast and accurate attention with asynchrony and low-precision. arXiv:2407.08608
2024 arXiv
-
[55]
Shen, K.; Guo, J.; Tan, X.; Tang, S.; Wang, R.; and Bian, J. 2023. A study on ReLU and softmax in transformer. arXiv:2302.06461
2023 arXiv
-
[56]
So, D.; Ma \'n ke, W.; Liu, H.; Dai, Z.; Shazeer, N.; and Le, Q. V. 2021. Searching for efficient transformers for language modeling. Advances in neural information processing systems, 34: 6010--6022
2021
-
[57]
Su, J. 2021. Viewing the scale operation of attention from the perspective of entropy invariance
2021
-
[58]
Su, J.; Ahmed, M.; Lu, Y.; Pan, S.; Bo, W.; and Liu, Y. 2024. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 568: 127063
2024
-
[59]
H.; Bai, S.; Yamada, M.; Morency, L.-P.; and Salakhutdinov, R
Tsai, Y.-H. H.; Bai, S.; Yamada, M.; Morency, L.-P.; and Salakhutdinov, R. 2019. Transformer Dissection: An Unified Understanding for Transformer's Attention via the Lens of Kernel. In Proceedings of the Conference on Empirical Methods in Natural Language Processing
2019
-
[60]
Veli c kovi \'c , P.; Perivolaropoulos, C.; Barbero, F.; and Pascanu, R. 2024. Softmax is not Enough (for Sharp Size Generalisation). arXiv:2410.01104
2024 arXiv
-
[61]
Wang, S.; Kobyzev, I.; Lu, P.; Rezagholizadeh, M.; and Liu, B. 2024. Resonance RoPE : Improving Context Length Generalization of Large Language Models. arXiv
2024
-
[62]
F.; and Gardner, M
Welbl, J.; Liu, N. F.; and Gardner, M. 2017. Crowdsourcing multiple choice science questions. arXiv:1707.06209
2017 arXiv
-
[63]
Wortsman, M.; Lee, J.; Gilmer, J.; and Kornblith, S. 2023. Replacing softmax with ReLU in vision transformers. arXiv:2309.08586
2023 arXiv
-
[64]
Wu, M.; Cheng, X.; Padon, O.; and Jia, Z. 2024. A Multi-Level Superoptimizer for Tensor Programs. arXiv:2405.05751
2024 arXiv
-
[65]
Yang, A.; Yang, B.; Hui, B.; Zheng, B.; Yu, B.; Zhou, C.; Li, C.; Li, C.; Liu, D.; Huang, F.; et al. 2024. Qwen2 technical report. arXiv:2407.10671
2024 arXiv
-
[66]
Zellers, R.; Holtzman, A.; Bisk, Y.; Farhadi, A.; and Choi, Y. 2019. Hellaswag: Can a machine really finish your sentence? arXiv:1905.07830
2019 arXiv
-
[67]
Zhang, Y.; Liu, Y.; Yuan, H.; Qin, Z.; Yuan, Y.; Gu, Q.; and Yao, A. C.-C. 2025. Tensor Product Attention Is All You Need. arXiv:2501.06425
2025
-
[68]
Zheng, C.; Gao, Y.; Shi, H.; Huang, M.; Li, J.; Xiong, J.; Ren, X.; Ng, M.; Jiang, X.; Li, Z.; and Li, Y. 2024. CAPE : Context-Adaptive Positional Encoding for Length Extrapolation. arXiv
2024
-
[69]
Zheng, H.; Yang, Z.; Liu, W.; Liang, J.; and Li, Y. 2015. Improving deep neural networks using softplus units. In 2015 International joint conference on neural networks (IJCNN), 1--4. IEEE
2015
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.