Pith. sign in

REVIEW 5 major objections 5 minor 69 references

Softplus Attention with Re-weighting Boosts Length Extrapolation in Large Language Models

T0 review · 5 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read By splitting attention into a Softplus normalisation stage and a sharpening re-weight stage, the paper shows a 124M-parameter GPT-2 can keep nearly flat validation loss at 16 times its 1K training length.

desk verdict A legitimate two-stage attention idea undermined by a tuned sharpening exponent and an unablated NTK scaling; the 16x extrapolation claim is as yet unsupported. read the letter →

arxiv 2501.13428 v6 pith:4FWCNJDN submitted 2025-01-23 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords softmax-freeattentionlengthextrapolationSoftplusactivationre-weightinglargelanguagemodelssmoothingnumericalstabilityl1-normalisation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Attention in transformers is usually identified with Softmax, but this paper argues that Softmax is two separate operations—making scores positive and normalising them—and that the normalisation is what actually matters. The authors replace the exponential with the numerically stable Softplus function, normalise each attention row with an $\ell^1$-norm, and scale the scores by $\log d \log N$, a per-row factor chosen from invariance-entropy reasoning. A second stage then re-weights the distribution by raising positive entries to a power $p$, which sharpens strong attention peaks and suppresses weak ones, directly countering the attention smoothing that makes long contexts fail. On a 124M-parameter GPT-2 trained with 1024-token sequences, the paper reports nearly constant validation loss out to 16x that length and non-zero passkey retrieval at 8K tokens, while Softmax collapses to zero accuracy beyond 1.5K tokens. If the result holds, length extrapolation can be improved by redesigning the attention computation itself, without post-hoc positional interpolation.

What carries the argument

The load-bearing machinery is a decomposition of Softmax into a positivity map followed by an $\ell^1$-normalisation, generalised to $\phi(x)/\lVert \phi(x)\rVert_1$ and rebuilt as two explicit stages. In the normalisation stage, LSSA, the paper uses $\phi(x)=\mathrm{Softplus}(x)$ and multiplies the score matrix by $\log d \log N$, where $N$ is a matrix whose row $i$ counts the tokens attended to up to position $i$, a scale factor derived from the idea that row-wise entropy should stay invariant as sequence length changes. In the sharpening stage, the normalised weights are passed through $A \leftarrow \mathrm{ReLU}_p(A \otimes N - O)$ and re-normalised, where $N$ again counts attended tokens, $O$ is an offset matrix, and $\mathrm{ReLU}_p$ masks non-positive entries and raises the rest to power $p$. The limit $\lim_{p\to\infty} (x_m - x_l) = 1$ for any $x_l < x_m$ is what guarantees the maximum weight tends to 1 and all smaller weights to 0, turning a smooth attention row into a sharp one and thereby preventing the attention smoothing that the paper identifies as the cause of poor length extrapolation.

What would settle it

Train the same 124M-parameter GPT-2 with LSSAR and with Softmax under identical conditions, switch off the test-time position-embedding stretch, and fix $p=15$ before inspecting any length sweep. If LSSAR's validation loss at 8K or 16K tokens then rises as steeply as Softmax's, the re-weighting mechanism is not the cause of the extrapolation; if Softmax also flattens when the stretch is present, the two-stage attention adds nothing beyond the shared test-time support.

Watch

Extended reading notes

Core claim

The paper's central discovery is that Softmax's contribution to a language model is its $\ell^1$-normalisation, not its positivity: inverting or re-centring the attention scores barely changes validation loss, while restoring $\ell^1$-normalisation recovers almost all performance. On that basis, the paper defines a general form $\phi(x)/\lVert \phi(x)\rVert_1$, instantiates it with $\phi=\mathrm{Softplus}$, and scales the score matrix by $\log d \log N$ so that each row's entropy stays invariant as the number of attended tokens grows. The resulting normalisation stage, LSSA, is then followed by a sharpening stage, $A \leftarrow \mathrm{ReLU}_p(A \otimes N - O)$ with a final $\ell^1$-normalisation, whose power $p$ drives the largest attention weight toward 1 and all others toward 0. The paper claims this two-stage mechanism, LSSAR, maintains a nearly constant validation loss at 16x the training token length and outperforms Softmax and Softmax-free baselines on long-context retrieval and downstream benchmarks.

Load-bearing premise

The claim stands on the assumption that the flat 16x-length loss comes from the new attention re-weighting rather than from the test-time position-embedding stretch (Dynamic NTK Scaling) applied to all models, or from choosing the sharpening strength $p=15$ after seeing the results.

Editorial extensions

If this is right

  • Trained on 1024-token contexts, LSSAR holds its validation loss nearly flat out to 16,384 tokens, so a small model can be deployed on much longer documents without retraining.
  • Passkey retrieval stays non-zero at 8K tokens under LSSAR, while Softmax attention drops to 0% beyond 1.5K tokens, showing the sharpening stage prevents the model from losing a single critical token in long contexts.
  • The re-weighting operation improves extrapolation for other normalised attention variants at $p=3$, indicating the sharpening stage can be added to existing attention designs rather than requiring a full architecture change.
  • Because the improvement is inside the attention computation, LSSAR combines with RoPE and Dynamic NTK Scaling, and the authors argue it should transfer to larger RoPE-based models by analogy with deeper networks.
  • Softplus's bounded derivative keeps training stable at large $p$ values, where Softmax-based re-weighting suffers gradient explosion, so the sharpening strength can be increased without numerical failure.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial extension: the sharpening stage appears transferable, since the paper applies it to Softmax, Sigmoid, and LSSA attention; a natural untested step is fine-tuning an already-pretrained model with re-weighting instead of training from scratch.
  • Editorial extension: the abstract's symbolic-regression claim about recovering Newton's gravitational law is not developed in the provided experimental section, so it should be read as a secondary, unverified assertion rather than part of the length-extrapolation evidence.
  • Editorial extension: the paper's Discussion supports larger models only by analogy between large $p$ and increased depth, so scaling to billions of parameters remains an extrapolation rather than a demonstrated result.
  • Editorial extension: a principled choice of $p$ from the desired attention entropy or the score distribution's tail could remove the empirical sweep and is not attempted; a length-dependent $p$ schedule may push the flat-loss region beyond 16x.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The paper proposes a two-stage redesign of self-attention. The first stage (LSSA) decomposes Softmax into a positive nonlinearity plus l1-normalization, replaces the exponential with Softplus, and applies a scale factor log d log N. The second stage (LSSAR) sharpens the normalized attention distribution through a power-based re-weighting operation. The authors train GPT-2-small (124M) on 1024-token sequences and report that LSSAR maintains a nearly constant validation loss up to 16x the training length, outperforms several Softmax-free attention variants on length extrapolation, succeeds on passkey retrieval, and improves downstream benchmarks. The abstract also claims, without any supporting experiments in the body, that the method recovers Newton's gravitational law via symbolic regression.

Significance. If the extrapolation claims are correct, the two-stage decomposition offers a useful design principle for attention and a potential alternative to positional-interpolation fixes. The paper contributes a clean decomposition experiment (Table 2), a broad comparison against ReLU- and Sigmoid-based attentions under shared settings (Table 4), and a release of code, all of which are valuable. However, the central headline result is currently contingent on two unestablished choices: Dynamic NTK Scaling is applied to every model at inference and is never ablated, and the sharpening exponent p=15 is selected from the very extrapolation lengths the paper aims to predict. The significance is therefore conditional on additional experiments that separate the attention mechanism from these test-time and tuning effects.

major comments (5)
  1. [Experiments, first paragraph; Introduction paragraph 2] Dynamic NTK Scaling is incorporated into every model at inference and is never ablated. Dynamic NTK Scaling is itself a post-hoc RoPE-based positional embedding adjustment, the same class of technique the Introduction says LSSAR renders unnecessary. Consequently, the claim that LSSAR 'fundamentally improves length extrapolation' through its two-stage attention mechanism is not supported: the gains could arise from the interaction between the re-weighting and the unablated NTK scaling. Please provide an ablation of LSSAR (and, for completeness, Softmax) with and without Dynamic NTK Scaling, and report the extrapolation losses in all four configurations.
  2. [Ablation Study for Re-weighting Mechanism; Table 4; Fig. 2] The parameter p=15 is selected from the extrapolation lengths themselves. The ablation sweeps p over 1-15, 50, and 100 using validation loss at 1K-16K, and the paper then reports LSSAR(p=15) as the headline result. Table 4 shows that at p=3, which has the best training-length loss (1K: 3.1782 vs 3.1905 for p=15), extrapolation fails: 4K loss is 5.4056 and 8K loss is 6.3007, worse than both unsharpened LSSA (5.9403 at 8K) and Softmax (6.2823 at 8K). Equation (6) only establishes behavior as p approaches infinity and offers no principle for choosing p=15. The near-constant-loss claim is therefore a result of tuning p on the evaluation lengths. Please fix p a priori using a validation split that does not overlap the extrapolation lengths, or derive a principled selection criterion for p.
  3. [Abstract] The abstract states that 'symbolic regression experiments demonstrate that our method enables models to recover Newton's gravitational law from orbital trajectory sequences.' No symbolic regression experiments, orbital trajectory data, or Newton's-law results appear anywhere in the manuscript. This is a claimed contribution that is entirely missing from the body. Either add the experiments and their results, or remove the claim from the abstract.
  4. [Experiments; Tables 2-4; Fig. 2] No confidence intervals or repeated seeds are reported. All loss values appear to come from single runs. The central 'nearly constant' validation loss of LSSAR(p=15) (3.1905, 3.1930, 3.2291, 3.3171 at 1K, 2K, 4K, 8K) may be within ordinary run-to-run variability for a 124M-parameter model. Please report means and standard deviations over at least 3 seeds and, if possible, a paired comparison that tests whether the loss increase from 1K to 8K is statistically distinguishable from that of the baselines.
  5. [Length Scaled Softplus Attention, Eq. (4)] The scale factor log d log N is introduced with the statement that it 'ensures entropy invariance' across sequence lengths, but no derivation or citation for this specific combination is provided. As written, the factor is an additional free parameter of LSSA. Please either supply a derivation showing why log d log N preserves entropy invariance, or explicitly present it as an empirical design choice and ablate its contribution to the extrapolation results.
minor comments (5)
  1. [Attention Re-weighting Mechanism, Eq. (6)] Equation (6) appears to misstate the quantity being computed: the left-hand side 'xm - xl' is the original score difference, but the limit expression is the difference between the re-weighted values after the power transformation. The intended statement is that the re-weighted distance approaches 1 as p goes to infinity, not that the original score difference equals 1. Please correct the notation.
  2. [Downstream Evaluation, Table 5] The SummScreen ROUGE-1 scores of 1.682 and 6.309 are far below typical ROUGE-1 values for summarization benchmarks. Please clarify the metric configuration or verify that the reported numbers are not the result of an evaluation error.
  3. [Comparison with State-of-the-Art Softmax-Free Attention Functions] The text repeatedly refers to 'state-of-the-art Softmax-free alternatives,' but the comparison set contains only three ReLU-based and two Sigmoid-based attention variants, not the full range of recent Softmax-free attention methods. Consider softening the characterization or expanding the baseline list.
  4. [Fig. 2] The x-axis labels '1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 50100' are cramped and difficult to read. Use a log scale or split the axis to show the p=50 and p=100 points clearly.
  5. [Ablation Study for Re-weighting Mechanism] The text says 'optimal results observed around p = 15,' but the sweep includes p=50 and p=100 with noticeably higher loss, so 'around p=15' should be clarified to mean, for example, p between 10 and 20, rather than implying a plateau that extends to p=100.

Circularity Check

1 steps flagged · score 6.0 of 10

The near-constant extrapolation loss is reported after p=15 is selected from the same extrapolation lengths in Fig. 2; at p=3 LSSAR fails to extrapolate, so the central claim is a tuned result rather than an independent prediction.

  1. fitted input called prediction [Ablation Study for Re-weighting Mechanism (Fig. 2 and Table 4)]
    "increasing p values generally improve the performance of LSSA across different sequence lengths, with optimal results observed around p = 15. The validation loss remains relatively stable as the sequence length increases, underscoring LSSA’s strong scalability. ... We compared LSSAR (p = 15) against the standard Softmax attention baseline."

    The central claim that LSSAR maintains nearly constant validation loss at 16x the training length is not an independent prediction: the sharpening exponent p is a hyperparameter in Eq. (5), and p=15 is chosen as optimal on the very sequence lengths (1K to 16K) used to report the near-constant loss. The paper's own Table 4 shows that with p=3, LSSAR loses extrapolation stability (loss 5.4056 at 4K and 6.3007 at 8K), so the flat-loss behavior is a property of the tuned configuration, not of the two-stage attention mechanism per se. Eq. (6) only characterizes p approaching infinity and provides no principled reason to select p=15. Reporting the curve after optimizing p on that same curve converts a tuned design into the appearance of a predicted extrapolation result.

full rationale

The comparison against Softmax, Sigmoid, and ReLU baselines under shared training settings is genuine evidence and not tautological, and there is no load-bearing self-citation chain: the only self-citation (Gao and Pavel 2017) supports background material on softmax properties. The circular component is confined to the headline extrapolation result: p is optimized on the extrapolation lengths in Fig. 2, and the near-constant-loss claim is then reported for that same p=15 configuration. Since p=3, which is better at the training length, fails to extrapolate, the result is best described as a tuned design rather than a prediction from first principles. Additionally, all models use Dynamic NTK Scaling at inference, a post-hoc RoPE extrapolation fix that is never ablated, so the paper's inference that the attention mechanism alone 'fundamentally' improves extrapolation is not isolated from this positional-embedding adjustment. The passkey and downstream results at p=15 are not fitted to those particular tasks and retain some independent content, but the central length-extrapolation claim itself partially reduces to a hyperparameter selection on the reported curve, yielding a circularity score of 6 rather than a full 8 or 10.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The central claims rest on hand-chosen scaling, a tuned re-weighting exponent, and assumptions about what causes extrapolation failure. There is no fully derived theory linking softplus normalization and sharpening to length extrapolation.

free parameters (2)
  • Re-weighting exponent p = 15 (main results); 3, 50, 100 also tested
    Chosen after sweeping p in Fig 2; the near-constant loss claim at 2K-16K depends on p=15.
  • Length scale factor log d log N = log 64 * log N (with d=64)
    Chosen by hand from an entropy-invariance heuristic; no derivation is supplied and the central LSSA formulation depends on it.
assumptions (4)
  • domain assumption The crucial component of softmax is l1-normalisation, not positivity or non-negativity.
    Supported only by small loss differences in Tables 1 and 2 (e.g., 3.1911 vs 3.1954) without significance testing; used to justify replacing exp with softplus.
  • domain assumption Attention smoothing is the cause of length extrapolation failure and sharpening fixes it.
    A hypothesis; Eq. (6) only shows power sharpening converges to the maximum, not that this produces extrapolation.
  • domain assumption Dynamic NTK Scaling is neutral or orthogonal to the attention comparison.
    Applied to all models at inference; no ablation without it, so the attention-only contribution is not isolated.
  • ad hoc to paper The log d log N scale factor preserves entropy invariance across sequence lengths.
    Presented as an extension of Su 2021 and Chiang-Cholak 2022, but no derivation is given for the row-dependent form.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Softplus Attention with Re-weighting Boosts Length Extrapolation in Large Language Models." pith.science (2026). https://pith.science/paper/4FWCNJDN

@misc{pith2026250113428,
  author       = {Pith},
  title        = {Pith review of: Softplus Attention with Re-weighting Boosts Length Extrapolation in Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4FWCNJDN}},
  note         = {Machine review of arXiv:2501.13428}
}
abstract

Large language models have achieved remarkable success in recent years, primarily due to self-attention. However, traditional Softmax attention suffers from numerical instability and reduced performance as the number of inference tokens increases. This work addresses these issues by proposing a new design principle for attention, viewing it as a two-stage process. The first stage (normalisation) refines standard attention by replacing Softmax with the more numerically stable Softplus followed by $l_{1}$-normalisation. Furthermore, we introduce a dynamic scale factor based on invariance entropy. We show that this novel attention mechanism outperforms conventional Softmax attention, and state-of-the-art Softmax-free alternatives. Our second proposal is to introduce a second processing stage (sharpening) which consists of a re-weighting mechanism that amplifies significant attentional weights while diminishing weaker ones. This enables the model to concentrate more effectively on relevant tokens, mitigating the attention sink phenomenon, and fundamentally improving length extrapolation. This novel, two-stage, replacement for self-attention is shown to ensure numerical stability and dramatically improve length extrapolation, maintaining a nearly constant validation loss at 16$\times$ the training length while achieving superior results on challenging long-context retrieval tasks and downstream benchmarks. Furthermore, symbolic regression experiments demonstrate that our method enables models to recover Newton's gravitational law from orbital trajectory sequences, providing evidence that appropriate attention mechanisms are crucial for foundation models to develop genuine physical world models. Our code is available at https://github.com/iminfine/freeattn.

Figures

Figures reproduced from arXiv: 2501.13428 by the authors.

Figure 1
Figure 1. Comparison of Softmax attention and the proposed [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 3
Figure 3. Passkey retrieval accuracy for LSSAR (p = 15) and standard Softmax attention. Accuracy is averaged over 100 trials with the passkey placed at random positions within the sequence. Long Context Passkey Retrieval To offer a more direct and challenging evaluation of length extrapolation, we employed the long-context passkey re￾trieval task (Mohtashami and Jaggi 2023). This ”needle-in￾a-haystack” test is specifically de… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

69 extracted references · 40 canonical work pages

  1. [1]

    , " * write output.state after.block = add.period write newline

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al

    Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023. GPT -4 technical report. arXiv:2303.08774

  4. [4]

    Arora, S.; Eyuboglu, S.; Zhang, M.; Timalsina, A.; Alberti, S.; Zinsley, D.; Zou, J.; Rudra, A.; and R \'e , C. 2024. Simple linear attention language models balance the recall-throughput tradeoff. arXiv:2402.18668

  5. [5]

    R.; et al

    AUEB, T. R.; et al. 2016. One-vs-each approximation to softmax for scalable estimation of probabilities. Advances in Neural Information Processing Systems, 29

  6. [6]

    R.; and Hinton, G

    Ba, J.; Kiros, J. R.; and Hinton, G. E. 2016. Layer normalization. In Proceedings of the 33rd International Conference on Machine Learning (ICML), 198--206. PMLR

  7. [7]

    Bai, Y.; Chen, F.; Wang, H.; Xiong, C.; and Mei, S. 2024. Transformers as statisticians: Provable in-context learning with in-context algorithm selection. Advances in neural information processing systems, 36

  8. [8]

    Bisk, Y.; Zellers, R.; Gao, J.; Choi, Y.; et al. 2020. Piqa: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI conference on artificial intelligence, volume 34, 7432--7439

Show all 69 references
  1. [9]

    by parts

    bloc97. 2023 a . Add NTK -Aware interpolation "by parts" correction. GitHub Pull Request

  2. [10]

    bloc97. 2023 b . NTK -Aware Scaled RoPE allows LLaMA models to have extended (8k+) context size without any fine-tuning and minimal perplexity degradation. Reddit post

  3. [11]

    Chen, M.; Chu, Z.; Wiseman, S.; and Gimpel, K. 2021. SummScreen: A dataset for abstractive screenplay summarization. arXiv:2104.07091

  4. [12]

    Chen, S.; Wong, S.; Chen, L.; and Tian, Y. 2023. Extending Context Window of Large Language Models via Positional Interpolation

  5. [13]

    J.; and Rudnicky, A

    Chi, T.-C.; Fan, T.-H.; Ramadge, P. J.; and Rudnicky, A. 2022. KERPLE: Kernelized Relative Positional Embedding for Length Extrapolation. In Advances in Neural Information Processing Systems, volume 35, 8386--8399

  6. [14]

    Chiang, D.; and Cholak, P. 2022. Overcoming a Theoretical Limitation of Self-Attention. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7654--7664

  7. [15]

    Clark, P.; Cowhey, I.; Etzioni, O.; Khot, T.; Sabharwal, A.; Schoenick, C.; and Tafjord, O. 2018. Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge. arXiv:1803.05457

  8. [16]

    DeepSeek-AI; Guo, D.; Yang, D.; Zhang, H.; Song, J.; Zhang, R.; Xu, R.; Zhu, Q.; Ma, S.; Wang, P.; Bi, X.; Zhang, X.; Yu, X.; Wu, Y.; Wu, Z. F.; Gou, Z.; Shao, Z.; Li, Z.; Gao, Z.; Liu, A.; Xue, B.; Wang, B.; Wu, B.; Feng, B.; Lu, C.; Zhao, C.; Deng, C.; Zhang, C.; Ruan, C.; D...

  9. [17]

    P.; Caron, M.; Geirhos, R.; Alabdul mohsin, I.; Jenatton, R.; Beyer, L.; Tschannen, M.; Arnab, A.; Wang, X.; Ruiz, C

    Dehghani, M.; Djolonga, J.; Mustafa, B.; Padlewski, P.; Heek, J.; Gilmer, J.; Steiner, A. P.; Caron, M.; Geirhos, R.; Alabdul mohsin, I.; Jenatton, R.; Beyer, L.; Tschannen, M.; Arnab, A.; Wang, X.; Ruiz, C. R.; Minderer, M.; Puigcerver, J.; Evci, U.; Kumar, M.; van Steenkiste...

  10. [18]

    Dubey, A.; Jauhri, A.; Pandey, A.; Kadian, A.; Al-Dahle, A.; Letman, A.; Mathur, A.; Schelten, A.; Yang, A.; Fan, A.; et al. 2024. The Llama 3 herd of models. arXiv:2407.21783

  11. [19]

    emozilla. 2023. Dynamically Scaled RoPE further increases performance of long context LLaMA with zero fine-tuning. Reddit post

  12. [20]

    Fu, H.; Guo, T.; Bai, Y.; and Mei, S. 2024. What can a single attention layer learn? a study through the random features lens. Advances in Neural Information Processing Systems, 36

  13. [21]

    Gao, B.; and Pavel, L. 2017. On the properties of the softmax function with application in game theory and reinforcement learning. arXiv:1704.00805

  14. [22]

    Gao, L.; Tow, J.; Abbasi, B.; Biderman, S.; Black, S.; DiPofi, A.; Foster, C.; Golding, L.; Hsu, J.; Le Noac'h, A.; Li, H.; McDonell, K.; Muennighoff, N.; Ociepa, C.; Phang, J.; Reynolds, L.; Schoelkopf, H.; Skowron, A.; Sutawika, L.; Tang, E.; Thite, A.; Wang, B.; Wang, K.; a...

  15. [23]

    Golovneva, O.; Wang, T.; Weston, J.; and Sukhbaatar, S. 2024. Contextual Position Encoding: Learning to Count What's Important. arXiv

  16. [24]

    Han, D.; Pan, X.; Han, Y.; Song, S.; and Huang, G. 2023. Flatten transformer: Vision transformer using focused linear attention. In Proceedings of the IEEE/CVF international conference on computer vision, 5961--5971

  17. [25]

    He, Z.; Feng, G.; Luo, S.; Yang, K.; He, D.; Xu, J.; Zhang, Z.; Yang, H.; and Wang, L. 2024. Two Stones Hit One Bird: Bilevel Positional Encoding for Better Length Extrapolation. arXiv

  18. [26]

    Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2020. Measuring massive multitask language understanding

  19. [27]

    Hendrycks, D.; and Gimpel, K. 2016. Gaussian Error Linear Units ( GELUs ). arXiv:1606.08415

  20. [28]

    R.; Pawar, S

    Henry, A.; Dachapally, P. R.; Pawar, S. S.; and Chen, Y. 2020. Query-key normalization for transformers. In Findings of the Association for Computational Linguistics: EMNLP 2020, 4246--4253. Association for Computational Linguistics

  21. [29]

    G.; Zhu, M.; Chen, B.; Kalenichenko, D.; Wang, W.; Weyand, T.; Andreetto, M.; and Adam, H

    Howard, A. G.; Zhu, M.; Chen, B.; Kalenichenko, D.; Wang, W.; Weyand, T.; Andreetto, M.; and Adam, H. 2017. MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVP...

  22. [30]

    Hron, J.; Bahri, Y.; Sohl-Dickstein, J.; and Novak, R. 2020. Infinite attention: NNGP and NTK for deep attention networks. In International Conference on Machine Learning, 4376--4386. PMLR

  23. [31]

    Hua, W.; Dai, Z.; Liu, H.; and Le, Q. 2022. Transformer quality in linear time. In International conference on machine learning, 9099--9117. PMLR

  24. [32]

    Huang, Z.; Liang, D.; Xu, P.; and Xiang, B. 2020. Improve Transformer Models with Better Relative Position Embeddings. In Findings of the Association for Computational Linguistics: EMNLP 2020, 3327--3335. Online: Association for Computational Linguistics

  25. [33]

    Jang, E.; Gu, S.; and Poole, B. 2016. Categorical reparameterization with gumbel-softmax. arXiv:1611.01144

  26. [34]

    kaiokendev. 2023. Things I'm learning while training superhot. Accessed: [Insert Access Date]

  27. [35]

    Katharopoulos, A.; Vyas, A.; Pappas, N.; and Fleuret, F. 2020. Transformers are RNN : Fast autoregressive transformers with linear attention. In International conference on machine learning, 5156--5165. PMLR

  28. [36]

    N.; Das, P.; and Reddy, S

    Kazemnejad, A.; Padhi, I.; Ramamurthy, K. N.; Das, P.; and Reddy, S. 2023. The Impact of Positional Encoding on Length Generalization in Transformers. arXiv:2305.19466

  29. [37]

    Kiyono, S.; Kobayashi, S.; Suzuki, J.; and Inui, K. 2021. SHAPE: Shifted Absolute Position Embedding for Transformers. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 3309--3321. Online and Punta Cana, Dominican Republic: Association ...

  30. [38]

    Lai, Z.; Lim, L.-H.; and Liu, Y. 2024. Attention is a smoothed cubic spline. arXiv:2408.09624

  31. [39]

    Li, S.; You, C.; Guruganesh, G.; Ainslie, J.; Ontanon, S.; Zaheer, M.; Sanghai, S.; Yang, Y.; Kumar, S.; and Bhojanapalli, S. 2023. Functional Interpolation for Relative Positions Improves Long Context Transformers

  32. [40]

    Li, Z.; Bhojanapalli, S.; Zaheer, M.; Reddi, S.; and Kumar, S. 2022. Robust training of neural networks using scale invariant architectures. In International Conference on Machine Learning, 12656--12684. PMLR

  33. [41]

    Likhomanenko, T.; Xu, Q.; Synnaeve, G.; Collobert, R.; and Rogozhnikov, A. 2021. CAPE : Encoding Relative Positions with Continuous Augmented Positional Embeddings. In Advances in Neural Information Processing Systems, volume 34, 16079--16092. Curran Associates, Inc

  34. [42]

    Liu, A.; Feng, B.; Xue, B.; Wang, B.; Wu, B.; Lu, C.; Zhao, C.; Deng, C.; Zhang, C.; Ruan, C.; et al. 2024. Deepseek-v3 technical report. arXiv:2412.19437

  35. [43]

    Liu, X.; Yan, H.; Zhang, S.; An, C.; Qiu, X.; and Lin, D. 2023. Scaling Laws of RoPE-based Extrapolation. arXiv

  36. [44]

    Liu, Z.; Hu, H.; Lin, Y.; Yao, Z.; Xie, Z.; Wei, Y.; Ning, J.; Cao, Y.; Zhang, Z.; Dong, L.; Wei, F.; and Guo, B. 2022. Swin Transformer V2: Scaling Up Capacity and Resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 11999--...

  37. [45]

    Lu, J.; Yao, J.; Zhang, J.; Zhu, X.; Xu, H.; Gao, W.; Xu, C.; Xiang, T.; and Zhang, L. 2021. Soft: Softmax-free transformer with linear complexity. Advances in Neural Information Processing Systems, 34: 21297--21309

  38. [46]

    Misra, D. 2019. Mish: A Self Regularized Non-Monotonic Neural Activation Function. arXiv:1908.08681

  39. [47]

    Mohtashami, A.; and Jaggi, M. 2023. Random-access infinite context length for transformers. Advances in Neural Information Processing Systems, 36: 54567--54585

  40. [48]

    B.; Lozhkov, A.; Mitchell, M.; Raffel, C.; Werra, L

    Penedo, G.; Kydlíček, H.; allal, L. B.; Lozhkov, A.; Mitchell, M.; Raffel, C.; Werra, L. V.; and Wolf, T. 2024. The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale. arXiv:2406.17557

  41. [49]

    Qi, X.; Ye, J.; He, Y.; Li, C.-G.; Zi, B.; Dai, X.; Zou, Q.; and Xiao, R. 2024. Stable-Transformer: Towards a Stable Transformer Training. Accessed: 2024-11-15

  42. [50]

    Qin, Z.; Sun, W.; Deng, H.; Li, D.; Wei, Y.; Lv, B.; Yan, J.; Kong, L.; and Zhong, Y. 2022. cosFormer : Rethinking Softmax In Attention. In International Conference on Learning Representations

  43. [51]

    Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; and Sutskever, I. 2019. Language Models are Unsupervised Multitask Learners

  44. [52]

    Ramapuram, J.; Danieli, F.; Dhekane, E.; Weers, F.; Busbridge, D.; Ablin, P.; Likhomanenko, T.; Digani, J.; Gu, Z.; Shidani, A.; et al. 2024. Theory, Analysis, and Best Practices for Sigmoid Self-Attention. arXiv:2409.04431

  45. [53]

    E.; Hinton, G

    Rumelhart, D. E.; Hinton, G. E.; and Williams, R. J. 1986. Learning Internal Representations by Error Propagation. Nature, 323(6088): 533--536

  46. [54]

    Shah, J.; Bikshandi, G.; Zhang, Y.; Thakkar, V.; Ramani, P.; and Dao, T. 2024. Flashattention-3: Fast and accurate attention with asynchrony and low-precision. arXiv:2407.08608

  47. [55]

    Shen, K.; Guo, J.; Tan, X.; Tang, S.; Wang, R.; and Bian, J. 2023. A study on ReLU and softmax in transformer. arXiv:2302.06461

  48. [56]

    So, D.; Ma \'n ke, W.; Liu, H.; Dai, Z.; Shazeer, N.; and Le, Q. V. 2021. Searching for efficient transformers for language modeling. Advances in neural information processing systems, 34: 6010--6022

  49. [57]

    Su, J. 2021. Viewing the scale operation of attention from the perspective of entropy invariance

  50. [58]

    Su, J.; Ahmed, M.; Lu, Y.; Pan, S.; Bo, W.; and Liu, Y. 2024. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 568: 127063

  51. [59]

    H.; Bai, S.; Yamada, M.; Morency, L.-P.; and Salakhutdinov, R

    Tsai, Y.-H. H.; Bai, S.; Yamada, M.; Morency, L.-P.; and Salakhutdinov, R. 2019. Transformer Dissection: An Unified Understanding for Transformer's Attention via the Lens of Kernel. In Proceedings of the Conference on Empirical Methods in Natural Language Processing

  52. [60]

    Veli c kovi \'c , P.; Perivolaropoulos, C.; Barbero, F.; and Pascanu, R. 2024. Softmax is not Enough (for Sharp Size Generalisation). arXiv:2410.01104

  53. [61]

    Wang, S.; Kobyzev, I.; Lu, P.; Rezagholizadeh, M.; and Liu, B. 2024. Resonance RoPE : Improving Context Length Generalization of Large Language Models. arXiv

  54. [62]

    F.; and Gardner, M

    Welbl, J.; Liu, N. F.; and Gardner, M. 2017. Crowdsourcing multiple choice science questions. arXiv:1707.06209

  55. [63]

    Wortsman, M.; Lee, J.; Gilmer, J.; and Kornblith, S. 2023. Replacing softmax with ReLU in vision transformers. arXiv:2309.08586

  56. [64]

    Wu, M.; Cheng, X.; Padon, O.; and Jia, Z. 2024. A Multi-Level Superoptimizer for Tensor Programs. arXiv:2405.05751

  57. [65]

    Yang, A.; Yang, B.; Hui, B.; Zheng, B.; Yu, B.; Zhou, C.; Li, C.; Li, C.; Liu, D.; Huang, F.; et al. 2024. Qwen2 technical report. arXiv:2407.10671

  58. [66]

    Zellers, R.; Holtzman, A.; Bisk, Y.; Farhadi, A.; and Choi, Y. 2019. Hellaswag: Can a machine really finish your sentence? arXiv:1905.07830

  59. [67]

    Zhang, Y.; Liu, Y.; Yuan, H.; Qin, Z.; Yuan, Y.; Gu, Q.; and Yao, A. C.-C. 2025. Tensor Product Attention Is All You Need. arXiv:2501.06425

  60. [68]

    Zheng, C.; Gao, Y.; Shi, H.; Huang, M.; Li, J.; Xiong, J.; Ren, X.; Ng, M.; Jiang, X.; Li, Z.; and Li, Y. 2024. CAPE : Context-Adaptive Positional Encoding for Length Extrapolation. arXiv

  61. [69]

    Zheng, H.; Yang, Z.; Liu, W.; Liang, J.; and Li, Y. 2015. Improving deep neural networks using softplus units. In 2015 International joint conference on neural networks (IJCNN), 1--4. IEEE

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.