Pith. sign in

Paper Citation Record · LEDGER

Scalable-Softmax Is Superior for Attention

As of 10 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 16 inbound Pith citation observations for arXiv:2501.19399.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.19399 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T20:34:01.273009Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:41:54.830690Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T17:09:59.046151Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2ed10e6f-91e7-4e9f-ae4b-3f33a355d427 · outbound

This paper cites write newline.

Scalable-Softmax Is Superior for Attention write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.134673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.134673Z digest=sha256:79879c98d36d912a814fcb15422a9d10ce2444caaa7bf9a4028bc3761adf2867

Observation 4b958027-fc05-4726-83fd-f03e1142615a · outbound

This paper cites Etc: Encoding long and structured inputs in transformers.

Scalable-Softmax Is Superior for Attention Etc: Encoding long and structured inputs in transformers

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.721755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.140337Z digest=sha256:b47fd0f76770a9c527b14b6cdcea02a9b242af27e9c282c8c3566a56b2358fe3

Observation ba82097c-5b3d-40ee-80d0-9979505729d3 · outbound

This paper cites Needle in a haystack - pressure testing llms, 2023.

Scalable-Softmax Is Superior for Attention Needle in a haystack - pressure testing llms, 2023

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.709283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.144584Z digest=sha256:47168b03b15bc4a1f5500d20fb27638e2862a4985db7a829e6710b6dd54695ef

Observation 92477d2e-ba56-47c1-9b9b-52b202cd2518 · outbound

This paper cites Longformer: The Long-Document Transformer.

Scalable-Softmax Is Superior for Attention Longformer: The Long-Document Transformer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.149151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.149151Z digest=sha256:19637d0d7c217ca79d01ae962cbf4ed2138dd47becd5f24b1556759d45b7a9d3

Observation af892ac0-26ae-4e92-9a58-f6922ec37907 · outbound

This paper cites an unresolved cited work.

Scalable-Softmax Is Superior for Attention Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-09T20:34:01.696917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.155036Z digest=sha256:12f5fcf758803cb269fb31ae7b1d539e080ca2d26be3b621a80c282068fa0447

Observation b27334a6-9fb7-4152-b4ed-c1dc68e81729 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Scalable-Softmax Is Superior for Attention Generating Long Sequences with Sparse Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.160153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.160153Z digest=sha256:980cc92498a857e895530a37fc8470e5c76148fd973acb1a2c24d7657acd8026

Observation e666f589-e24b-4c15-9bc4-2e326da9c711 · outbound

This paper cites Redpajama: An open source recipe to reproduce llama training dataset, April 2023.

Scalable-Softmax Is Superior for Attention Redpajama: An open source recipe to reproduce llama training dataset, April 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.684935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.165155Z digest=sha256:cb44f397c8b48ac39b189db61c7a68c825fdae695e484133858145af4def2428

Observation 2e378df9-9919-4659-9baa-865950801315 · outbound

This paper cites GMAT: Global Memory Augmentation for Transformers.

Scalable-Softmax Is Superior for Attention GMAT: Global Memory Augmentation for Transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.170481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.170481Z digest=sha256:eb563481913239b4d8241f25eee4319e1ece4bb3d15a02c06efa8cb1a067280f

Observation 5084afc2-b111-4f59-a5a8-6eeef06233c2 · outbound

This paper cites Needle in a haystack - pressure testing llms, 2023.

Scalable-Softmax Is Superior for Attention Needle in a haystack - pressure testing llms, 2023

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.670341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.174657Z digest=sha256:c5b5f1abe21a5f4ad45377d3d404b2663a4dfdc9213f0e6c728c82ad3d97a8f5

Observation 55b2139c-c9c9-4bba-9c9d-e91538b5d273 · outbound

This paper cites The impact of positional encoding on length generalization in transformers.

Scalable-Softmax Is Superior for Attention The impact of positional encoding on length generalization in transformers

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.656333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.178471Z digest=sha256:36a813c0131cf2e064bf0c1a8e5e474acf7aef6f1fe760ebc2efdcb8e76e446e

Observation 42a47d50-98d9-4644-ae9a-52cec5fac23c · outbound

This paper cites Reformer: The Efficient Transformer.

Scalable-Softmax Is Superior for Attention Reformer: The Efficient Transformer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.182826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.182826Z digest=sha256:f5f20d6f61bba1c2a89fe7ddc200f823bc9c47a1e1a0e0e7323ff4c9807b2939

Observation dbe4dab1-c88c-4c3c-9d77-8444c61fdecd · outbound

This paper cites Gradient-based learning applied to document recognition.

Scalable-Softmax Is Superior for Attention Gradient-based learning applied to document recognition

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.187475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.187475Z digest=sha256:a291d52919d1bbc6a78cd50b6e7f704b1b66e5b0f22521f3c07bbcbb7c25c25d

Observation 088bc4d1-82ce-41e8-9d7f-dff5b338909a · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

Scalable-Softmax Is Superior for Attention World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.191643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.191643Z digest=sha256:6b4238418d002f271b6d25eab0d8ec25b1be67896b936031097a8d5e77158e87

Observation 7fabc220-9369-4818-93e1-06316fc0aade · outbound

This paper cites Scaling laws of ro PE -based extrapolation.

Scalable-Softmax Is Superior for Attention Scaling laws of ro PE -based extrapolation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.632605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.196011Z digest=sha256:aeb867636dad1fa1774189a1c11bd0d2ee08264bca0c8c92a45b86fe5980c9e3

Observation c0c7beda-2a5e-4b4d-8d56-43c9b9386cb7 · outbound

This paper cites and Hutter, F.

Scalable-Softmax Is Superior for Attention and Hutter, F

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.620244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.200339Z digest=sha256:e10cd8d2fa2fce4f7360945fde2a484137ebbda208167a906b42b37bea003d08

Observation 80e93c09-6eeb-4279-a265-fa4c0baf0979 · outbound

This paper cites Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation.

Scalable-Softmax Is Superior for Attention Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.204290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.204290Z digest=sha256:34c1bd2ba77d7cbad4af5bd6d4cdb97194decb46a9462f9990916d2a9cfcf5cb

Observation c07752c3-592f-4efa-a1bc-8f42f540947f · outbound

This paper cites Language models are unsupervised multitask learners.

Scalable-Softmax Is Superior for Attention Language models are unsupervised multitask learners

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.608076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.208531Z digest=sha256:c438e0f076fe03db9a7d835728e6403b2ed6f1b6f4257676454e6b7fdf3624ee

Observation 672dcff2-cfd2-4416-9f5c-87c364774549 · outbound

This paper cites SQ u AD : 100,000+ questions for machine comprehension of text.

Scalable-Softmax Is Superior for Attention SQ u AD : 100,000+ questions for machine comprehension of text

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.594186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.212024Z digest=sha256:d79b76abc08eac850847cc16d74b3b05ea42ed3423d05be67dcd37b84b783358

Observation 1a5ae387-fcab-438b-8b31-4ce5dcee1a3e · outbound

This paper cites Know what you don't know: Unanswerable questions for SQ u AD.

Scalable-Softmax Is Superior for Attention Know what you don't know: Unanswerable questions for SQ u AD

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.580066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.216919Z digest=sha256:1da88472c2b66cb8cb4b8040c88b92c0bf2e5111d51433b4ba63f2dfea79ee25

Observation ae3452d9-32d1-47c4-827d-ca61e191e457 · outbound

This paper cites Searching for Activation Functions.

Scalable-Softmax Is Superior for Attention Searching for Activation Functions

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.220920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.220920Z digest=sha256:81b1868e619ee920e1ba2b5bf613f44f224fe1e26e042e692d8991ea4a1bd8b2

Observation 03b547a5-d657-48d4-a3c3-832377568d2d · outbound

This paper cites Efficient content-based sparse attention with routing transformers.

Scalable-Softmax Is Superior for Attention Efficient content-based sparse attention with routing transformers

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.565825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.225849Z digest=sha256:f2a385850aae9c10727765fdf8b0797da8d231e7984a60c9cd492e04eff9680c

Observation 22a26b0f-8430-46a5-aed1-e915167855ad · outbound

This paper cites Self-attention with relative position representations.

Scalable-Softmax Is Superior for Attention Self-attention with relative position representations

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.552563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.229542Z digest=sha256:fdf650865f81b3ac36ceea0c73dc20707d0a4d5edb146c7d3ec67b3e6229c6ef

Observation fc8e0108-2bb8-48fb-9c18-460b87a6c619 · outbound

This paper cites GLU Variants Improve Transformer.

Scalable-Softmax Is Superior for Attention GLU Variants Improve Transformer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.233103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.233103Z digest=sha256:abe2bf937bc395cb25dcb22be1c894592558d346a926f43d05bd3ac5bc3f42b8

Observation af07bdd1-1197-44fe-b133-9f926d5b3124 · outbound

This paper cites R., Hestness, J., and Dey, N.

Scalable-Softmax Is Superior for Attention R., Hestness, J., and Dey, N

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.537092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.237096Z digest=sha256:0fc14ad0d10f26e60532f1823c93e5436ba697d925e39ec6da5c31bc03ec6707

Observation 7896d758-26e6-412c-820a-1a804a63671a · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

Scalable-Softmax Is Superior for Attention Roformer: Enhanced transformer with rotary position embedding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.240997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.240997Z digest=sha256:1dcb67a715744e383461f0c45c0327221b237e4b63f660ebab185764f98aba2e

Observation f87f8846-8bc2-4896-8e5a-4a95b40b0be6 · outbound

This paper cites Adaptive attention span in transformers.

Scalable-Softmax Is Superior for Attention Adaptive attention span in transformers

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.515253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.244520Z digest=sha256:0c86e5d97ae3e279aa16b0d8873ec305a1259a2a2f9ace203703bd214838516c

Observation e7d417f5-166c-4484-b821-73942fafef59 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Scalable-Softmax Is Superior for Attention Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.249086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.249086Z digest=sha256:3e72d10ce5ac5e77e11c9e4edf418dd86cef34cc90ee1254aa05ad8e893d70fe

Observation d767d1a9-4b1d-47ec-9edc-964290771642 · outbound

This paper cites N., Kaiser, ., and Polosukhin, I.

Scalable-Softmax Is Superior for Attention N., Kaiser, ., and Polosukhin, I

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.502711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.253366Z digest=sha256:9847b1d7d331fac5edc1342c1939d6badf1a26f3f87aa68f92934e2c3b36f257

Observation c72e8e0f-57f3-4f34-b4ce-a6db9d23ebcd · outbound

This paper cites Length Generalization of Causal Transformers without Position Encoding.

Scalable-Softmax Is Superior for Attention Length Generalization of Causal Transformers without Position Encoding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.257094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.257094Z digest=sha256:184a2f41fc5104ba12bfce6861092cd601c20738be17336e834653d8bfa8df21

Observation 8bbe99af-eec7-405e-b837-61811ca41f61 · outbound

This paper cites V., and Zhou, D.

Scalable-Softmax Is Superior for Attention V., and Zhou, D

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.490700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.261120Z digest=sha256:b3b93fe06b2aabc31c173e365073163d2b428a78197c8d00bfd0b8a278c79db4

Observation 7cb6b1d2-436f-4fab-b038-103e2a046339 · outbound

This paper cites Differential Transformer.

Scalable-Softmax Is Superior for Attention Differential Transformer

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.265036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.265036Z digest=sha256:ef2bbe5fec35e469e33d1a8f10f448be5b04e6e8ad517223487c49c0ba411ec6

Observation 451730d9-5fce-48c8-bec8-057ca56429ac · outbound

This paper cites A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., and Ahmed, A.

Scalable-Softmax Is Superior for Attention A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., and Ahmed, A

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.473502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.269080Z digest=sha256:6272d6e56e80382e0c319627480c7749a19c6515b669856ed12384f13258f2d7

Observation 8e252d1a-4a8f-4e76-9075-1c1ec35fb147 · outbound

This paper cites and Sennrich, R.

Scalable-Softmax Is Superior for Attention and Sennrich, R

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.273009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.273009Z digest=sha256:bf091b20bdedb3e797bbd21cc7fc97d046334be61f1eac9bbe31381be01c4b95

Pith citing papers

Observation 6280ffc3-a12f-40fe-9030-78aab512c9ec · inbound

On the Mathematical Impossibility of Safe Universal Approximators cites this paper.

On the Mathematical Impossibility of Safe Universal Approximators Scalable-Softmax Is Superior for Attention

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T20:41:54.830690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:41:54.830690Z digest=sha256:a20004086e45edb79bea2a7563e72b77d3ad6ba6e1634eaa55920307cf2a02bc

Observation ecce2904-108a-4489-bac0-607d95f91640 · inbound

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs cites this paper.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Scalable-Softmax Is Superior for Attention

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:55.007461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:55.007461Z digest=sha256:dd02a7e0dfb615b734b89ef0e1da6ad003f54f309518454e4f3a92bab8d3d35a

Observation e8ec1c17-07c2-4575-b098-96245eb944d3 · inbound

Critical attention scaling in long-context transformers cites this paper.

Critical attention scaling in long-context transformers Scalable-Softmax Is Superior for Attention

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:28.867756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:28.867756Z digest=sha256:1a887523e02b538798ab71a562f35c769290df362c8e13ce58bd8cae63f057df

Observation 6ad89e38-fc74-4506-91b7-74b7c349c60d · inbound

Ministral 3 cites this paper.

Ministral 3 Scalable-Softmax Is Superior for Attention

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:12:24.763656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T19:12:24.627033Z digest=sha256:b89d5d882455b7aff414ba953b7c99673ee888f6dc22471f24926a0cd0bdb32c

Observation 547c0705-9c85-459c-bb4d-3885f7f86b6c · inbound

Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language Modeling cites this paper.

Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language Modeling Scalable-Softmax Is Superior for Attention

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-03T09:57:48.435605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:57:48.435605Z digest=sha256:92b2d59e9157f88f3dd6799bd623506fdc1886027e125a421589ae51c4de2c52

Observation febb80e5-4542-49f1-b831-16c14a806b94 · inbound

MemDLM: Memory-Enhanced DLM Training cites this paper.

MemDLM: Memory-Enhanced DLM Training Scalable-Softmax Is Superior for Attention

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:43:24.547738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T00:42:22.989588Z digest=sha256:5d847f9c9009fb4d9c1ba8168bbcecdb6c9b28669da210a607f00f775639fa0a

Observation 8214c335-6e25-4887-94e1-bca7aca29e4e · inbound

Screening Is Enough cites this paper.

Screening Is Enough Scalable-Softmax Is Superior for Attention

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:18:20.924558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T22:17:13.803778Z digest=sha256:1dd744046a8f725821c8a223f9a28d6169ebab27e2b2d80767e9e5a3eb22d540

Observation 7e276543-d219-40ab-870a-e97c1ee32e87 · inbound

ExtraVAR: Stage-Aware RoPE Remapping for Resolution Extrapolation in Visual Autoregressive Models cites this paper.

ExtraVAR: Stage-Aware RoPE Remapping for Resolution Extrapolation in Visual Autoregressive Models Scalable-Softmax Is Superior for Attention

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:26:28.046243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T03:01:32.762236Z digest=sha256:487ad129c07f76566de94127f023c8a886ab96319cc2b284fa6cb2ce45da65d6

Observation 3a902776-13d8-420a-9934-45d5a0eca3c9 · inbound

Phoenix-VL 1.5 Medium Technical Report cites this paper.

Phoenix-VL 1.5 Medium Technical Report Scalable-Softmax Is Superior for Attention

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.890914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:41:27.144815Z digest=sha256:dc5dd3077a5ebc18016fd3b30dec59dae578921dd3fe5bbc71ccee4934063d8d

Observation 87dc7db5-a504-499b-8cda-600fe95b4f51 · inbound

A Unified Framework for Critical Scaling of Inverse Temperature in Self-Attention cites this paper.

A Unified Framework for Critical Scaling of Inverse Temperature in Self-Attention Scalable-Softmax Is Superior for Attention

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:57:53.542250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-14T19:56:02.462786Z digest=sha256:92bbc66c07238bc7ddaa407849e2ae7b12fe5ea3a184bf41f66c93011e2a7394

Observation 059d601f-3e12-4ecd-b912-7690e1872ae0 · inbound

TabPFN-3: Technical Report cites this paper.

TabPFN-3: Technical Report Scalable-Softmax Is Superior for Attention

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:10:06.047420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T06:05:16.199413Z digest=sha256:a936657d1be9169f0a818196899790adb5ae05386e68d07b55b479a5497219be

Observation 45c62d5c-a0b1-49e2-907c-bd68b278b9b6 · inbound

TabPFN-3: Technical Report cites this paper.

TabPFN-3: Technical Report Scalable-Softmax Is Superior for Attention

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:25:46.516167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T21:37:08.627791Z digest=sha256:1fd564571485c7a3e2aa30742a99c69056da7e378a61f49b20f65272051f9f2a

Observation 48068909-0dd8-46d2-a385-e779d598b381 · inbound

Most Transformer Modifications Still Do Not Transfer at 1-3B: A 2020-2026 Update to Narang et al. (2021) with Downstream Evaluation and a Noise Floor cites this paper.

Most Transformer Modifications Still Do Not Transfer at 1-3B: A 2020-2026 Update to Narang et al. (2021) with Downstream Evaluation and a Noise Floor Scalable-Softmax Is Superior for Attention

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:19:42.039766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-21T06:15:47.451870Z digest=sha256:3118b05c8be260250cbb529948bb6e5c2e5432e173360e693d4f1864994f23e9

Observation d54cc909-4148-41a9-af19-ec29a218b84a · inbound

ATMA: Length-Invariant Language Modeling via Polar Attention and Gated-Delta Compression Memory cites this paper.

ATMA: Length-Invariant Language Modeling via Polar Attention and Gated-Delta Compression Memory Scalable-Softmax Is Superior for Attention

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:09:59.047893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T23:52:52.985532Z digest=sha256:0e03593f7cb4c119497dff286e4a5a306ebb7f09254ccb91cb7682ed895bd2ef

Observation 5924af4a-ecc5-4c0c-a99f-4170508d4130 · inbound

ATMA: Length-Invariant Language Modeling via Polar Attention and Gated-Delta Compression Memory cites this paper.

ATMA: Length-Invariant Language Modeling via Polar Attention and Gated-Delta Compression Memory Scalable-Softmax Is Superior for Attention

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:44:37.730997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T09:38:43.250420Z digest=sha256:2924cb70ac365e1ce3641e5dd820814087ced51934ce3df058056bca143ad6d7

Observation 91be52e2-032d-4042-a936-f8b56196e403 · inbound

Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale cites this paper.

Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale Scalable-Softmax Is Superior for Attention

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:48:55.216961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T20:43:40.938898Z digest=sha256:4fb7f5cad1bff035665af0ce2ad737193bf5ba641efb01536cb2276a00408db4