Pith. sign in

Paper Citation Record · LEDGER

Scalable-Softmax Is Superior for Attention

As of 11 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 16 inbound Pith citation observations for arXiv:2501.19399.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.19399 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T20:34:01.273009Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:41:54.830690Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T17:09:59.046151Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2ed10e6f-91e7-4e9f-ae4b-3f33a355d427 · outbound

This paper cites write newline.

Scalable-Softmax Is Superior for Attention write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.134673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.134673Z digest=sha256:79879c98d36d912a814fcb15422a9d10ce2444caaa7bf9a4028bc3761adf2867

Observation 4b958027-fc05-4726-83fd-f03e1142615a · outbound

This paper cites Etc: Encoding long and structured inputs in transformers.

Scalable-Softmax Is Superior for Attention Etc: Encoding long and structured inputs in transformers

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.721755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.140337Z digest=sha256:725900d68a925afc2dce70252f9daffce166ecf81e64dd852e9a9a6433bf15fd

Observation ba82097c-5b3d-40ee-80d0-9979505729d3 · outbound

This paper cites Needle in a haystack - pressure testing llms, 2023.

Scalable-Softmax Is Superior for Attention Needle in a haystack - pressure testing llms, 2023

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.709283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.144584Z digest=sha256:a3c63d7be91ce3be5ce123fda1ae5124cf6dee5573a9d674ef511ab3b91ca278

Observation 92477d2e-ba56-47c1-9b9b-52b202cd2518 · outbound

This paper cites Longformer: The Long-Document Transformer.

Scalable-Softmax Is Superior for Attention Longformer: The Long-Document Transformer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.149151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.149151Z digest=sha256:19637d0d7c217ca79d01ae962cbf4ed2138dd47becd5f24b1556759d45b7a9d3

Observation af892ac0-26ae-4e92-9a58-f6922ec37907 · outbound

This paper cites an unresolved cited work.

Scalable-Softmax Is Superior for Attention Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-09T20:34:01.696917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.155036Z digest=sha256:fa2e278660b5558131b4947f30033bd4640f0de69933c40db88d2af822893b63

Observation b27334a6-9fb7-4152-b4ed-c1dc68e81729 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Scalable-Softmax Is Superior for Attention Generating Long Sequences with Sparse Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.160153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.160153Z digest=sha256:980cc92498a857e895530a37fc8470e5c76148fd973acb1a2c24d7657acd8026

Observation e666f589-e24b-4c15-9bc4-2e326da9c711 · outbound

This paper cites Redpajama: An open source recipe to reproduce llama training dataset, April 2023.

Scalable-Softmax Is Superior for Attention Redpajama: An open source recipe to reproduce llama training dataset, April 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.684935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.165155Z digest=sha256:580e074541949911c6e0d90fd3ff776b7842bb5b40ebafe88ebf9e520e08cf15

Observation 2e378df9-9919-4659-9baa-865950801315 · outbound

This paper cites GMAT: Global Memory Augmentation for Transformers.

Scalable-Softmax Is Superior for Attention GMAT: Global Memory Augmentation for Transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.170481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.170481Z digest=sha256:a4414b337cffc81a956f495e121c056c0cdd544d650680a795ceb04b6ed7e7c9

Observation 5084afc2-b111-4f59-a5a8-6eeef06233c2 · outbound

This paper cites Needle in a haystack - pressure testing llms, 2023.

Scalable-Softmax Is Superior for Attention Needle in a haystack - pressure testing llms, 2023

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.670341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.174657Z digest=sha256:a34837e896a95a2921002bb30ec2027651c7dd04411360eb6af36067a1779bf5

Observation 55b2139c-c9c9-4bba-9c9d-e91538b5d273 · outbound

This paper cites The impact of positional encoding on length generalization in transformers.

Scalable-Softmax Is Superior for Attention The impact of positional encoding on length generalization in transformers

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.656333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.178471Z digest=sha256:c8ad6b5fe564712e74f9f87d37a89702953487482c5fd51795810154ea12c901

Observation 42a47d50-98d9-4644-ae9a-52cec5fac23c · outbound

This paper cites Reformer: The Efficient Transformer.

Scalable-Softmax Is Superior for Attention Reformer: The Efficient Transformer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.182826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.182826Z digest=sha256:f5f20d6f61bba1c2a89fe7ddc200f823bc9c47a1e1a0e0e7323ff4c9807b2939

Observation dbe4dab1-c88c-4c3c-9d77-8444c61fdecd · outbound

This paper cites Gradient-based learning applied to document recognition.

Scalable-Softmax Is Superior for Attention Gradient-based learning applied to document recognition

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.187475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.187475Z digest=sha256:a291d52919d1bbc6a78cd50b6e7f704b1b66e5b0f22521f3c07bbcbb7c25c25d

Observation 088bc4d1-82ce-41e8-9d7f-dff5b338909a · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

Scalable-Softmax Is Superior for Attention World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.191643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.191643Z digest=sha256:6b4238418d002f271b6d25eab0d8ec25b1be67896b936031097a8d5e77158e87

Observation 7fabc220-9369-4818-93e1-06316fc0aade · outbound

This paper cites Scaling laws of ro PE -based extrapolation.

Scalable-Softmax Is Superior for Attention Scaling laws of ro PE -based extrapolation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.632605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.196011Z digest=sha256:a5af9513d12e7ccab323c1f2cdb33258bbb74b67a7becef75287e72656f5c8cf

Observation c0c7beda-2a5e-4b4d-8d56-43c9b9386cb7 · outbound

This paper cites and Hutter, F.

Scalable-Softmax Is Superior for Attention and Hutter, F

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.620244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.200339Z digest=sha256:6cfdee399f01c38968fc0a1b0292a50b1d06d7613fa018359e471414355d8e20

Observation 80e93c09-6eeb-4279-a265-fa4c0baf0979 · outbound

This paper cites Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation.

Scalable-Softmax Is Superior for Attention Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.204290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.204290Z digest=sha256:34c1bd2ba77d7cbad4af5bd6d4cdb97194decb46a9462f9990916d2a9cfcf5cb

Observation c07752c3-592f-4efa-a1bc-8f42f540947f · outbound

This paper cites Language models are unsupervised multitask learners.

Scalable-Softmax Is Superior for Attention Language models are unsupervised multitask learners

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.608076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.208531Z digest=sha256:3182456757cb267f544db5bd410cc1285a0a4c5ab2c07ea978105dbc759ac524

Observation 672dcff2-cfd2-4416-9f5c-87c364774549 · outbound

This paper cites SQ u AD : 100,000+ questions for machine comprehension of text.

Scalable-Softmax Is Superior for Attention SQ u AD : 100,000+ questions for machine comprehension of text

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.594186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.212024Z digest=sha256:7ac1b36ad2406b3ff9ba1ae202879d50b06cffa8d7165542debbf09e1f06e3fe

Observation 1a5ae387-fcab-438b-8b31-4ce5dcee1a3e · outbound

This paper cites Know what you don't know: Unanswerable questions for SQ u AD.

Scalable-Softmax Is Superior for Attention Know what you don't know: Unanswerable questions for SQ u AD

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.580066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.216919Z digest=sha256:3b81576b3f21d6d9173d0b53cd57d7c01de873ff2a30d69a0aa04a8da17e42e6

Observation ae3452d9-32d1-47c4-827d-ca61e191e457 · outbound

This paper cites Searching for Activation Functions.

Scalable-Softmax Is Superior for Attention Searching for Activation Functions

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.220920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.220920Z digest=sha256:81b1868e619ee920e1ba2b5bf613f44f224fe1e26e042e692d8991ea4a1bd8b2

Observation 03b547a5-d657-48d4-a3c3-832377568d2d · outbound

This paper cites Efficient content-based sparse attention with routing transformers.

Scalable-Softmax Is Superior for Attention Efficient content-based sparse attention with routing transformers

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.565825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.225849Z digest=sha256:8db114f22d1d551f67fba38eb098e6488a156e0a45b46293dc4c54cee2617f39

Observation 22a26b0f-8430-46a5-aed1-e915167855ad · outbound

This paper cites Self-attention with relative position representations.

Scalable-Softmax Is Superior for Attention Self-attention with relative position representations

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.552563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.229542Z digest=sha256:151dc977d2972ccd05e2705d198daa9f1edc3df404c79342df0dadc39683a611

Observation fc8e0108-2bb8-48fb-9c18-460b87a6c619 · outbound

This paper cites GLU Variants Improve Transformer.

Scalable-Softmax Is Superior for Attention GLU Variants Improve Transformer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.233103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.233103Z digest=sha256:abe2bf937bc395cb25dcb22be1c894592558d346a926f43d05bd3ac5bc3f42b8

Observation af07bdd1-1197-44fe-b133-9f926d5b3124 · outbound

This paper cites R., Hestness, J., and Dey, N.

Scalable-Softmax Is Superior for Attention R., Hestness, J., and Dey, N

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.537092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.237096Z digest=sha256:57ecacac7b00a6e7164ab6e53123b6f2c7746c1605925690d15379de265bb50e

Observation 7896d758-26e6-412c-820a-1a804a63671a · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

Scalable-Softmax Is Superior for Attention Roformer: Enhanced transformer with rotary position embedding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.240997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.240997Z digest=sha256:1dcb67a715744e383461f0c45c0327221b237e4b63f660ebab185764f98aba2e

Observation f87f8846-8bc2-4896-8e5a-4a95b40b0be6 · outbound

This paper cites Adaptive attention span in transformers.

Scalable-Softmax Is Superior for Attention Adaptive attention span in transformers

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.515253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.244520Z digest=sha256:2dc1f551a307284c7aa7caea6d85feb3430d58a2bb160004b6abf86b0949e52c

Observation e7d417f5-166c-4484-b821-73942fafef59 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Scalable-Softmax Is Superior for Attention Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.249086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.249086Z digest=sha256:3e72d10ce5ac5e77e11c9e4edf418dd86cef34cc90ee1254aa05ad8e893d70fe

Observation d767d1a9-4b1d-47ec-9edc-964290771642 · outbound

This paper cites N., Kaiser, ., and Polosukhin, I.

Scalable-Softmax Is Superior for Attention N., Kaiser, ., and Polosukhin, I

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.502711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.253366Z digest=sha256:c430d371c51aac3fa648a58d90fdbaeec93f021a56436342c27a06e3e8b94af3

Observation c72e8e0f-57f3-4f34-b4ce-a6db9d23ebcd · outbound

This paper cites Length Generalization of Causal Transformers without Position Encoding.

Scalable-Softmax Is Superior for Attention Length Generalization of Causal Transformers without Position Encoding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.257094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.257094Z digest=sha256:184a2f41fc5104ba12bfce6861092cd601c20738be17336e834653d8bfa8df21

Observation 8bbe99af-eec7-405e-b837-61811ca41f61 · outbound

This paper cites V., and Zhou, D.

Scalable-Softmax Is Superior for Attention V., and Zhou, D

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.490700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.261120Z digest=sha256:6cd214ed624d7c797ad8e4b00d6bfcaf9a701851590647f8689f40ab40b2359b

Observation 7cb6b1d2-436f-4fab-b038-103e2a046339 · outbound

This paper cites Differential Transformer.

Scalable-Softmax Is Superior for Attention Differential Transformer

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.265036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.265036Z digest=sha256:ef2bbe5fec35e469e33d1a8f10f448be5b04e6e8ad517223487c49c0ba411ec6

Observation 451730d9-5fce-48c8-bec8-057ca56429ac · outbound

This paper cites A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., and Ahmed, A.

Scalable-Softmax Is Superior for Attention A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., and Ahmed, A

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T20:34:01.473502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T20:34:01.269080Z digest=sha256:a9ab0366ce973f08e5f783200599d50618f156222cd657a821c86a65ded748d0

Observation 8e252d1a-4a8f-4e76-9075-1c1ec35fb147 · outbound

This paper cites and Sennrich, R.

Scalable-Softmax Is Superior for Attention and Sennrich, R

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.273009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.273009Z digest=sha256:bf091b20bdedb3e797bbd21cc7fc97d046334be61f1eac9bbe31381be01c4b95

Pith citing papers

Observation 6280ffc3-a12f-40fe-9030-78aab512c9ec · inbound

On the Mathematical Impossibility of Safe Universal Approximators cites this paper.

On the Mathematical Impossibility of Safe Universal Approximators Scalable-Softmax Is Superior for Attention

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T20:41:54.830690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:41:54.830690Z digest=sha256:18a6d6fbe2ff0bb2c6a8763d8ac31be6c7ecdc9f656214fa4c6dc9903856ee86

Observation ecce2904-108a-4489-bac0-607d95f91640 · inbound

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs cites this paper.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Scalable-Softmax Is Superior for Attention

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:55.007461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:55.007461Z digest=sha256:dd02a7e0dfb615b734b89ef0e1da6ad003f54f309518454e4f3a92bab8d3d35a

Observation e8ec1c17-07c2-4575-b098-96245eb944d3 · inbound

Critical attention scaling in long-context transformers cites this paper.

Critical attention scaling in long-context transformers Scalable-Softmax Is Superior for Attention

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:28.867756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:28.867756Z digest=sha256:1a887523e02b538798ab71a562f35c769290df362c8e13ce58bd8cae63f057df

Observation 6ad89e38-fc74-4506-91b7-74b7c349c60d · inbound

Ministral 3 cites this paper.

Ministral 3 Scalable-Softmax Is Superior for Attention

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:12:24.763656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T19:12:24.627033Z digest=sha256:280cbb79d463805c2cc4f45e7f5278429a6fa010b0024ae53e0610515bf98f2c

Observation 547c0705-9c85-459c-bb4d-3885f7f86b6c · inbound

Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language Modeling cites this paper.

Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language Modeling Scalable-Softmax Is Superior for Attention

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-03T09:57:48.435605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:57:48.435605Z digest=sha256:92b2d59e9157f88f3dd6799bd623506fdc1886027e125a421589ae51c4de2c52

Observation febb80e5-4542-49f1-b831-16c14a806b94 · inbound

MemDLM: Memory-Enhanced DLM Training cites this paper.

MemDLM: Memory-Enhanced DLM Training Scalable-Softmax Is Superior for Attention

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:43:24.547738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T00:42:22.989588Z digest=sha256:1e15a37b888807916d635adf423e91f714a2b0eb58013c6e265d4b61cbb86023

Observation 8214c335-6e25-4887-94e1-bca7aca29e4e · inbound

Screening Is Enough cites this paper.

Screening Is Enough Scalable-Softmax Is Superior for Attention

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:18:20.924558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T22:17:13.803778Z digest=sha256:4fb282943a24c243c5790953864a44eae857ffd49d7dd875faed1a412ef6dffa

Observation 7e276543-d219-40ab-870a-e97c1ee32e87 · inbound

ExtraVAR: Stage-Aware RoPE Remapping for Resolution Extrapolation in Visual Autoregressive Models cites this paper.

ExtraVAR: Stage-Aware RoPE Remapping for Resolution Extrapolation in Visual Autoregressive Models Scalable-Softmax Is Superior for Attention

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:26:28.046243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:01:32.762236Z digest=sha256:ad33e87063dc45e8efdcc5319396e9594309fbcc4c20d34e82d7c5c65b093cf1

Observation 3a902776-13d8-420a-9934-45d5a0eca3c9 · inbound

Phoenix-VL 1.5 Medium Technical Report cites this paper.

Phoenix-VL 1.5 Medium Technical Report Scalable-Softmax Is Superior for Attention

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.890914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:41:27.144815Z digest=sha256:5772d076fa1d63e58eada88589598a04855faec605519f038369ca3137336232

Observation 87dc7db5-a504-499b-8cda-600fe95b4f51 · inbound

A Unified Framework for Critical Scaling of Inverse Temperature in Self-Attention cites this paper.

A Unified Framework for Critical Scaling of Inverse Temperature in Self-Attention Scalable-Softmax Is Superior for Attention

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:57:53.542250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-14T19:56:02.462786Z digest=sha256:b0f5a646363e9f00fa4ea9f3119bd0cc2515e27c0264353e4cfb12535a7cb3a2

Observation 059d601f-3e12-4ecd-b912-7690e1872ae0 · inbound

TabPFN-3: Technical Report cites this paper.

TabPFN-3: Technical Report Scalable-Softmax Is Superior for Attention

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:10:06.047420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T06:05:16.199413Z digest=sha256:41bb007bca8a8380ace8abe605b972ca35578ead79e6a15ddff35fd314c3e5cc

Observation 45c62d5c-a0b1-49e2-907c-bd68b278b9b6 · inbound

TabPFN-3: Technical Report cites this paper.

TabPFN-3: Technical Report Scalable-Softmax Is Superior for Attention

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:25:46.516167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T21:37:08.627791Z digest=sha256:f95baf4072d4a9392b8a372ef546deca14509b5cb36c2b725535c3c1cff13bbe

Observation 48068909-0dd8-46d2-a385-e779d598b381 · inbound

Most Transformer Modifications Still Do Not Transfer at 1-3B: A 2020-2026 Update to Narang et al. (2021) with Downstream Evaluation and a Noise Floor cites this paper.

Most Transformer Modifications Still Do Not Transfer at 1-3B: A 2020-2026 Update to Narang et al. (2021) with Downstream Evaluation and a Noise Floor Scalable-Softmax Is Superior for Attention

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:19:42.039766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-21T06:15:47.451870Z digest=sha256:7d6c08a3277489507a32d744c822b1fa60eecbb3ae328d7449724d393bed4bf8

Observation d54cc909-4148-41a9-af19-ec29a218b84a · inbound

ATMA: Long-Context Language Modeling via Polar Attention and Gated-Delta Compression Memory cites this paper.

ATMA: Long-Context Language Modeling via Polar Attention and Gated-Delta Compression Memory Scalable-Softmax Is Superior for Attention

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:09:59.047893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-25T23:52:52.985532Z digest=sha256:716731fef355e41e28fe17b5f3c09f2d9aa8e5e5e7e9bd21e397eb53f7a27f97

Observation 5924af4a-ecc5-4c0c-a99f-4170508d4130 · inbound

ATMA: Long-Context Language Modeling via Polar Attention and Gated-Delta Compression Memory cites this paper.

ATMA: Long-Context Language Modeling via Polar Attention and Gated-Delta Compression Memory Scalable-Softmax Is Superior for Attention

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:44:37.730997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T09:38:43.250420Z digest=sha256:e55111707844a47d4190806c9f6be7631dbe0efd90f99d07ebdc405ba31465b4

Observation 91be52e2-032d-4042-a936-f8b56196e403 · inbound

Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale cites this paper.

Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale Scalable-Softmax Is Superior for Attention

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:48:55.216961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-03T20:43:40.938898Z digest=sha256:067c85f52d7e71177deef480bd169a89ba5816028ef37c1d306861ba8cd22294