Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T20:34:01.273009Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 16 inbound Pith citation observations for arXiv:2501.19399.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T20:34:01.273009Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:41:54.830690Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T17:09:59.046151Z
33 of 33 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2ed10e6f-91e7-4e9f-ae4b-3f33a355d427 · outbound
Scalable-Softmax Is Superior for Attention write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b958027-fc05-4726-83fd-f03e1142615a · outbound
Scalable-Softmax Is Superior for Attention Etc: Encoding long and structured inputs in transformers
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ba82097c-5b3d-40ee-80d0-9979505729d3 · outbound
Scalable-Softmax Is Superior for Attention Needle in a haystack - pressure testing llms, 2023
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 92477d2e-ba56-47c1-9b9b-52b202cd2518 · outbound
Scalable-Softmax Is Superior for Attention Longformer: The Long-Document Transformer
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af892ac0-26ae-4e92-9a58-f6922ec37907 · outbound
Scalable-Softmax Is Superior for Attention Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b27334a6-9fb7-4152-b4ed-c1dc68e81729 · outbound
Scalable-Softmax Is Superior for Attention Generating Long Sequences with Sparse Transformers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e666f589-e24b-4c15-9bc4-2e326da9c711 · outbound
Scalable-Softmax Is Superior for Attention Redpajama: An open source recipe to reproduce llama training dataset, April 2023
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2e378df9-9919-4659-9baa-865950801315 · outbound
Scalable-Softmax Is Superior for Attention GMAT: Global Memory Augmentation for Transformers
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5084afc2-b111-4f59-a5a8-6eeef06233c2 · outbound
Scalable-Softmax Is Superior for Attention Needle in a haystack - pressure testing llms, 2023
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 55b2139c-c9c9-4bba-9c9d-e91538b5d273 · outbound
Scalable-Softmax Is Superior for Attention The impact of positional encoding on length generalization in transformers
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 42a47d50-98d9-4644-ae9a-52cec5fac23c · outbound
Scalable-Softmax Is Superior for Attention Reformer: The Efficient Transformer
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbe4dab1-c88c-4c3c-9d77-8444c61fdecd · outbound
Scalable-Softmax Is Superior for Attention Gradient-based learning applied to document recognition
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 088bc4d1-82ce-41e8-9d7f-dff5b338909a · outbound
Scalable-Softmax Is Superior for Attention World Model on Million-Length Video And Language With Blockwise RingAttention
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fabc220-9369-4818-93e1-06316fc0aade · outbound
Scalable-Softmax Is Superior for Attention Scaling laws of ro PE -based extrapolation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c0c7beda-2a5e-4b4d-8d56-43c9b9386cb7 · outbound
Scalable-Softmax Is Superior for Attention and Hutter, F
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 80e93c09-6eeb-4279-a265-fa4c0baf0979 · outbound
Scalable-Softmax Is Superior for Attention Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c07752c3-592f-4efa-a1bc-8f42f540947f · outbound
Scalable-Softmax Is Superior for Attention Language models are unsupervised multitask learners
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 672dcff2-cfd2-4416-9f5c-87c364774549 · outbound
Scalable-Softmax Is Superior for Attention SQ u AD : 100,000+ questions for machine comprehension of text
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1a5ae387-fcab-438b-8b31-4ce5dcee1a3e · outbound
Scalable-Softmax Is Superior for Attention Know what you don't know: Unanswerable questions for SQ u AD
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ae3452d9-32d1-47c4-827d-ca61e191e457 · outbound
Scalable-Softmax Is Superior for Attention Searching for Activation Functions
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03b547a5-d657-48d4-a3c3-832377568d2d · outbound
Scalable-Softmax Is Superior for Attention Efficient content-based sparse attention with routing transformers
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 22a26b0f-8430-46a5-aed1-e915167855ad · outbound
Scalable-Softmax Is Superior for Attention Self-attention with relative position representations
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fc8e0108-2bb8-48fb-9c18-460b87a6c619 · outbound
Scalable-Softmax Is Superior for Attention GLU Variants Improve Transformer
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af07bdd1-1197-44fe-b133-9f926d5b3124 · outbound
Scalable-Softmax Is Superior for Attention R., Hestness, J., and Dey, N
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7896d758-26e6-412c-820a-1a804a63671a · outbound
Scalable-Softmax Is Superior for Attention Roformer: Enhanced transformer with rotary position embedding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f87f8846-8bc2-4896-8e5a-4a95b40b0be6 · outbound
Scalable-Softmax Is Superior for Attention Adaptive attention span in transformers
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e7d417f5-166c-4484-b821-73942fafef59 · outbound
Scalable-Softmax Is Superior for Attention Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d767d1a9-4b1d-47ec-9edc-964290771642 · outbound
Scalable-Softmax Is Superior for Attention N., Kaiser, ., and Polosukhin, I
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c72e8e0f-57f3-4f34-b4ce-a6db9d23ebcd · outbound
Scalable-Softmax Is Superior for Attention Length Generalization of Causal Transformers without Position Encoding
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bbe99af-eec7-405e-b837-61811ca41f61 · outbound
Scalable-Softmax Is Superior for Attention V., and Zhou, D
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7cb6b1d2-436f-4fab-b038-103e2a046339 · outbound
Scalable-Softmax Is Superior for Attention Differential Transformer
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 451730d9-5fce-48c8-bec8-057ca56429ac · outbound
Scalable-Softmax Is Superior for Attention A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., and Ahmed, A
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8e252d1a-4a8f-4e76-9075-1c1ec35fb147 · outbound
Scalable-Softmax Is Superior for Attention and Sennrich, R
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6280ffc3-a12f-40fe-9030-78aab512c9ec · inbound
On the Mathematical Impossibility of Safe Universal Approximators Scalable-Softmax Is Superior for Attention
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecce2904-108a-4489-bac0-607d95f91640 · inbound
AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Scalable-Softmax Is Superior for Attention
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8ec1c17-07c2-4575-b098-96245eb944d3 · inbound
Critical attention scaling in long-context transformers Scalable-Softmax Is Superior for Attention
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ad89e38-fc74-4506-91b7-74b7c349c60d · inbound
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 547c0705-9c85-459c-bb4d-3885f7f86b6c · inbound
Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language Modeling Scalable-Softmax Is Superior for Attention
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation febb80e5-4542-49f1-b831-16c14a806b94 · inbound
MemDLM: Memory-Enhanced DLM Training Scalable-Softmax Is Superior for Attention
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8214c335-6e25-4887-94e1-bca7aca29e4e · inbound
Screening Is Enough Scalable-Softmax Is Superior for Attention
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7e276543-d219-40ab-870a-e97c1ee32e87 · inbound
ExtraVAR: Stage-Aware RoPE Remapping for Resolution Extrapolation in Visual Autoregressive Models Scalable-Softmax Is Superior for Attention
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3a902776-13d8-420a-9934-45d5a0eca3c9 · inbound
Phoenix-VL 1.5 Medium Technical Report Scalable-Softmax Is Superior for Attention
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 87dc7db5-a504-499b-8cda-600fe95b4f51 · inbound
A Unified Framework for Critical Scaling of Inverse Temperature in Self-Attention Scalable-Softmax Is Superior for Attention
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 059d601f-3e12-4ecd-b912-7690e1872ae0 · inbound
TabPFN-3: Technical Report Scalable-Softmax Is Superior for Attention
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 45c62d5c-a0b1-49e2-907c-bd68b278b9b6 · inbound
TabPFN-3: Technical Report Scalable-Softmax Is Superior for Attention
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 48068909-0dd8-46d2-a385-e779d598b381 · inbound
Most Transformer Modifications Still Do Not Transfer at 1-3B: A 2020-2026 Update to Narang et al. (2021) with Downstream Evaluation and a Noise Floor Scalable-Softmax Is Superior for Attention
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d54cc909-4148-41a9-af19-ec29a218b84a · inbound
ATMA: Length-Invariant Language Modeling via Polar Attention and Gated-Delta Compression Memory Scalable-Softmax Is Superior for Attention
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5924af4a-ecc5-4c0c-a99f-4170508d4130 · inbound
ATMA: Length-Invariant Language Modeling via Polar Attention and Gated-Delta Compression Memory Scalable-Softmax Is Superior for Attention
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 91be52e2-032d-4042-a936-f8b56196e403 · inbound
Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale Scalable-Softmax Is Superior for Attention
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.