Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:19:59.124623Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2507.08637.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:19:59.124623Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
22 of 22 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fd8dd414-dd24-4ab0-97ee-9560488ff616 · outbound
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a081d3e3-d084-4854-b1c2-170057c786fb · outbound
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) On the Use of ArXiv as a Dataset
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 109a4e65-e163-46fb-aa3f-825e0a6d4ab8 · outbound
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c854729-b3e7-41cf-982e-9683292a4243 · outbound
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e5daeec-0946-4337-a372-26d9a9fed0fc · outbound
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5daa5ea-e5a2-4432-8753-c109029e39dd · outbound
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Efficiently Modeling Long Sequences with Structured State Spaces
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ba0c6f1-76da-44b3-b17b-a3107a7cd05a · outbound
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Mega: Moving Average Equipped Gated Attention
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4682cb19-479f-4f81-9aa7-bd7cfd362be3 · outbound
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) RWKV: Reinventing RNNs for the Transformer Era
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 892858da-22bc-4cfa-b2b6-4f01e6cb0adf · outbound
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Random Feature Attention
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8da21433-da9a-4d10-a229-6a4651f9ca91 · outbound
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Compressive Transformers for Long-Range Sequence Modelling
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c7cea95-11b6-462b-8566-4de0ae71d0e9 · outbound
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Attention Is All You Need
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42f71596-37b0-4ead-bc7a-d3545bbf0d89 · outbound
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Linformer: Self-Attention with Linear Complexity
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf4b19f1-f964-4045-9576-875964532731 · outbound
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Tian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang, Liang Sun, and Rong Jin
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 92a9b74e-4d6c-4b38-b474-fc4fc78b8561 · outbound
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e9d6b606-b5d2-4735-880d-d1ec74f053fb · outbound
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Here, R ∈ Rd×m has independent and identically distributed Gaussian entries, E ϕ(q)T ϕ(k) − κ(q, k) ≤ C√m , where κ refers to the softmax kernel
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3d6ca54a-980e-4ec0-b14d-1495f126d14d · outbound
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Unresolved cited work
Reference 1999
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cf960b86-006c-4967-81c9-d1efade079cf · outbound
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Generating Long Sequences with Sparse Transformers
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cde1ea8-d033-463a-a18e-5d66e60c7560 · outbound
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Longformer: The Long-Document Transformer
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec2fe666-c143-432c-9d6c-f682d7b3ee8b · outbound
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) FNet: Mixing Tokens with Fourier Transforms
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14bbb4a8-30e5-4edf-bad3-72df948be2a4 · outbound
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Peter L Bartlett and Shahar Mendelson
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9ae1559d-846f-472c-85b1-5fd428862434 · outbound
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Bert: Pre-training of deep bidirectional transformers for language understanding
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 45f548ce-8e1f-4ee9-8e50-e8e145103771 · outbound
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) On the Properties of the Softmax Function with Application in Game Theory and Reinforcement Learning
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.