Pith. sign in

Paper Citation Record · LEDGER

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA)

As of 17 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2507.08637.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.08637 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:19:59.124623Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fd8dd414-dd24-4ab0-97ee-9560488ff616 · outbound

This paper cites an unresolved cited work.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:19:59.425992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:19:59.120424Z digest=sha256:4b15dd0c36ea20127fdbe531c78bddf256446fa4c7ea362f935cc73accb87aa1

Observation a081d3e3-d084-4854-b1c2-170057c786fb · outbound

This paper cites On the Use of ArXiv as a Dataset.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) On the Use of ArXiv as a Dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:59.002779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:59.002779Z digest=sha256:92d8c4620d0ec84ea69402e3e1f4f1117031bea0bbcb79f3451a1654beecc94f

Observation 109a4e65-e163-46fb-aa3f-825e0a6d4ab8 · outbound

This paper cites Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:59.026663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:59.026663Z digest=sha256:6cc7712c2c54749697a10862b24cd43299da2bd0db87fc17973b83b4128f793b

Observation 7c854729-b3e7-41cf-982e-9683292a4243 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:59.030523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:59.030523Z digest=sha256:82408747c36407277d7298047b011d6e1b25cbca7a53bb96470646d7e70a2795

Observation 5e5daeec-0946-4337-a372-26d9a9fed0fc · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:59.039722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:59.039722Z digest=sha256:5b65d61cca7ab1251f9fa300942047b8d0fca5dba9172c2fd13040cba1190d44

Observation c5daa5ea-e5a2-4432-8753-c109029e39dd · outbound

This paper cites Efficiently Modeling Long Sequences with Structured State Spaces.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Efficiently Modeling Long Sequences with Structured State Spaces

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:59.055602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:59.055602Z digest=sha256:df457ec11fc6a2b0cc2151247561def47c0558074d4f4187ceffd9efd6b0574d

Observation 2ba0c6f1-76da-44b3-b17b-a3107a7cd05a · outbound

This paper cites Mega: Moving Average Equipped Gated Attention.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Mega: Moving Average Equipped Gated Attention

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:59.066945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:59.066945Z digest=sha256:f6651beb17636220d123da601493b1ce1bdea0219101517e428d2b426393039c

Observation 4682cb19-479f-4f81-9aa7-bd7cfd362be3 · outbound

This paper cites RWKV: Reinventing RNNs for the Transformer Era.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) RWKV: Reinventing RNNs for the Transformer Era

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:59.080608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:59.080608Z digest=sha256:50784392d5568f8f9433a899114ad437575997ec3b7e0cecd47bbf4bc0322a49

Observation 892858da-22bc-4cfa-b2b6-4f01e6cb0adf · outbound

This paper cites Random Feature Attention.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Random Feature Attention

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:59.085521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:59.085521Z digest=sha256:89a41289200020c9bc744b1647d1da06b0c9885ac77442ff94e9123cc9582fbd

Observation 8da21433-da9a-4d10-a229-6a4651f9ca91 · outbound

This paper cites Compressive Transformers for Long-Range Sequence Modelling.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Compressive Transformers for Long-Range Sequence Modelling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:59.088847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:59.088847Z digest=sha256:a69fa437b5a274db66d04532facf1135275c1449d03466ebc8c9b6d3b409b46e

Observation 5c7cea95-11b6-462b-8566-4de0ae71d0e9 · outbound

This paper cites Attention Is All You Need.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Attention Is All You Need

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:59.102425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:59.102425Z digest=sha256:06fa42a11a0beabd981f4282acae73b462e81ee877d433df032356f9bb3969a8

Observation 42f71596-37b0-4ead-bc7a-d3545bbf0d89 · outbound

This paper cites Linformer: Self-Attention with Linear Complexity.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Linformer: Self-Attention with Linear Complexity

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:59.106177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:59.106177Z digest=sha256:2aa25c3c5be72cae35bb9ae8010afe19c3d8c18f66e2ecf5f496606631f70d30

Observation bf4b19f1-f964-4045-9576-875964532731 · outbound

This paper cites Tian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang, Liang Sun, and Rong Jin.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Tian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang, Liang Sun, and Rong Jin

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.454835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:19:59.112324Z digest=sha256:daaf2fa6a629c6e65681dce11ff276e9193e9d5b29288061ac09100dedd38a1b

Observation 92a9b74e-4d6c-4b38-b474-fc4fc78b8561 · outbound

This paper cites an unresolved cited work.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:19:59.443101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:19:59.116208Z digest=sha256:9f18e28da8e9737b30b07a30f9315522eeed3173dae68ac9cb25374ac9b968fd

Observation e9d6b606-b5d2-4735-880d-d1ec74f053fb · outbound

This paper cites Here, R ∈ Rd×m has independent and identically distributed Gaussian entries, E ϕ(q)T ϕ(k) − κ(q, k) ≤ C√m , where κ refers to the softmax kernel.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Here, R ∈ Rd×m has independent and identically distributed Gaussian entries, E ϕ(q)T ϕ(k) − κ(q, k) ≤ C√m , where κ refers to the softmax kernel

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.413911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:19:59.124623Z digest=sha256:00ebaf96803224d6180314c295286acf5f41581f66ee8f53b43ec6987881f610

Observation 3d6ca54a-980e-4ec0-b14d-1495f126d14d · outbound

This paper cites an unresolved cited work.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Unresolved cited work

Reference 1999

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:19:59.467682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:19:59.075723Z digest=sha256:5b0172b92421fad57a1288d9998af965361ab66d15151ac030e2eb585594d676

Observation cf960b86-006c-4967-81c9-d1efade079cf · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Generating Long Sequences with Sparse Transformers

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:58.997598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:58.997598Z digest=sha256:788093d38e532f4dfb3e23e8b3a98cf476a2dcbe715cdf08d2f6ee25562c042f

Observation 1cde1ea8-d033-463a-a18e-5d66e60c7560 · outbound

This paper cites Longformer: The Long-Document Transformer.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Longformer: The Long-Document Transformer

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:58.989559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:58.989559Z digest=sha256:2900a606bbef8ce7a6ed0d4757d77894c97ad79710d0afde4151bfedeafeb597

Observation ec2fe666-c143-432c-9d6c-f682d7b3ee8b · outbound

This paper cites FNet: Mixing Tokens with Fourier Transforms.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) FNet: Mixing Tokens with Fourier Transforms

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:59.062242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:59.062242Z digest=sha256:5df9db139ad1abc963a1c8392f8d3540e786208b74a980420506192b25488782

Observation 14bbb4a8-30e5-4edf-bad3-72df948be2a4 · outbound

This paper cites Peter L Bartlett and Shahar Mendelson.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Peter L Bartlett and Shahar Mendelson

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.512551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:19:58.976895Z digest=sha256:9a09503ec7cc36e2250a9e1f32dcaeba3a543d93addd84cb87c478696309d1bb

Observation 9ae1559d-846f-472c-85b1-5fd428862434 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.489106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:19:59.035017Z digest=sha256:4a0a797ba2fee9a48a006b98db9226ee19cce9c45dcdd421cd7d0796dae67186

Observation 45f548ce-8e1f-4ee9-8e50-e8e145103771 · outbound

This paper cites On the Properties of the Softmax Function with Application in Game Theory and Reinforcement Learning.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) On the Properties of the Softmax Function with Application in Game Theory and Reinforcement Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:59.044342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:59.044342Z digest=sha256:3e7de7048066f734f55403110b866e3bbfc330e79c08106f6fcadd30783b726d

Pith citing papers

No inbound Pith citation observations are available.