Pith. sign in

Paper Citation Record · LEDGER

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing

As of 12 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 4 inbound Pith citation observations for arXiv:2501.00658.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.00658 v2

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:52:33.772825Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:58:54.578841Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T04:25:56.146025Z

Reference resolution

79 of 79 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved56
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation db420ba5-9de6-4c93-839a-53348ab29c1a · outbound

This paper cites The Hidden Attention of Mamba Models.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing The Hidden Attention of Mamba Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.343979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.343979Z digest=sha256:9a1be29e64e0f5893d2486a963670ad0330ba81bcfd885f441b6e32c0d2a9184

Observation e0d11af2-eaeb-46be-b550-ef666c06c5e9 · outbound

This paper cites DeciMamba: Exploring the Length Extrapolation Potential of Mamba.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing DeciMamba: Exploring the Length Extrapolation Potential of Mamba

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.368101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.368101Z digest=sha256:f837f84454949203b4eb9a1efcdf1c9abcce6160e87d5242e7fd84de5299b23d

Observation b3726c25-d3c1-4271-b87f-f66b23dcd435 · outbound

This paper cites Language models are few-shot learners.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Language models are few-shot learners

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.373579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.373579Z digest=sha256:0b7aef01f061924fdaf4c31112ff5afc2460a0e7c13ed3cf3ac6642929c2728c

Observation 0b53bfbc-b713-4d11-b675-4c8034e50fd4 · outbound

This paper cites Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.379010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.379010Z digest=sha256:ca5112cebbbd26b2ea455f3f1d7c40ef8ec3f4b15a159ad6636b8edb86538f43

Observation e648a296-7d88-4cd8-9573-f06a8ec462e6 · outbound

This paper cites Rethinking Attention with Performers.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Rethinking Attention with Performers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.391034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.391034Z digest=sha256:7d697f706dcd1659ba036c347e134b28fe94378628ced3e0fd77276f05e88caf

Observation ad2c096e-eb0b-4ada-97f1-f0301739c813 · outbound

This paper cites Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.396665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.396665Z digest=sha256:e330eba70b2e4f1a2ee2d2a7d540289c817bc74648e20843a6be244730f08e22

Observation ff742bf8-2803-48b4-990b-dd784d4ac900 · outbound

This paper cites Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.407687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.407687Z digest=sha256:1a2782a1556e8172ffc2ba374443fbe945b74bc0fd25c53ecd27703d59ac31e8

Observation 46ff30e7-9068-4de8-a393-2d3f1182e06a · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:35.269547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:52:33.413833Z digest=sha256:a3412640336a0c2457c5096cd1c07a5247e15bde4ff46ff66b9fcf55d0c60669

Observation 65edb53b-c766-4914-a89f-1d00d1585662 · outbound

This paper cites Were RNNs All We Needed?.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Were RNNs All We Needed?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.425335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.425335Z digest=sha256:7bf5ede9a701ebc598fc01599346435d5aa37f1d1abc377fb23b9016cc456136

Observation 8701a239-5a12-4284-9d32-4c65459f1f28 · outbound

This paper cites Hungry Hungry Hippos: Towards Language Modeling with State Space Models.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Hungry Hungry Hippos: Towards Language Modeling with State Space Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.430458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.430458Z digest=sha256:be672c3a86050f222afe6a83f284fa9f7ce24e3b7de57e8998c9ed77088748ae

Observation f44c4f39-53a2-4934-a11d-6ed4523c579c · outbound

This paper cites Efficiently Modeling Long Sequences with Structured State Spaces.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Efficiently Modeling Long Sequences with Structured State Spaces

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.445941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.445941Z digest=sha256:35edb9437f0ce3e6b871115cc99556c2e4ffc57da4c1cc0a96eeb4e5b2aca0e6

Observation 52e4a487-a12e-4691-b1b3-24fed19f0800 · outbound

This paper cites Long short-term memory.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Long short-term memory

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:35.253392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:52:33.450952Z digest=sha256:b1a8cbc092d8633df3b855271ce90515fd0908f4325a4e3a08c33f83024179d5

Observation 196423ba-964a-463b-87b4-7a607dd2f3bc · outbound

This paper cites Repeat After Me: Transformers are Better than State Space Models at Copying.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.461185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.461185Z digest=sha256:c70aeb3420fb9f9592eb26a4c7324470739df3e8658328a4ac8e1e76b2ed7dc3

Observation 7a60ab9c-7c21-4ff1-a378-ee88380691d1 · outbound

This paper cites Mistral 7B.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Mistral 7B

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.466271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.466271Z digest=sha256:6c2be9bb5c0ffb66a9b2ebfffd68203fa21e903bf91945b72187f25c79d2aef9

Observation 4d0cfb62-0c7d-4ea3-8edd-387a9418404a · outbound

This paper cites Scaling Laws for Neural Language Models.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Scaling Laws for Neural Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.471740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.471740Z digest=sha256:411b7e85b77e341100b027587183f1f562e4f862c8b725ccf8f3fca09f243488

Observation cd04fae6-c460-4fb2-9b81-ae140044975a · outbound

This paper cites Semi-Supervised Classification with Graph Convolutional Networks.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Semi-Supervised Classification with Graph Convolutional Networks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.477934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.477934Z digest=sha256:4638bd229eb7060c7303e7f475561c6bd1d716f8c7d6199fea8cb4313da1a140

Observation 981a46df-05ba-48ed-ba67-ae8ce609bb7c · outbound

This paper cites Jamba: A Hybrid Transformer-Mamba Language Model.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Jamba: A Hybrid Transformer-Mamba Language Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.498531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.498531Z digest=sha256:a8c5ca2e62ae0cb175e85136028354ca049e9318c9113ac93a3acb4f8a4fd644

Observation 604fde47-106f-4e2e-8fe9-8f6ca2dc52de · outbound

This paper cites Longhorn: State Space Models are Amortized Online Learners.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Longhorn: State Space Models are Amortized Online Learners

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.504123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.504123Z digest=sha256:180cccb7db130629fa2f7842b696701773975fb88fe381f9f7f81d17ef540a0d

Observation 5a25b695-3eb7-42c2-bead-ee61821c716f · outbound

This paper cites Mega: Moving Average Equipped Gated Attention.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Mega: Moving Average Equipped Gated Attention

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.509637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.509637Z digest=sha256:869e3bb538d8a241c2c7841ab94de2855898863a558a4c815991cc4203a59553

Observation 6b830b8c-4773-4ce7-a0a6-2fdf0ffdabd9 · outbound

This paper cites Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.515259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.515259Z digest=sha256:63fdd5254494b29fe3dbf84e451966fbcc10bf562e29257bcab22e8674ce7bc9

Observation f7c675f7-ccbc-472d-b38e-bf0e31e13161 · outbound

This paper cites The Illusion of State in State-Space Models.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing The Illusion of State in State-Space Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.520485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.520485Z digest=sha256:30d366e77de9cd9365e8758ecd85acf0f216c252eaa9ef5faa2ecdcee2456ebe

Observation 5d2451af-a931-47a5-b4f3-c10f01d81e91 · outbound

This paper cites In-context Learning and Induction Heads.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing In-context Learning and Induction Heads

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.525899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.525899Z digest=sha256:3c53a01b661a7cf10579a9f27c46f05c33fcfd9e5e9360d21a404399b81ce2cb

Observation 94bd2d67-7ceb-4d68-b90d-af485a0d876c · outbound

This paper cites Graph neural networks exponentially lose expressive power for node classification.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Graph neural networks exponentially lose expressive power for node classification

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:35.237621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:52:33.531018Z digest=sha256:60aff0e733400bd3f8c7013ef9cef5729b44eb0e69e8c840789b0d5b09be04d4

Observation bb41ab24-0105-441a-8d79-be64415d5e43 · outbound

This paper cites Can Mamba Learn How to Learn? A Comparative Study on In-Context Learning Tasks.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Can Mamba Learn How to Learn? A Comparative Study on In-Context Learning Tasks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.536068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.536068Z digest=sha256:b85817664cd551859a89e355e8b52005a7ddb1e5127269884a2300d280f1fce1

Observation aa7d45c6-9e40-4381-abdd-c0db406053e1 · outbound

This paper cites RWKV: Reinventing RNNs for the Transformer Era.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing RWKV: Reinventing RNNs for the Transformer Era

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.541213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.541213Z digest=sha256:58bfd2e730d719232432fe95e51c705f778351ffdf7be1fe722d1fbd12993661

Observation f1fe9ae7-d4f5-4272-806a-a7f3dccc2fcd · outbound

This paper cites Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.546904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.546904Z digest=sha256:f06558a9e85bbfd5553488e9ec6ab6cd7e84d89f2999873d41ff39e783a8d42f

Observation 31794b19-e32f-462e-a7be-4115a891856a · outbound

This paper cites Ignore Previous Prompt: Attack Techniques For Language Models.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Ignore Previous Prompt: Attack Techniques For Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.552154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.552154Z digest=sha256:50ebc06d44d7d649a2d15d1f53924e9a8dab6f9748ef46dc249a6846e6ea650b

Observation 3313a84a-12f0-41a4-9f60-523c149c119c · outbound

This paper cites Mechanistic Design and Scaling of Hybrid Architectures.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Mechanistic Design and Scaling of Hybrid Architectures

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.557438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.557438Z digest=sha256:b5c7d85bb936edd864a55607ba1f9067b7459efa906f8696f7a6c368ccda2090

Observation 8b7c8b1c-75c6-4691-824a-824f3bc3daa1 · outbound

This paper cites Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.563025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.563025Z digest=sha256:355f08ee1c84de58d5d8c64cc91740b90be4aaa0ada2143580fd4ec2336b2287

Observation bb9dc605-5b32-4362-95c2-dfc742b0af19 · outbound

This paper cites HGRN2: Gated Linear RNNs with State Expansion.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing HGRN2: Gated Linear RNNs with State Expansion

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.568763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.568763Z digest=sha256:e5fcf0af0d8e21c395c491c75d78585f91936a748be8279105fac21497c35f8a

Observation a2869e78-4e65-4c9a-b094-5fb5bd046626 · outbound

This paper cites Revisiting Over-smoothing in BERT from the Perspective of Graph.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Revisiting Over-smoothing in BERT from the Perspective of Graph

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.574082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.574082Z digest=sha256:9aecc6a3f8b87390456ef481f11de1645af92698f74e1314b7229bd4cdbde2f5

Observation d590677e-9b9b-4c03-97ec-cee95a138271 · outbound

This paper cites Learning to (Learn at Test Time): RNNs with Expressive Hidden States.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Learning to (Learn at Test Time): RNNs with Expressive Hidden States

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.579902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.579902Z digest=sha256:84cbb939021e6eeb3b7126d5d975770f2ba2cf98b64ee8721d6e364f262c470c

Observation f119f4a3-66a9-40a9-afc9-cff89dbd82a2 · outbound

This paper cites A Length-Extrapolatable Transformer.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing A Length-Extrapolatable Transformer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.584955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.584955Z digest=sha256:635e14cb139bc5a7c159ec69886016cf64dd2dffc432dec328205b253bf024d3

Observation 38c5dc08-a5b4-4e9f-8c95-2a9bc85b24e6 · outbound

This paper cites Retentive Network: A Successor to Transformer for Large Language Models.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Retentive Network: A Successor to Transformer for Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.589977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.589977Z digest=sha256:68be2e998dc44ccc077132cc33e9026e5387341cfbf5ba9f58ae3b9cb5c14cc3

Observation 7ee84a7d-addb-4452-8dc3-4ffec2706397 · outbound

This paper cites Long Range Arena: A Benchmark for Efficient Transformers.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Long Range Arena: A Benchmark for Efficient Transformers

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.595036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.595036Z digest=sha256:ca8944dfc09834c19491023e6afdbdae3d9e7398ae8be47a9c7e9c439d4c3042

Observation 7232c5e4-5021-4556-a012-8dae7549c9aa · outbound

This paper cites Understanding over-squashing and bottlenecks on graphs via curvature.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Understanding over-squashing and bottlenecks on graphs via curvature

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.600245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.600245Z digest=sha256:8916e6df3b1ec2f9a3341eae3ea1b686c6f1856f1f8062e8a3fb287ce722f437

Observation 820c490b-6f6a-4707-8891-06174adc6209 · outbound

This paper cites Transformer Dissection: A Unified Understanding of Transformer's Attention via the Lens of Kernel.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Transformer Dissection: A Unified Understanding of Transformer's Attention via the Lens of Kernel

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.605463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.605463Z digest=sha256:8ae298c19a68d32f7bf8e8e78baa99189ec7edc6aa909b057122e9492ce79677

Observation c0f19b57-13a7-4776-a1a8-7d0ea1f40e1c · outbound

This paper cites The unreasonable effectiveness of the forget gate.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing The unreasonable effectiveness of the forget gate

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.610896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.610896Z digest=sha256:1847f834cc850d749844f57755e1e9caee997dad5b2f92abf61187d18bce6144

Observation 3eac356a-996c-4591-981c-374ea6ebbfaa · outbound

This paper cites An Empirical Study of Mamba-based Language Models.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing An Empirical Study of Mamba-based Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.615941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.615941Z digest=sha256:0227edf5230984f1205792b78b024a4059b5ac1bcfcebb0a72546afdc9f7abc2

Observation 9fae8463-f14f-42d4-8f5a-1a071cc6734e · outbound

This paper cites Anti-Oversmoothing in Deep Vision Transformers via the Fourier Domain Analysis: From Theory to Practice.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Anti-Oversmoothing in Deep Vision Transformers via the Fourier Domain Analysis: From Theory to Practice

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.621004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.621004Z digest=sha256:7b52d46af1fcb45bd10114ab7a603a0a4666e8676996c530ffd631b3f5a261aa

Observation ff99a6c3-f4c5-4123-957e-d1daa19e3488 · outbound

This paper cites A Non-Asymptotic Analysis of Oversmoothing in Graph Neural Networks.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing A Non-Asymptotic Analysis of Oversmoothing in Graph Neural Networks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.626180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.626180Z digest=sha256:672a656ec9344ccc3a3555890dd36e8d118ac90c13096a5446fb3fc0a81b92cc

Observation ab356f8d-85e6-4549-bff7-1b23830adf12 · outbound

This paper cites On the Role of Attention Masks and LayerNorm in Transformers.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing On the Role of Attention Masks and LayerNorm in Transformers

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.631796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.631796Z digest=sha256:1b15243bee0b690a12ef9be5dbdd355f596db9f459b56dc8dabe4127193111aa

Observation 4aa22575-10a5-4c71-99b3-af692f1008c1 · outbound

This paper cites Gated Linear Attention Transformers with Hardware-Efficient Training.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Gated Linear Attention Transformers with Hardware-Efficient Training

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.636905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.636905Z digest=sha256:f174bdcb8870d2b9765e7c0211b9c34bb988bfe7225490f6fb396f79ede7c22f

Observation 4ba87f36-9fb3-4c98-b0dc-8822888cfb0e · outbound

This paper cites Parallelizing Linear Transformers with the Delta Rule over Sequence Length.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.642183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.642183Z digest=sha256:ab68732f05fd20ebe311064bc30007179f019f1ceb8b103a448e1dad5d69f0b5

Observation 0ca922e9-e7f9-4072-98f4-df3502f21d96 · outbound

This paper cites The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.648381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.648381Z digest=sha256:4c3f706813bbc1d75b8c79ac3442308ab311d76241819851297b61df81bb839a

Observation 7bea6beb-642d-4d91-8fa1-7f41c625fcc8 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.653929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.653929Z digest=sha256:4537eea4c265cdcc7827569cf2f1a8dcfbf14a4f6b595c7af2684bb67d847253

Observation efe2d6c0-284e-4b5e-8d57-d9a70d40839b · outbound

This paper cites Arora et al.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Arora et al

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:35.221479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:52:33.659298Z digest=sha256:51a4f15de2aaff5f4953408928d8b4b172c32a99c46939f00bef5469c43f8d70

Observation cec37677-5009-4c35-9a1f-1d003b909a9d · outbound

This paper cites Linear Attention.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Linear Attention

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:35.204856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:52:33.664625Z digest=sha256:4d0984bcacb9f55df60b9f73dd67bfafd8d67607a9cbd8629ab1240f617c8c38

Observation 988a1212-ba25-431d-ae6f-15c8d3853831 · outbound

This paper cites Each layer of RetNet consists of a key, 16 Published as a conference paper at ICLR 2025 query, and value transformation, akin to linear attention.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Each layer of RetNet consists of a key, 16 Published as a conference paper at ICLR 2025 query, and value transformation, akin to linear attention

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:35.188224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:52:33.669461Z digest=sha256:747d86a88a99b5ebcfe18273ac4ab944b8fe7e2cb302fa0143e07751b9bcf27f

Observation ca99d2e5-b3ea-427a-8027-19f3f5abb175 · outbound

This paper cites Similar to Mamba (Gu & Dao, 2023), RetNet shares bt and ct across channels while assigning distinct ∆t for each channel when handling multi-channel inputs.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Similar to Mamba (Gu & Dao, 2023), RetNet shares bt and ct across channels while assigning distinct ∆t for each channel when handling multi-channel inputs

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:35.171024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:52:33.674872Z digest=sha256:6052e04b3699fc3bfa8f538a27c912f5135850e2bc35fd8b8caafbadc8e33b09

Observation 8be4f9ad-4423-4eff-b9f8-e7f7cebadf1f · outbound

This paper cites Its computational mechanism can be encompassed by our formulation in Eq.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Its computational mechanism can be encompassed by our formulation in Eq

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:35.153508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:52:33.679907Z digest=sha256:425fd2edbe4d599f42f38fadbdebb7604dbcefffc38d5386768b0942ba18a42e

Observation b2504547-851c-4e66-9a8b-e242a8c6d643 · outbound

This paper cites an unresolved cited work.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:52:35.137753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:52:33.684676Z digest=sha256:de88ac3c901e2907ebb2911f5f4fe6ab81f2482cefa1c0981643ee970a90ab56

Observation 9c86dd9e-3253-4af5-a355-f7c35df7859b · outbound

This paper cites an unresolved cited work.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:52:35.120957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:52:33.689634Z digest=sha256:1fdd1c45b9275e9f12fc4abd2a514b95ab6d281d419eea721ab41802210b3ff2

Observation b58c48aa-d56c-4ae0-a6a7-eb3cd43f4f80 · outbound

This paper cites In particular, the dimension of ht in Griffin is equal to the dimension of xt.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing In particular, the dimension of ht in Griffin is equal to the dimension of xt

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:35.099255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:52:33.694930Z digest=sha256:3a7606bdc922609582f85e1b3198c470a22b7e32cdd0342ef679df0136d93bef

Observation 2c4ae5a1-9287-43b2-be62-a1da2360be1e · outbound

This paper cites This design has quickly become a standard backbone for various SSMs (Gu & Dao, 2023; Beck et al., 2024).

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing This design has quickly become a standard backbone for various SSMs (Gu & Dao, 2023; Beck et al., 2024)

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:35.079906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:52:33.700455Z digest=sha256:d24abaed84a1a320fd84ee4890eccacbca320381d8fd14b32aea5fc2c17419c2

Observation 0ad87dfd-bb5e-4198-97de-4fbe5e2d05d6 · outbound

This paper cites attention.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing attention

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:35.062080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:52:33.708740Z digest=sha256:11ccec5e30f1ded500e205ff9adc7a27ebbcd60549f6cfb196d6e0c7240f102b

Observation f362ba20-f132-47cb-9635-ea2101b7fd0b · outbound

This paper cites We consider ϵ >0 small enough, thus, it is sufficient to consider the scenario when |ω| > Amax ≜ maxn∈[N ] |An,n|.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing We consider ϵ >0 small enough, thus, it is sufficient to consider the scenario when |ω| > Amax ≜ maxn∈[N ] |An,n|

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:35.044358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:52:33.714780Z digest=sha256:9120a619c6dbe929efce29f22c00815309351a65de8cdbfc5666ab0ef3a4f17c

Observation 60b5fc55-623d-445f-9c91-633d71208f2a · outbound

This paper cites Furthermore, let q = 1 − p.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Furthermore, let q = 1 − p

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:35.027899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:52:33.719862Z digest=sha256:b6fd613e2b2a8ff15a1afe50b947fa0d110aa16c341c9c6db1920bf8591736c1

Observation 0ff8a946-7cf4-42f8-86cc-c3260a97e1d5 · outbound

This paper cites less than.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing less than

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:35.010124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:52:33.724979Z digest=sha256:f2360d547fab15565da0cce8e3dc987211b293195339963721080bd4a3cb0a6f

Observation e199a764-7053-462f-a8e5-97568a4d0196 · outbound

This paper cites Needle in a Haystack.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Needle in a Haystack

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:34.875178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:52:33.731116Z digest=sha256:7697ce5ad9a0e5eaf2ceec52edba9138121c41a298a173f574acf21d0d2e8f79

Observation 92ff088c-6436-4679-92dd-0e678b6714e9 · outbound

This paper cites In SSMs, the class token must be positioned last to aggregate features from the entire sequence.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing In SSMs, the class token must be positioned last to aggregate features from the entire sequence

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:34.858842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:52:33.736023Z digest=sha256:414ba6946925fc889258a83e40393eb61695c17a92e8d4cdc9a84b5003fd9bc2

Observation 33ea289f-1c1b-45b4-92cf-b1dc038a563e · outbound

This paper cites In addition, our image classification setup differs from Tay et al.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing In addition, our image classification setup differs from Tay et al

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:34.840131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:52:33.741482Z digest=sha256:073fa0478bfe40a4b6b107b6eda90a671dd4674761faa44dc45a4e73dd90fcf3

Observation cef29079-19c9-46f9-b0d9-b9f933e1ea1d · outbound

This paper cites The models and training pipelines are built on Arora et al.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing The models and training pipelines are built on Arora et al

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:34.822951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:52:33.747493Z digest=sha256:7b66ee5cca41b641ac1f973761832c7018a95c4d05df2b49ef8e79ede9174c9e

Observation 688f22a6-b7d7-45c6-a3ba-1edfb924e787 · outbound

This paper cites For each curve in Fig.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing For each curve in Fig

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:34.803795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:52:33.752839Z digest=sha256:63c246d1eee796fbbb331098953522b9dbb62d9f35e946de808843a679535f03

Observation 6bf046c4-9441-4621-a89b-01a5b3a1aa9e · outbound

This paper cites The evaluation set is created by holding out a subset of 10M tokens from the training data.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing The evaluation set is created by holding out a subset of 10M tokens from the training data

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:34.781393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:52:33.757757Z digest=sha256:eb2c7148d35aeb16d51d3980cbe22e505ca3d84734520502b95394824137d096

Observation dc07023b-8caf-43bc-ad5c-0828f5bcd519 · outbound

This paper cites We test two block sizes {2048, 8192}.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing We test two block sizes {2048, 8192}

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:34.764751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:52:33.762879Z digest=sha256:da1c8286f74eb55375392aeffab8b9e94b359d01eb399620ecc8a9c04ddc39c2

Observation 682063cf-cf7c-4876-a815-9a389b0b523b · outbound

This paper cites 0 A −1000 # , At ≈.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing 0 A −1000 # , At ≈

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:34.748149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:52:33.767881Z digest=sha256:5a386d6a390c41575682e06f4aedc366875ec3ed125d44156e5569fc819f37a1

Observation 55b2cff7-9f66-465f-9b3f-8f1d9719e749 · outbound

This paper cites We consider 1-polarization mitigates locality most significantly, while deepening architecture only relieves recency mildly but deteriorates over-smoothing.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing We consider 1-polarization mitigates locality most significantly, while deepening architecture only relieves recency mildly but deteriorates over-smoothing

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:52:34.730788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:52:33.772825Z digest=sha256:249f704a6079c4cdf6093e1080425c68048511a01cbaf69e732b51aa65dc2002

Observation c812e534-2070-4e09-9dd2-d6c59daae234 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Training Compute-Optimal Large Language Models

Reference 1997

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.455866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.455866Z digest=sha256:83e34b7094cfca66a7e731bf66edbc2c980049b8c5ae858e7d8a566f65ff5ee8

Observation dbbac38c-b5cb-42f7-9537-2af0e4d67ef8 · outbound

This paper cites DataComp-LM: In search of the next generation of training sets for language models.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing DataComp-LM: In search of the next generation of training sets for language models

Reference 1998

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.483637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.483637Z digest=sha256:a6dd66769d704617c7823b05bed38b44972d477c384fe3d0e88a1368afd1165c

Observation 88595fe3-5e42-42e3-9ba2-de9ce0d109b8 · outbound

This paper cites On the Properties of Neural Machine Translation: Encoder-Decoder Approaches.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing On the Properties of Neural Machine Translation: Encoder-Decoder Approaches

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.384730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.384730Z digest=sha256:8bcf51fbfb1876074325d50eeb31927dedde318ad402572c1152761afde207dd

Observation 2e6c85fc-6a4a-4d53-8292-32b5122f7358 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.440979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.440979Z digest=sha256:5d0addc2697983a777dabf785e23bbdae961dac912759386a2c9307ca4dd37f3

Observation a86b3051-3238-40c5-b68c-7d4911d25845 · outbound

This paper cites What Makes Convolutional Models Great on Long Sequence Modeling?.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing What Makes Convolutional Models Great on Long Sequence Modeling?

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.489223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.489223Z digest=sha256:00a7a964c9159b86a79f04aeac76da1a5171ee0c505e90085d2caca6ca55d32e

Observation 694c31f0-3289-4e9a-9462-624096c29bfd · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.402204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.402204Z digest=sha256:8b1ab9499a7eb4e65426f2a59b5f60074cc918dd80ba8243a132407f46670877

Observation 02387a9e-f650-4752-9771-413844168932 · outbound

This paper cites Zoology: Measuring and Improving Recall in Efficient Language Models.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Zoology: Measuring and Improving Recall in Efficient Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.356788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.356788Z digest=sha256:d7617639d70acf08e9c06fbdf5659396eb5c312e1957b480ae791609e0abb3bc

Observation 30770a5f-7dcb-4956-a4ea-d4384b00656f · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.419780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.419780Z digest=sha256:ae4773e8ddb40b08c060a88474ba12f5155b5a63c903a58e8a490df06443dc28

Observation baced2b1-8fcc-4e09-adb2-d23b25c95fdb · outbound

This paper cites A mathematical perspective on Transformers.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing A mathematical perspective on Transformers

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.435966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.435966Z digest=sha256:f442208ff7a4f7848b10459221b7acaf19b800f0c4ac07d0a8e1572daf96ffb9

Observation 76ed0eb1-b777-4cee-a63a-b260004c07bf · outbound

This paper cites xLSTM: Extended Long Short-Term Memory.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing xLSTM: Extended Long Short-Term Memory

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.362373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.362373Z digest=sha256:9fcdec534cbea21d23cff3586d8331c517817b76a2edcb5e217ceedeb5391104

Observation 64837234-4575-4e0b-b303-0bafb2f189b9 · outbound

This paper cites On the Bottleneck of Graph Neural Networks and its Practical Implications.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing On the Bottleneck of Graph Neural Networks and its Practical Implications

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.351307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.351307Z digest=sha256:87a9e6675cf11c52f23080e5989c11d7df8847d0c13dae9eb84c4f452b611a7a

Pith citing papers

Observation 78136a9d-755a-4e54-a2d9-3a851978ff3c · inbound

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention cites this paper.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:54.578841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:54.578841Z digest=sha256:ddc92129b319497e9443fb62ec76796f399f94bae5bb527cdbc510c69042dc7a

Observation e6a00698-189c-4a82-abb3-014f0b22a923 · inbound

Echo: KV-Cache-Free Associative Recall with Spectral Koopman Operators cites this paper.

Echo: KV-Cache-Free Associative Recall with Spectral Koopman Operators Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:25:56.150123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T01:26:48.026538Z digest=sha256:135e37b328dfb5b8af0afccef751eb79d445274cb02c2465bc3f18e0ad2f83e2

Observation b59c1c1e-61d6-4d90-a7d9-14f50aeb86fb · inbound

The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale cites this paper.

The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T06:39:25.045565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:39:25.045565Z digest=sha256:60823a5fa0cca09d35134cb2cf6b7eddd5bbd6bde9eab09f22234a47d6701614

Observation 8b9a51b3-bdaf-4644-83ec-dc2a6ca9bd5e · inbound

The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale cites this paper.

The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T02:03:25.442040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:03:25.442040Z digest=sha256:cbb446b529edeee90c6ef3bf1a35268511bec8312a08a210511ab52a4ce5e6f3